GitHub's schedule event is explicitly best-effort: cron runs are delayed (community reports 29-60 min) or silently dropped during high-load windows, and the single most congested slot is the top of the hour (0 * * * *). Discussion #656 documented this hitting us in production (an hourly compliance sweep produced zero scheduled runs over a watched window; historical '05:00 UTC' daily runs fired 3-4 hours late). The danger is that 'successful but never ran' is indistinguishable from healthy, so backlogs sit untouched.
Owner decision (don-petry, #656): adopt Option 1 (move crons off the top of the hour) and Option 2 (lean into idempotency + slightly higher frequency at odd offsets) -- implement the offset as an organizational workflow standard, adjust all existing workflows across the org, and add follow-up analysis that reports the success rate of scheduled runs, incorporated into the fleet health daily check.
Scope of this epic: (1) offset every minute-0 cron in .github-private off the top of the hour; (2) codify the off-peak scheduling standard + a CI compliance check; (3) add scheduled-run reliability analysis to the daily Actions Fleet Monitor; (4) a human-led Phase-3 rollout that lifts the standard into the org source of truth (petry-projects/.github) and offsets scheduled workflows across the other org repos.
Deliberately NOT in scope (owner chose #1 + #2 only): Option #3 (external heartbeat -> workflow_dispatch) and Option #4 (event-chaining via workflow_run) backstops, and self-hosted runners. See #656 for the full options analysis.
Planned from idea discussion #656 by the BMAD Scrum Master initiative-planner. Inert until a maintainer adds initiative:auto.
GitHub's
scheduleevent is explicitly best-effort: cron runs are delayed (community reports 29-60 min) or silently dropped during high-load windows, and the single most congested slot is the top of the hour (0 * * * *). Discussion #656 documented this hitting us in production (an hourly compliance sweep produced zero scheduled runs over a watched window; historical '05:00 UTC' daily runs fired 3-4 hours late). The danger is that 'successful but never ran' is indistinguishable from healthy, so backlogs sit untouched.Owner decision (don-petry, #656): adopt Option 1 (move crons off the top of the hour) and Option 2 (lean into idempotency + slightly higher frequency at odd offsets) -- implement the offset as an organizational workflow standard, adjust all existing workflows across the org, and add follow-up analysis that reports the success rate of scheduled runs, incorporated into the fleet health daily check.
Scope of this epic: (1) offset every minute-0 cron in .github-private off the top of the hour; (2) codify the off-peak scheduling standard + a CI compliance check; (3) add scheduled-run reliability analysis to the daily Actions Fleet Monitor; (4) a human-led Phase-3 rollout that lifts the standard into the org source of truth (petry-projects/.github) and offsets scheduled workflows across the other org repos.
Deliberately NOT in scope (owner chose #1 + #2 only): Option #3 (external heartbeat -> workflow_dispatch) and Option #4 (event-chaining via workflow_run) backstops, and self-hosted runners. See #656 for the full options analysis.
Planned from idea discussion #656 by the BMAD Scrum Master initiative-planner. Inert until a maintainer adds
initiative:auto.