Skip to content

canary-rollout promote targets whatever next is when the run executes, not the commit its dry run showed — a queued override can force-move the wrong ring to an unverified commit #1223

Description

@don-petry

Summary

canary-rollout.yml -f command=promote has no way to say which commit or which ring the operator approved. It resolves the candidate as "whatever <agent>/next points at" and the ring as "the first ring not on that commit", both when the run executes. A dry run therefore verifies nothing about the real run that follows it: if next is re-cut in between, the real run silently does something different, and with override=true it force-moves a tag regardless of gate state.

The gap between the two is not small. A dispatched run can sit queued for many minutes behind org runner contention, and the scheduled tick re-cuts next every 4 hours whenever a watched path changed on the host's main.

Mechanism

Permalinks at cd0b1675:

  • scripts/canary-rollout.sh L1008-L1011: _frontier_state reads cand="$(channel_commit "$agent" next)" and picks the frontier as the first ring whose commit differs from it.
  • scripts/canary-rollout.sh L1246 and L1301-L1304: cmd_promote calls it fresh, then either prints the [DRY-RUN] would: … sha=$cand line or moves the tag to that same freshly-read $cand.
  • .github/workflows/canary-rollout.yml L64-L118: the dispatch inputs are command, agent, override, confirm, dry_run, plus ring and to, which only rollback uses. rollback is pinned to an immutable release; promote is pinned to nothing.
  • The manual promote runs in concurrency group canary-rollout-<agent> and the scheduled run in canary-rollout-fleet, so the scheduled autocut and a manual promote for the same agent do not serialize.

So after a re-cut, every ring is "not on the candidate" again and the frontier snaps back to ring0. A run the operator dispatched to move stable to commit A instead moves ring0 to commit B, exits green, and logs promoted.

Evidence (dev-lead, 2026-10-01, one-off override release to v139.39.0 = 51cb28ee)

All times UTC.

Time Event
20:01:43 petry-projects/.github-private#1977 merges to main (fb99072c). It changes scripts/template_stub_drift.sh; scripts/ is a default watched agent_ref_paths entry, so the next scheduled autocut will re-cut dev-lead/v139-next.
20:04:22 Dry run 36918476004 prints [DRY-RUN] would: gh api PATCH …/tags/dev-lead/v139-stable sha=51cb28ee….
20:04:35 Real run 36918930731 dispatched (override=true, dry_run=false). It stays queued: the #1977 merge set off a rebase storm, with 326 runs queued in .github-private at 20:11 and 78-86 in .github.
20:25:05 The run finally executes and logs promoted dev-lead/v139-stable -> 51cb28ee6c36.
20:33 Next 33 */4 * * * scheduled tick is due.

The run was queued for about 20 minutes and executed about 8 minutes before the tick that would re-cut next. Had the order been reversed, the run would have force-moved dev-lead/v139-ring0 to fb99072c, a commit that was never dry-run, never soaked and never looked at, while stable stayed where it was and the run reported success.

It came out right only because the operator noticed the merge, polled the next tag every 15 seconds and was ready to cancel the queued run. Nothing in the tooling provides that guard.

The same exposure exists without any queueing: there is always a gap between a dry run and the real dispatch, and the 2026-10-01 release needed three dry-run/real pairs (one per ring).

Acceptance criteria

  1. promote accepts an expected candidate commit and an expected ring, as a CLI flag and as workflow_dispatch inputs (for example expect_sha and expect_ring). A prefix of at least 8 hex characters is accepted for the commit.
  2. When either is supplied and the live candidate or the computed frontier differs, promote moves nothing, prints an ::error:: naming the expected and actual values, and exits non-zero so the run is red.
  3. override=true with dry_run=false and no expected commit is refused with a clear error. An override bypasses the gate, so it must at least name what it is overriding for.
  4. The dry-run output prints the exact follow-up dispatch command, including the expected commit and ring it just computed, so the real run can be copy-pasted.
  5. promote-all (the scheduled arm) is unchanged: it runs autocut and promotion in one job by design.
  6. tests/canary_rollout.bats covers: candidate changed between dry run and real run (refuses, no tag write), frontier changed (refuses, no tag write), both match (promotes), and override without an expected commit (refuses).
  7. The canary runbook's manual-promotion section documents the dry-run then pinned-real-run sequence.

Related, not duplicates

  • .github#1118 "canary-rollout evaluates only the newest candidate — every new cut cancels a pending ring1→stable human confirmation": the same "newest next wins" model, seen from the gate's side. Fixing it may change how the candidate is chosen, but an operator still needs to pin what they approved.
  • .github#1175 "Canary Rollout times out: 6 of the last 7 scheduled runs hit timeout-minutes: 30": the same runner contention that widened the window here.
  • .github#1177 "Canary Rollout operator gaps: no allow_pre_existing dispatch input, 'all agents promoted' summary despite BLOCKED gates": other missing dispatch inputs and misleading output on the same workflow.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions