You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
ADR-0007 collapse: standards-deploy RESURRECTS three collapsed stubs (double dispatch), and the agent rate-limit guard FAILS OPEN on a collapsed repo #1226
Two more consumers in this repo resolve agent identity in a way the ADR-0007 collapse breaks. Both bite the pilot repo (petry-projects/markets, petry-projects/.github-private#1729, companion petry-projects/markets#506). Sibling of #1224 (canary-rollout), same root cause: identity keyed on per-role workflow name/path, which the ingress replaces with jobs.
Found by sweeping both org-infra repos for identity consumers. Ten scripts in .github-private were already made ingress-aware (#1724/#1725/#1726/#1727); nothing in this repo was.
Gap 1 — deploy-standard-workflows.sh re-creates the deleted stubs (double dispatch)
SKIP_REPOS=(".github" ".github-private"), so markets is a deploy target. DEPLOYABLE_WORKFLOWS contains three of the five stubs the collapse deletes:
After markets lands the collapse, the next standards sweep deploys those three templates back from standards/workflows/. The repo then has bothagent-ingress.yml and the per-role stubs subscribing to the same events — so dev-lead, pr-review-mention and pr-auto-review each dispatch twice per event. That is double token spend and two racing claims on the same PR, not a cosmetic drift.
.github-private's #1726 taught fleet_stub_drift / fleet_stub_remediate never to resurrect a collapsed stub. This script is a second, unpatched resurrector in a different repo.
Secondary: the header notes that a workflow which is universal-required and deployable but neither opted-in nor self-managed is "drift: the sweep skips it and the audit flags it required-but-missing forever". A collapsed repo is exactly that state for these three — so the compliance audit will flag markets non-compliant permanently.
Acceptance criteria
A repo that has .github/workflows/agent-ingress.yml with a job for role R does not receive standards/workflows/R.yml from the sweep.
Detection reads the repo's actual state (ingress present + role job present), not a hardcoded repo list — fan-out is incremental, so the set of collapsed repos changes.
The compliance audit treats a role served by an ingress job as present, not required-but-missing.
Tests: a collapsed repo is not re-seeded for a collapsed role; a non-collapsed repo still receives the stub; a repo with an ingress that lacks a given role job does still receive that role's stub.
Gap 2 — the agent rate-limit guard fails OPEN (no throttle)
runs="$(gh run list --workflow "$agent_type" --json status --limit 1000 2>/dev/null || true)"if [ -z"$runs" ];then
arl_log "warning: run enumeration for '${agent_type}' returned no data (treating concurrency as 0)"
scripts/agent-rate-limit-gate.sh → argate_fetch_runs() does the same and degrades to [] with "treating as empty (degraded)".
On a collapsed repo there is no workflow named for the agent, so enumeration returns nothing and concurrency is counted as 0. The guard then never throttles that repo. This is worse than #1224: canary-rollout loses evidence, while this one actively permits runaway concurrency — and the fail-open is deliberate for transient API errors, so a permanent absence is indistinguishable from a blip.
A permanent "no such workflow" is distinguished from a transient API failure. The former must not silently read as zero — surface it, and fail closed (or explicitly report unresolved) rather than allowing.
AGENT_RATE_LIMITS_ORG_REPOS tallying works across a mixed fleet of collapsed and non-collapsed repos.
Tests: collapsed repo contributes its in-flight ingress role jobs to the tally; a transient failure still degrades permissively with its existing warning; a permanent missing-workflow does not read as 0.
Sequencing
Gap 1 must land beforemarkets merges its collapse — a resurrected stub double-dispatches immediately, and the sweep is scheduled. Gap 2 should land before too, since the pilot is a live agent repo. markets#506 is deliberately unlabelled pending ADR-0010 acceptance, so there is time.
May be split into two issues if that suits implementation better; they share only the root cause.
Summary
Two more consumers in this repo resolve agent identity in a way the ADR-0007 collapse breaks. Both bite the pilot repo (
petry-projects/markets, petry-projects/.github-private#1729, companion petry-projects/markets#506). Sibling of #1224 (canary-rollout), same root cause: identity keyed on per-role workflow name/path, which the ingress replaces with jobs.Found by sweeping both org-infra repos for identity consumers. Ten scripts in
.github-privatewere already made ingress-aware (#1724/#1725/#1726/#1727); nothing in this repo was.Gap 1 —
deploy-standard-workflows.shre-creates the deleted stubs (double dispatch)SKIP_REPOS=(".github" ".github-private"), somarketsis a deploy target.DEPLOYABLE_WORKFLOWScontains three of the five stubs the collapse deletes:After
marketslands the collapse, the next standards sweep deploys those three templates back fromstandards/workflows/. The repo then has bothagent-ingress.ymland the per-role stubs subscribing to the same events — sodev-lead,pr-review-mentionandpr-auto-revieweach dispatch twice per event. That is double token spend and two racing claims on the same PR, not a cosmetic drift..github-private's #1726 taughtfleet_stub_drift/fleet_stub_remediatenever to resurrect a collapsed stub. This script is a second, unpatched resurrector in a different repo.Secondary: the header notes that a workflow which is universal-required and deployable but neither opted-in nor self-managed is "drift: the sweep skips it and the audit flags it required-but-missing forever". A collapsed repo is exactly that state for these three — so the compliance audit will flag
marketsnon-compliant permanently.Acceptance criteria
.github/workflows/agent-ingress.ymlwith a job for role R does not receivestandards/workflows/R.ymlfrom the sweep.Gap 2 — the agent rate-limit guard fails OPEN (no throttle)
scripts/lib/agent-rate-limit.sh→arl_count_concurrent_runs():scripts/agent-rate-limit-gate.sh→argate_fetch_runs()does the same and degrades to[]with "treating as empty (degraded)".On a collapsed repo there is no workflow named for the agent, so enumeration returns nothing and concurrency is counted as 0. The guard then never throttles that repo. This is worse than #1224: canary-rollout loses evidence, while this one actively permits runaway concurrency — and the fail-open is deliberate for transient API errors, so a permanent absence is indistinguishable from a blip.
Acceptance criteria
.github-private'srun-attribution.sh.AGENT_RATE_LIMITS_ORG_REPOStallying works across a mixed fleet of collapsed and non-collapsed repos.Sequencing
Gap 1 must land before
marketsmerges its collapse — a resurrected stub double-dispatches immediately, and the sweep is scheduled. Gap 2 should land before too, since the pilot is a live agent repo. markets#506 is deliberately unlabelled pending ADR-0010 acceptance, so there is time.May be split into two issues if that suits implementation better; they share only the root cause.