Skip to content

apply-repo-settings can never reach the fleet: stub not deployable, ring tags missing, canary deadlocked on a vendored script #1045

Description

@don-petry

Summary

The apply-repo-settings reusable already does everything the org needs — it runs apply-rulesets.sh and apply-repo-settings.sh under the classic admin PAT. No repo in the fleet can reach it. Four independent gaps sit between the working reusable and the repos that need it, and they form a cycle.

This is the root cause of the recurring ruleset-drift-pr-quality-* findings across broodly, TalkTerm, google-app-scripts and bmad-bgreat-suite.


The chain

Gap 1 — the stub is not deployable. standards/workflows/apply-repo-settings.yml exists, but apply-repo-settings.yml is absent from DEPLOYABLE_WORKFLOWS in scripts/deploy-standard-workflows.sh:

pr-review-mention.yml  dev-lead.yml  agent-shield.yml  auto-rebase.yml
dependabot-automerge.yml  dependabot-rebase.yml  dependency-audit.yml
add-to-project.yml  initiative-driver.yml  pr-auto-review.yml  feature-ideation.yml

The template is never deployed to anyone. The weekly standards-deploy.yml sweep cannot pick it up.

Gap 2 — the channel tags the stub needs do not exist. ring_tier_for_repo puts TalkTerm and bmad-bgreat-suite in ring1, the rest of the fleet at stable. Present tags:

next  ring0  v0.0.1–v0.0.5  v0.1.0  v1-next  v1-ring0  v1.0.0

apply-repo-settings/v1-ring1 and apply-repo-settings/v1-stable are both missing. The stub pins @apply-repo-settings/v1-stable. Deploying it today would pin a nonexistent ref — a fleet-wide startup_failure.

Gap 3 — the canary will not promote those tags. Blocker #1042 holds apply-repo-settings at ring0->ring1, citing TalkTerm run 33446111321.

Gap 4 — that run fails because the repo vendors a stale script. TalkTerm, markets and bmad-bgreat-suite each carry their own .github/workflows/apply-repo-settings.yml doing run: bash scripts/apply-repo-settings.sh with GH_TOKEN: ${{ secrets.GH_PAT_WORKFLOWS }}:

TalkTerm scripts/apply-repo-settings.sh   320 lines   403-handling: 0   apply-rulesets refs: 0
.github  scripts/apply-repo-settings.sh   497 lines   403-handling: 4

The vendored copy has no 403 tolerance and never invokes apply-rulesets.sh. It fails on the check-suites PATCH every run — 30/30 across the three repos — and performs no ruleset convergence at all.

The cycle: canary held → TalkTerm fails → vendored stale script → cannot adopt the stub → stub pins a tag that does not exist → tag not cut → canary held.


Ordered unblock — the order is load-bearing

Doing these out of order breaks the fleet. In particular Step 4 before Step 3 pins every repo to a nonexistent tag.

Step 1 — break the cycle at the vendored repos (maintainer)

Make TalkTerm, markets and bmad-bgreat-suite stop failing. Either sync canonical scripts/apply-repo-settings.sh into each, or remove the vendored workflow + script pending stub adoption. Removing is the smaller change and loses nothing: the vendored path has never successfully applied anything.

Verify: a run of apply-repo-settings.yml in TalkTerm reaches success, or the workflow is gone.

Step 2 — let the canary clear

#1042 auto-closes when the gate passes; it regenerates each run and must not be edited by hand.

Step 3 — cut the channel tags (maintainer)

apply-repo-settings/v1-ring1 and apply-repo-settings/v1-stable, via scripts/cut-release.sh <agent> <ver> --channel <tier> --promote --push. Dry-run first.

Verify: both refs resolve.

Step 4 — add the stub to DEPLOYABLE_WORKFLOWS (agent-safe, once Step 3 is done)

One line in scripts/deploy-standard-workflows.sh. Gate a test on it so the list and standards/workflows/ cannot drift apart again — every template that is meant to deploy should be listed, and a template that is deliberately excluded (ci.yml, sonarcloud.yml) should say so in a comment.

Step 5 — deploy and confirm the drift actually heals

Run standards-deploy.yml, let the stub PRs merge, then confirm convergence against live state, not CI colour:

for r in broodly TalkTerm google-app-scripts bmad-bgreat-suite; do
  echo -n "$r "
  gh api repos/petry-projects/$r/rulesets --jq '.[]|select(.name=="pr-quality")|.id' \
  | xargs -I{} gh api repos/petry-projects/$r/rulesets/{} \
      --jq '.rules[]|select(.type=="pull_request")|.parameters
            |"codeowner=\(.require_code_owner_review) lastpush=\(.require_last_push_approval) dismissStale=\(.dismiss_stale_reviews_on_push)"'
done

All three must read true everywhere. Current state — broodly false/false/false, TalkTerm true/false/false, google-app-scripts true/false/false, bmad-bgreat-suite true/true/false.


Not released to dev-lead

Only Step 4 is safe for an agent, and only after Step 3. Steps 1 and 3 need maintainer judgement and admin credentials; an agent doing Step 4 early would pin the fleet to a missing ref. Label deliberately withheld.

Related: #1038 (the reconciler abort itself — PR #1044 in review), #1036 (findings closed while live), #1037 (routing decision), .github-private#1620 (zero-diff PRs), canary blocker #1042.

Filed from a manual compliance sweep, 2026-09-01.


ORDERED UNBLOCK — REVISED 2026-09-07. Supersedes the five steps above.

Two of the four gaps in the original report have closed, one was misdiagnosed, and the sequence has changed. Use this list, not the one above.

Gap status

Gap Original claim Status
1 stub absent from DEPLOYABLE_WORKFLOWS Still open, but it was never the cause of the pin failure — see below
2 ring tags missing Closed — v1-ring1 created 2026-09-02 (by hand; the automation defect is #1065)
3 canary deadlocked Closed — #1042 cleared; next→ring0 and ring0→ring1 both promoted
4 vendored stale scripts Neutralised — all three set disabled_manually 2026-09-01

The misdiagnosis, corrected

I attributed the TalkTerm#489 / bmad#463 mis-pin to --workflow bypassing DEPLOYABLE_WORKFLOWS. That was wrong. The cause was RING_REUSABLES in scripts/lib/ring-pins.sh omitting apply-repo-settings, so ring_is_ring_reusable returned false, emit_ref_for returned empty, and the template deployed verbatim with its hardcoded @apply-repo-settings/v1-stable. That would have happened through the normal fleet sweep too.

Fixed in #1088 (merged): RING_REUSABLES now equals the registry (16/16), and the deploy refuses any computed ref that does not resolve on the host.

Revised sequence

Step A — deploy the stub to the ring1 repos. Now safe: emit_ref_for computes @apply-repo-settings/v1-ring1 (exists, edbac2551e4c), and #1088's assert-exists check refuses anything that does not resolve.

gh workflow run standards-deploy.yml -R petry-projects/.github \
  -f target_repo=TalkTerm -f target_workflow=apply-repo-settings.yml -f dry_run=true

Verify the PR pins v1-ring1 before merging — read the file on the branch, do not infer it. Repeat for bmad-bgreat-suite. Per #1038 AC5″, the same change should remove each repo's vendored caller and local scripts/apply-repo-settings.sh.

markets is included at this step by maintainer decision (2026-09-07): standards/canary-rings.json is authoritative, so markets is ring1. ring_tier_for_repo in scripts/lib/ring-pins.sh still maps it to stable and must be corrected first, or markets will compute the non-existent v1-stable and be refused.

Step B — feed the gate. Manual workflow_dispatch of apply-repo-settings.yml in the ring1 repos, per the registry note: "ring members feed the gate via a manual workflow_dispatch after the stub lands. Do NOT paper over a starved gate with --override."

Success criterion is the live read, not the run's colour:

gh api repos/petry-projects/TalkTerm/rulesets --jq '.[]|select(.name=="pr-quality")|.id' \
| xargs -I{} gh api repos/petry-projects/TalkTerm/rulesets/{} \
  --jq '.rules[]|select(.type=="pull_request")|.parameters
        |{require_code_owner_review, require_last_push_approval, dismiss_stale_reviews_on_push}'

Step C — promote ring1→stable (dwell 12h, sample_min: 1, now satisfiable). Creates stable / v1-stable.

Step D — add apply-repo-settings.yml to DEPLOYABLE_WORKFLOWS (Gap 1, still open) so the weekly sweep carries the stub to the rest of the fleet. Only after Step C — before it, stable-tier repos compute the non-existent v1-stable and #1088 will refuse them.

Step E — confirm convergence across broodly, google-app-scripts, bmad-bgreat-suite and markets with the live read above. All three parameters true everywhere.

Live state 2026-09-10 — Step C will not fire on its own

Verified against the scheduled promoter, not inferred. apply-repo-settings/v1-stable still does
not exist
, yet the 05:07Z Canary Rollout run (34439789612, conclusion: success) logs:

──────── agent: apply-repo-settings ────────
nothing to promote — apply-repo-settings is fully rolled out.

The promoter reads the bare channel set (next/ring0/ring1/stable, all present) and
concludes rollout is complete, while consumers pin the major-scoped v1-stable, which was
never created. So Step C is not merely pending — the automation believes it is already done and
will keep reporting success every 4h without ever creating the tag. This is #1065 exactly, one
tier further along than the v1-ring1 case that was patched by hand on 2026-09-02.

persona-mention is the same defect, worse: it has ring0, ring1, v1-next, v1-ring0,
v1-ring1 and releases through v1.5.0, but no next, no stable, no v1-stable at all.

The fix is already in flight: PR #1097 implements #1065. Landing it is what unblocks Step C
here, and #1045 in turn unblocks .github-private#1740. Until #1097 merges, the only way to
complete Step C is another hand-created tag — which is what the standing caution below warns
against treating as evidence the pipeline works.

For the record, one thing this is not: a "no release was ever cut" problem. Every registered
ring agent has channel tags; they live in the agent's host repo (cross-repo reusables in
petry-projects/.github, dev-lead/pr-review/ci-failure-analyst in .github-private), so an
inventory of only one repo will appear to show 14 agents with no tags. The genuine gaps are the two
partially-seeded agents named above.

Standing caution

TalkTerm currently reads true/true/true, but that came from two hand-applied edits by don-petry (ruleset history: 2026-09-04T16:56Z, 2026-09-06T17:03Z), not from the reconciler. A compliant end state is not evidence the pipeline works. Step B must demonstrate convergence by the reusable, or the drift returns and the audit re-files it.


Step C done, 2026-09-13

apply-repo-settings/v1-stable now exists at a781ad9d5e50 (release v1.4.0), equal to v1-ring1. It was created by a gated canary-rollout promote apply-repo-settings (run 34738419511, 04:58Z, no override) after PR #1097 (#1065) fixed channel-major resolution. The dry run (34737895599) showed exactly that one move.

Step D is filed as #1128 (add apply-repo-settings.yml to DEPLOYABLE_WORKFLOWS). It is held until the org Actions queue drains (#1126). Step E (live-read convergence) stays here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dev-lead:hands-offOpt an item out of the dev-lead persona automation entirely

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions