Replies: 1 comment
|
📋 Initiative planned by the BMAD Scrum Master (Bob). Epic #1723 — Implement ADR-0007: collapse per-role caller stubs into one agent-ingress.yml per repo 6 stories created (inert — labelled
Open questions for review:
Review the epic and its sub-issue DAG, adjust as needed, then add |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Status: approved. The decision is recorded as ADR-0007 (
accepted) in#1710. This Discussion is open for the deeper review — the ADR
is append-only, so anything that changes direction here becomes ADR-0008 superseding 0007,
not an edit.
The question
Could an org- or repo-level webhook invoke our persona / agent workflows directly across
the fleet, replacing the per-repo thin caller stubs (ADR-0001)? Specifically: can we trigger
agents more easily across all repos from one place; are all our action types webhookable; and
do authN / authZ / secrets work the same way as they do inside a workflow?
The decided approach
Collapse each consumer repo's per-role caller stubs into one
.github/workflows/agent-ingress.yml. Do not stand up an off-GitHub webhook receiver at thecurrent fleet size (8 consumer repos per
scripts/lib/consumer-manifest.json, 9–13 pinnedrefs each).
The reframe that decided it: a stub is doing three separable jobs — ingress (subscribe to
the event), enrollment (this repo opted in), and the ring pin (ADR-0002). A webhook
replaces only the first, and the first is the cheapest to fix without one.
Answers to the three questions
1. Can webhooks trigger agents more easily across all repos? Not by themselves — a
webhook is not an executor. An org webhook delivers JSON to an endpoint we would have to
host; it cannot start a workflow. It can only call back into GitHub via
repository_dispatch, or run the agent on compute we own. And we already run hub-and-spoke:persona-mention.yml→repository_dispatch→persona-runner.ymlserves all nine personasfrom one runtime (Bridge A,
docs/agentic-interaction-model.md§5). The centralization isalready in hand. What is not consolidated is the ingress file — and consolidating a
doorbell needs no receiver.
2. Are all our action types webhookable? No — roughly 40% at best. From the CI-verified
§4 classification table (35 agentic rows):
initiative-planner.ymlisworkflow_dispatch-only)Gaps beyond the count:
pull_request_target(dependabot-automerge,add-to-project) is anActions-only construct — a receiver sees plain
pull_requestand inherits full secretaccess on fork PRs, which Actions currently withholds;
check_suite:requested/rerequestedis GitHub Apps only, not org webhooks; required checks (
holdout-guard,duplicate-decl-gate,test-deletion-guard) must stay in-repo because a hub run posts nocheck on the target PR unless we synthesize commit statuses; and
workflow_callis not anevent at all, so the reusable tier is unaffected either way.
3. Do authN / authZ / secrets work the same? No — four deltas.
GITHUB_TOKEN. Today it is per-run, repo-scoped, auto-expiring, and shaped by thejob's
permissions:block. A GitHub App installation token is genuinely better scopedthan a PAT and would fix the fine-grained-PAT single-owner limit §5 already flags for
cross-org — but the read/write split in
persona-runner-reusable.yml(the agent holdsgithub.tokenand physically cannot post; the workflow holds the PAT and posts) isenforced by the runtime today and would have to be hand-rebuilt.
secrets: inherithas noanalogue;
CLAUDE_CODE_OAUTH_TOKEN,GH_PAT_DON_PETRY,GOOGLE_API_KEYwould need asecond store and a second rotation path. Keep execution in Actions and this delta is zero.
GITHUB_TOKENtriggernothing — is a free loop-guard. A receiver has none: every comment it posts is a live
event it receives back. After
docs/postmortems/2026-06-pr-860-runaway.md(1,481acknowledgements in 4.5 hours), the marker guard would go from belt-and-braces to sole
defense.
don-petry(runtime.identity.credential, checkedby
verify-persona-identity.sh);donpetry-botis deliberately a separate review-onlyaccount. An App changes the actor on every comment, approval and push, and
approval-by-actor is already load-bearing under the last-push rule. That is its own
change, not a rider on an ingress change.
What ADR-0007 actually binds
on:block is the union of the collapsed roles' triggers.
if:may act as an event filter and only as an event filter —github.event_name,github.event.action, payload fields as a pure predicate. Anythingreaching for repo state is logic and belongs in the reusable. One job per role is required,
not stylistic: each role needs its own
permissions:block and its own channel pin, and asingle dispatching job would have to grant the union of every role's permissions.
tags become per-job pins in the ingress.
pull_request_targetroles keep their own stub; Class 2/3 timers keeptheirs; non-agentic CI and gate workflows are out of scope.
silent on ingress. That silence was the actual finding.
Costs we accepted with our eyes open
agent-ingress.ymldisables everyevent-driven agent in that repo, where today a broken
dev-lead.ymlleaves pr-reviewalive.
caller_stub_freeze.shbecomes correspondingly load-bearing.by it. Same trade ADR-0006 made for personas, made knowingly a second time.
run_workflow;collapsing many workflow names into one reproduces the attribution problem ADR-0006 fact 3
names for personas. Job names must carry the role, and monitors keying on workflow name
must retarget to job level in the same change.
paths:filters move down a tier (dependency-advisory.yml), so union subscriptionstarts more runs with mostly-skipped jobs — free in minutes, not in log legibility.
first collapse, or protection blocks merges on a check that no longer exists.
Not foreclosed
The ADR states the reopening triggers for the webhook option: the fleet outgrowing a stub
fanout (roughly, when the eight-file edit is itself the bottleneck), a need for events Actions
cannot subscribe to inline, or cross-org routing that fine-grained PATs cannot span. Reopening
means a new ADR, because it moves the boundary this one just set.
Delivery scope (refined 2026-09-08)
Scope of this initiative — S1: make the ingress collapsible, then pilot it once
This initiative delivers the enabling work plus a single-repo pilot. Fleet-wide rollout
across all 8 consumer repos is deliberately a separate, later initiative gated on the
pilot's evidence.
agent-ingress.ymltemplate + router shape. Author the template under thestandards path the other org templates use, with one job per role, each carrying its own
permissions:block and its own ADR-0002 channel pin. No repo adopts it in this story.if:-as-event-filter checker. ADR-0007 loosens ADR-0001 in exactly one place, sothat loosening gets an automated check: a job-level
if:in a caller stub may referenceonly
github.event_name,github.event.action, and event-payload fields used as a purepredicate. Anything reaching for repo state fails the check. Pure-logic + bats per
ADR-0004.
validate-caller-inputsabout multi-ref forwarding. The [#1052 A] CI guard: forwarded caller inputs ⊆ declared inputs at the pinned @ref (+ #1034 regression) #1253 check assumes onestub forwards to one pinned reusable; the ingress forwards to several. Extend it to
validate every
uses:/with:pair in one file against its own pinned ref's declaredworkflow_call.inputs.caller_stub_freeze.shto the ingress path. One file, one byte-identitybaseline. The freeze becomes more load-bearing under ADR-0007, not less.
fleet_stub_drift.shandtemplate_stub_drift.sh. Both key on per-role stubpaths today; they must classify ALIGNED / DRIFTED / MISSING against the single ingress
path without losing per-role visibility in the report.
fleet_monitor.sh,pr_review_health.shand the health scans sample byrun_workflow; collapsing manyworkflow names into one reproduces the attribution problem ADR-0006 fact 3 names for
personas. Job names must carry the role and the monitors must key on them.
petry-projects/.github-privatebehind one ingress. The dogfood repoonly. Includes the branch-protection required-check renames, which must land in the same
change or protection blocks merges on a check name that no longer exists. Excludes
pull_request_targetroles and Class 2/3 timers per the ADR carve-outs.table in
docs/agentic-interaction-model.mdandinteraction-contracts/*.ymlmust movein the same PR as the pilot collapse, or
validate-interaction-modelfails.Initiative success metric
Primary: in the pilot repo, the count of Class 1 agentic caller-stub files drops from its
current value to 1, and every collapsed role demonstrably fires on its subscribed event
at least once after the collapse — evidence is one linked run per collapsed role, recorded
on the pilot story before it closes. A collapse that reduces file count but silently stops a
role from firing is a failure, not a partial success.
Secondary: zero net change in required-check outcomes on the pilot repo's next 10 merged
PRs (no check lost, no check newly blocking).
Cost cap
≤ 60 agent runs and ≤ USD 40 of engine spend across the whole initiative. The pilot is
one repo; if the pilot alone exceeds half that budget, stop and re-scope rather than
continuing to the fleet-rollout initiative.
Prerequisites (known, not open questions)
it as the governing decision, so it is a hard
blocked_byfor story 1.story 7 cannot be done by an agent alone and needs a human step.
Open questions — advisory, do NOT bake into acceptance criteria
These are genuinely unresolved. Per the planner's rubric they belong here, not inside an AC,
and they are non-blocking — the S1 scope above is buildable without answering them.
The cost is total per-repo blast radius, which is the cost I am least comfortable with.
Story 7's pilot is the evidence that should settle it — if the answer changes, that is
ADR-0008 superseding 0007, not an edit.
channel-skew defect (feat: implement issue #1031 — fix(#844): plumb LSP_PILOT_ENABLED env + decouple capture from LSP wiring so the pilot A/B runs in CI #1034 / epic fix(ci): prevent channel-skew workflow breakage — guard forwarded inputs vs pinned refs + codify the pinning rules (post-#1034) #1052) with more pinned surfaces per file? Story 3 is
where this would first show up.
Out of S1 scope by construction — this is the fleet-rollout initiative's problem, and
story 7 should record what the single-repo rename actually cost as input to it.
Explicitly out of scope
new ADR, not a story here.
feature-ideation.yml's broken tooling checkout — real, tracked separately.it does not ride along on an ingress change.
All reactions