Workstream of melodic-software/github-iac#78. Decision record from the 2026-07-16 capacity incident (228 queued jobs, 2h+ backlog, oldest job queued 11:59Z unreleased at 14:00Z).
Problem
The per-run fan-out shape optimizes for the wrong constraint. It was built to dodge GitHub-hosted per-job 1-minute billing rounding; on the self-hosted fleet the scarce resource is worker slots, and each micro-job pays ~20-30s of container spin-up + JIT registration + checkout for 7-60s of work. Measured today: ~40-50% of fleet throughput consumed by select-runner preflights alone; dotfiles spends 16 selector jobs + 16 lanes per run (37 jobs, legacy Shape B); standards ~30 jobs/run; main-push bursts (4 pushes in 15s from distribution sync) enqueue 4 full runs with zero collapse.
Target shape (research-verified 2026-07-16, official sources)
Per consumer repo:
- One selector (Shape A, already live in medley/claude-code-plugins) — completes the epic-mandated single preflight; migrate dotfiles off Shape B.
- Consolidated lanes instead of micro-jobs: a
hygiene lane running all lint/hygiene checks as steps in one worker, using continue-on-error: true + id: per check step and a final aggregation step failing on any steps.<id>.outcome == 'failure' (run-everything-fail-at-end; outcome/conclusion semantics per contexts reference). Language/build lanes only where toolchain or resource isolation genuinely differs (e.g. dotnet build/test). standards' fixture matrix folds into a fixtures lane.
- Parallel steps where I/O-bound (changelog 2026-06-25):
background: true/wait/parallel: — runner v2.335.1 (current fleet image) has the engine; max 10 concurrent background steps; a failing background step surfaces at the join; NOT usable inside composite actions. On 2-CPU workers this buys log clarity and overlap for network-bound checks (lychee), not CPU throughput — default to sequential steps.
ci-status fan-in unchanged as the single required check (job-level check granularity; a job skipped by if: reports Success for required-check purposes — keep in-workflow conditionals, never workflow-level paths: filters on required workflows).
- Main-push burst collapse: keep PR groups (
cancel-in-progress: true) as-is; give main pushes group: <workflow>-${{ github.ref }} with cancel-in-progress: false — the documented default (queue: single) then cancels the superseded pending run while the in-progress run finishes (concurrency docs). Do NOT use queue: max here (it serializes every push — deploy semantics, not lint semantics).
- Per-check visibility inside a lane: per-step collapsible logs +
$GITHUB_STEP_SUMMARY (1MiB/step, 20 summaries/job). Annotation counts (10/step, 50/job) are community-reported only — do not design around them.
Sizing interplay
Consolidated lanes get a larger default worker (2 CPU / 4GiB — CPU parity with hosted private-repo runners at 2 vCPU/8GB, RAM bounded by the per-host VM budget; provisioning#133/#134). Fewer, beefier workers replace many 2GiB micro-workers; setup-* toolchain installs amortize once per lane instead of once per micro-job, with actions/cache for dependencies (image already ships zstd for hosted cache parity).
Sequencing
Reusable workflow/template changes here → standards runner-policy/component sync as needed → consumer PRs (dotfiles first: Shape A migration + consolidation together; standards fixtures lane; medley last — it already has detect-changes and the most lanes). Selector contract and approvedSelectorReferences governance unchanged.
Workstream of melodic-software/github-iac#78. Decision record from the 2026-07-16 capacity incident (228 queued jobs, 2h+ backlog, oldest job queued 11:59Z unreleased at 14:00Z).
Problem
The per-run fan-out shape optimizes for the wrong constraint. It was built to dodge GitHub-hosted per-job 1-minute billing rounding; on the self-hosted fleet the scarce resource is worker slots, and each micro-job pays ~20-30s of container spin-up + JIT registration + checkout for 7-60s of work. Measured today: ~40-50% of fleet throughput consumed by
select-runnerpreflights alone; dotfiles spends 16 selector jobs + 16 lanes per run (37 jobs, legacy Shape B); standards ~30 jobs/run; main-push bursts (4 pushes in 15s from distribution sync) enqueue 4 full runs with zero collapse.Target shape (research-verified 2026-07-16, official sources)
Per consumer repo:
hygienelane running all lint/hygiene checks as steps in one worker, usingcontinue-on-error: true+id:per check step and a final aggregation step failing on anysteps.<id>.outcome == 'failure'(run-everything-fail-at-end;outcome/conclusionsemantics per contexts reference). Language/build lanes only where toolchain or resource isolation genuinely differs (e.g. dotnet build/test). standards' fixture matrix folds into a fixtures lane.background: true/wait/parallel:— runner v2.335.1 (current fleet image) has the engine; max 10 concurrent background steps; a failing background step surfaces at the join; NOT usable inside composite actions. On 2-CPU workers this buys log clarity and overlap for network-bound checks (lychee), not CPU throughput — default to sequential steps.ci-statusfan-in unchanged as the single required check (job-level check granularity; a job skipped byif:reports Success for required-check purposes — keep in-workflow conditionals, never workflow-levelpaths:filters on required workflows).cancel-in-progress: true) as-is; give main pushesgroup: <workflow>-${{ github.ref }}withcancel-in-progress: false— the documented default (queue: single) then cancels the superseded pending run while the in-progress run finishes (concurrency docs). Do NOT usequeue: maxhere (it serializes every push — deploy semantics, not lint semantics).$GITHUB_STEP_SUMMARY(1MiB/step, 20 summaries/job). Annotation counts (10/step, 50/job) are community-reported only — do not design around them.Sizing interplay
Consolidated lanes get a larger default worker (2 CPU / 4GiB — CPU parity with hosted private-repo runners at 2 vCPU/8GB, RAM bounded by the per-host VM budget; provisioning#133/#134). Fewer, beefier workers replace many 2GiB micro-workers; setup-* toolchain installs amortize once per lane instead of once per micro-job, with
actions/cachefor dependencies (image already ships zstd for hosted cache parity).Sequencing
Reusable workflow/template changes here → standards runner-policy/component sync as needed → consumer PRs (dotfiles first: Shape A migration + consolidation together; standards fixtures lane; medley last — it already has detect-changes and the most lanes). Selector contract and
approvedSelectorReferencesgovernance unchanged.