Problem
Drop cancelled from the failure-investigator's own FAILURE_CONCLUSIONS set — a single Dependabot bulk-PR batch just produced 60+ false-positive "failures" in one 6h window because every GitHub-Actions-cancelled run is counted as a failure. That's 3 real failures (failure/timed_out) buried under ~63 phantom ones.
Why: Dependabot opened/synced many PRs almost simultaneously (2026-08-01T07:56:12Z–07:58:00Z: dependabot/npm_and_yarn/docs/pdfjs-dist-6.2.108, dependabot/docker/alpine-3.24, dependabot/npm_and_yarn/docs/astrojs/markdown-remark-7.2.2, dependabot/npm_and_yarn/docs/astro-7.1.5, and more). Every PR-triggered workflow (Smoke Gemini, Smoke Antigravity, Smoke Agent variants, Code Refiner, Changeset Generator, etc.) got cancelled or skipped within 1–2s, before executing any agent logic — and the prefetch counted every single one as a failed run.
Affected workflows and run IDs
- Smoke Gemini — §30690759438 (
cancelled, 1.0s)
- Smoke Antigravity — §30690759518 (
cancelled)
- Smoke Agent: scoped/approved — §30690733233 (
cancelled, 2.0s)
- Code Refiner — §30690747762 (
cancelled)
- Changeset Generator — §30690733575, §30690754996 (
cancelled)
- ~60 of the 68 run IDs in today's window share this exact signature (same 07:56–07:58Z burst); only 3 are genuine
failure/timed_out conclusions.
Probable root cause
Fix .github/workflows/aw-failure-investigator.md:88:
FAILURE_CONCLUSIONS = {"failure", "timed_out", "startup_failure", "cancelled"}
This has no distinction between a genuinely-interrupted run and one superseded by GitHub's own concurrency-group cancellation before any work started.
Evidence: audit-diff between §30690759438 and §30690733233 shows zero firewall/domain drift and 0 token usage on both sides. audit on §30690759438 confirms duration: 1.0s, turns: 0, tool_types: 0, error_count: 0. Neither run executed any agent logic — this is scheduling noise, not a functional regression.
Proposed remediation
Exclude cancelled from FAILURE_CONCLUSIONS by default, or add a pre-execution filter that drops cancelled runs under ~10s duration (the signature of a pre-start concurrency-supersede) before they enter failed_run_ids.
Success criteria
The next Dependabot bulk-PR event does not inflate failed_run_ids with cancelled matrix runs — the investigator's 6h count tracks only genuine failure/timed_out/startup_failure conclusions.
Parent: #49245
Analyzed runs: 30690759438, 30690759518, 30690733233, 30690747762, 30690733575, 30690754996 (plus the full 07:56–07:58Z burst subset of today's 68 failed_run_ids)
Related to #49245
Generated by 🔍 [aw] Failure Investigator (6h) · agent · 191.1 AIC · ⌖ 25.1 AIC · ⊞ 6.8K · ◷
Problem
Drop
cancelledfrom the failure-investigator's ownFAILURE_CONCLUSIONSset — a single Dependabot bulk-PR batch just produced 60+ false-positive "failures" in one 6h window because every GitHub-Actions-cancelled run is counted as a failure. That's 3 real failures (failure/timed_out) buried under ~63 phantom ones.Why: Dependabot opened/synced many PRs almost simultaneously (2026-08-01T07:56:12Z–07:58:00Z:
dependabot/npm_and_yarn/docs/pdfjs-dist-6.2.108,dependabot/docker/alpine-3.24,dependabot/npm_and_yarn/docs/astrojs/markdown-remark-7.2.2,dependabot/npm_and_yarn/docs/astro-7.1.5, and more). Every PR-triggered workflow (Smoke Gemini, Smoke Antigravity, Smoke Agent variants, Code Refiner, Changeset Generator, etc.) got cancelled or skipped within 1–2s, before executing any agent logic — and the prefetch counted every single one as a failed run.Affected workflows and run IDs
cancelled, 1.0s)cancelled)cancelled, 2.0s)cancelled)cancelled)failure/timed_outconclusions.Probable root cause
Fix
.github/workflows/aw-failure-investigator.md:88:This has no distinction between a genuinely-interrupted run and one superseded by GitHub's own concurrency-group cancellation before any work started.
Evidence:
audit-diffbetween §30690759438 and §30690733233 shows zero firewall/domain drift and 0 token usage on both sides.auditon §30690759438 confirmsduration: 1.0s,turns: 0,tool_types: 0,error_count: 0. Neither run executed any agent logic — this is scheduling noise, not a functional regression.Proposed remediation
Exclude
cancelledfromFAILURE_CONCLUSIONSby default, or add a pre-execution filter that dropscancelledruns under ~10s duration (the signature of a pre-start concurrency-supersede) before they enterfailed_run_ids.Success criteria
The next Dependabot bulk-PR event does not inflate
failed_run_idswith cancelled matrix runs — the investigator's 6h count tracks only genuinefailure/timed_out/startup_failureconclusions.Parent: #49245
Analyzed runs: 30690759438, 30690759518, 30690733233, 30690747762, 30690733575, 30690754996 (plus the full 07:56–07:58Z burst subset of today's 68 failed_run_ids)
Related to #49245