Skip to content

[aw-failures] Drop "cancelled" from FAILURE_CONCLUSIONS — Dependabot PR bursts are creating 60+ false-positive failures per cycl [Content truncated due to length] #49582

Description

@github-actions

Problem

Drop cancelled from the failure-investigator's own FAILURE_CONCLUSIONS set — a single Dependabot bulk-PR batch just produced 60+ false-positive "failures" in one 6h window because every GitHub-Actions-cancelled run is counted as a failure. That's 3 real failures (failure/timed_out) buried under ~63 phantom ones.

Why: Dependabot opened/synced many PRs almost simultaneously (2026-08-01T07:56:12Z–07:58:00Z: dependabot/npm_and_yarn/docs/pdfjs-dist-6.2.108, dependabot/docker/alpine-3.24, dependabot/npm_and_yarn/docs/astrojs/markdown-remark-7.2.2, dependabot/npm_and_yarn/docs/astro-7.1.5, and more). Every PR-triggered workflow (Smoke Gemini, Smoke Antigravity, Smoke Agent variants, Code Refiner, Changeset Generator, etc.) got cancelled or skipped within 1–2s, before executing any agent logic — and the prefetch counted every single one as a failed run.

Affected workflows and run IDs

  • Smoke Gemini — §30690759438 (cancelled, 1.0s)
  • Smoke Antigravity — §30690759518 (cancelled)
  • Smoke Agent: scoped/approved — §30690733233 (cancelled, 2.0s)
  • Code Refiner — §30690747762 (cancelled)
  • Changeset Generator — §30690733575, §30690754996 (cancelled)
  • ~60 of the 68 run IDs in today's window share this exact signature (same 07:56–07:58Z burst); only 3 are genuine failure/timed_out conclusions.

Probable root cause

Fix .github/workflows/aw-failure-investigator.md:88:

FAILURE_CONCLUSIONS = {"failure", "timed_out", "startup_failure", "cancelled"}

This has no distinction between a genuinely-interrupted run and one superseded by GitHub's own concurrency-group cancellation before any work started.

Evidence: audit-diff between §30690759438 and §30690733233 shows zero firewall/domain drift and 0 token usage on both sides. audit on §30690759438 confirms duration: 1.0s, turns: 0, tool_types: 0, error_count: 0. Neither run executed any agent logic — this is scheduling noise, not a functional regression.

Proposed remediation

Exclude cancelled from FAILURE_CONCLUSIONS by default, or add a pre-execution filter that drops cancelled runs under ~10s duration (the signature of a pre-start concurrency-supersede) before they enter failed_run_ids.

Success criteria

The next Dependabot bulk-PR event does not inflate failed_run_ids with cancelled matrix runs — the investigator's 6h count tracks only genuine failure/timed_out/startup_failure conclusions.


Parent: #49245
Analyzed runs: 30690759438, 30690759518, 30690733233, 30690747762, 30690733575, 30690754996 (plus the full 07:56–07:58Z burst subset of today's 68 failed_run_ids)
Related to #49245

Generated by 🔍 [aw] Failure Investigator (6h) · agent · 191.1 AIC · ⌖ 25.1 AIC · ⊞ 6.8K · ◷

  • expires on Aug 8, 2026, 5:27 AM UTC-08:00

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions