Skip to content

[aw-failures] [aw] Failure Investigator Report — 2026-08-05 (6h) #50518

Description

@github-actions

Patch the Kiro network allow-list first — it's the confirmed P0 root cause

audit-diff on 30974811087 vs 30977137021 proves shared/kiro.md's allow-list is missing two live domains that kiro-cli calls on every run — this is a one-line config fix, do it before anything else below.

None of the 4 untracked P0/P1 clusters below have existing coverage among the 20 open agentic-workflows issues, so this is a new parent report. One sub-issue (Cluster B) is filed as a tracked child; Clusters C/D/E are documented inline only — create_issue is capped at 2 calls per run, so they're not yet separate issues (flagging so they aren't silently dropped).

Failure clusters (last 6h)

Cluster Severity Signature Runs Status
A — Kiro/Cursor secret validation + E009 lock mismatch 4 runs, 02:33–04:13Z Stale — fixed by 80b45ff, no open issue references it
B P0 Kiro/Cursor CLI auth fails intermittently 4 runs New — sub-issue filed
C P1 check_token_telemetry false-fails on successful Kiro/Cursor agent runs 2 runs New — documented inline
D P1 Daily Container Image Scan: Critical CVEs across 9 third-party MCP images 1 run New — documented inline
E P2 Smoke Crush: crush run rejects its own prompt arg ("No prompt provided") 1 run New — documented inline

Cluster B (P0) — Kiro/Cursor auth breaks because the firewall allow-list and secret handling don't match reality

Fix the allow-list in shared/kiro.md today — one sentence why: audit-diff shows kiro-cli making live calls to q.us-east-1.amazonaws.com (42 requests, 100% denied) and client-telemetry.us-east-1.amazonaws.com (36 requests, 100% denied) in the run where the agent step still passed, neither of which is in network.defaults.

Evidence

network.defaults in shared/kiro.md currently lists only:

codewhisperer.us-east-1.amazonaws.com
cognito-identity.us-east-1.amazonaws.com
prod.us-east-1.telemetry.kiro.aws.dev
prod.assets.shortbread.aws.dev
*.kiro.dev (wildcard)

audit-diff (base=30974811087, compare=30977137021) output:

{"domain":"q.us-east-1.amazonaws.com:443","status":"new","is_anomaly":true,"anomaly_note":"new denied domain","run2_blocked":42,"run2_status":"denied"}
{"domain":"client-telemetry.us-east-1.amazonaws.com:443","status":"volume_changed","run1_blocked":2,"run2_blocked":36,"volume_change":"+1700%"}

Separately, both engines' harness-scripts (shared/kiro.md, shared/cursor.md) read process.env.KIRO_API_KEY / process.env.CURSOR_API_KEY directly, but the AWF entrypoint step "Unsetting sensitive tokens from parent shell environment" renames these to SECRET_KIRO_API_KEY / SECRET_CURSOR_API_KEY before the harness runs, with no fallback read of the SECRET_-prefixed name in either script. Confirmed in job logs for 30975982770 (Kiro: error: Failed to open URL, falls back to kiro-cli login --use-device-flow) and 30974539538 (Cursor: Error: Authentication required... or set CURSOR_API_KEY environment variable.).

Both mechanisms are intermittent-failure-shaped (some runs of the same workflow succeed, e.g. 30977137021, 30975573886), consistent with a race/retry-dependent network call plus an auth fallback path that only triggers when that call fails.

Root cause: (1) shared/kiro.md network allow-list is stale relative to kiro-cli's actual runtime call graph; (2) neither Kiro nor Cursor harness-script falls back to the SECRET_-prefixed env var name the AWF entrypoint renames auth secrets to.

Remediation:

  1. Add q.us-east-1.amazonaws.com and client-telemetry.us-east-1.amazonaws.com to network.defaults in shared/kiro.md.
  2. In both shared/kiro.md and shared/cursor.md harness-scripts, read process.env.KIRO_API_KEY || process.env.SECRET_KIRO_API_KEY (and the Cursor equivalent) before invoking the CLI.

Success criteria: 10 consecutive Smoke Kiro + Smoke Cursor runs with zero auth failures and zero denied entries for the two domains above in audit-diff/firewall logs.


Cluster C (P1) — check_token_telemetry is structurally wrong for provider-native engines

Stop importing shared/token-telemetry-check.md into Kiro/Cursor smoke workflows — one sentence why: the check asserts AWF's api-proxy token-usage.jsonl is non-empty, but Kiro and Cursor never route through the api-proxy sidecar, so the file is empty by design, not by failure.

Evidence

Runs 30977137021 (Kiro) and 30975573886 (Cursor): agent job = success, check_token_telemetry job = failure in both. Log line confirms design intent: [INFO] API proxy enabled: OpenAI=false, Anthropic=false, Copilot=false, Gemini=false, Vertex=false / [WARN] API proxy enabled but no API keys found in environment. The check has no branch for provider-native engines.

Root cause: shared/token-telemetry-check.md hardcodes an AWF-proxy-routing assumption that doesn't hold for provider-native engines (Kiro, Cursor; likely also any future non-proxied engine).

Remediation: Either drop the import in smoke-kiro.md/smoke-cursor.md, or make the check provider-aware (skip the token-usage.jsonl assertion when engine.provider.name is a non-proxied provider, keep the agent_usage.json non-zero-token check).

Success criteria: Smoke Kiro/Cursor runs stop failing check_token_telemetry while agent succeeds; provider-proxied engines (Claude/Codex/Copilot/Gemini) keep the existing check unchanged.


Cluster D (P1) — Daily Container Image Scan is failing on real Critical CVEs across 9 pinned MCP server images

Bump or re-pin the flagged images now — one sentence why: the scan gate (: error: [Critical]) is working correctly; it found genuine Critical-severity CVEs in github-mcp-server, serena-mcp-server, grafana/mcp-grafana, node:lts-alpine, mcp/ast-grep, mcp/arxiv-mcp-server, mcp/context7, mcp/memory, and ghcr.io/fabio-rovai/open-ontologies.

Evidence (representative lines, run (a href="https://github.com/github/gh-aw/actions/runs/30980674254")30980674254(/a))
serena-mcp-server:sha-891c160: [Critical] CVE-2025-68121 golang-1.24-go@1.24.4-1 (no fix yet, GO-2026-4337 stdlib advisory)
serena-mcp-server:sha-891c160: [Critical] CVE-2026-42010/CVE-2026-33845 libgnutls30t64 (fix: 3.8.9-3+deb13u4)
serena-mcp-server:sha-891c160: [Critical] GHSA-2w6w-674q-4c4q handlebars@4.7.7 (fix: 4.7.9)
node:lts-alpine: [Critical] GHSA-23hp-3jrh-7fpw tar@7.5.16 (fix: 7.5.19)
mcp/context7, mcp/memory: [Critical] GHSA-23hp-3jrh-7fpw tar@6.2.1/7.4.3 (fix: 7.5.19), CVE-2025-55130 node (fix: 20.20.0+/22.22.0+/24.13.0+/25.3.0+)
mcp/ast-grep, mcp/arxiv-mcp-server: [Critical] CVE-2025-15467/CVE-2026-34182/CVE-2026-31789 libcrypto3/libssl3
grafana/mcp-grafana, mcp/arxiv-mcp-server: [Critical] CVE-2026-8376/CVE-2026-13221/CVE-2026-42496/CVE-2026-12087/CVE-2026-57433 perl-base

50+ distinct Critical findings total across these 9 images — full list in job logs.

Root cause: these are third-party/upstream MCP server images pinned to versions (or floating tags like latest/lts-alpine) that have accumulated new Critical CVEs since last scan; not a gh-aw code defect.

Remediation: bump each pinned image to its latest patched tag where a fix exists (tar, libssl3/libcrypto3, perl-base, libgnutls30t64, handlebars all have fixed versions available); for golang-1.24-go CVE-2025-68121 (no fix yet) and any other no-fix CVEs, add a time-boxed scan exception with a tracking date rather than leaving the gate red indefinitely.

Success criteria: Daily Container Image Security Scan passes with zero [Critical] findings, or documented time-boxed exceptions for no-fix-available CVEs only.


Cluster E (P2) — Smoke Crush: harness passes a prompt the CLI refuses to see

Deprioritize but track — one sentence why: single occurrence, experimental engine, but the harness-script's own arg construction looks correct, so this is worth a focused repro before assuming it's transient.

Evidence

Run 30967390017: shared/crush.md harness-script runs spawnSync(command, [...commandArgs, "--model", "awf-proxy/<model>", prompt], ...) i.e. crush run --quiet --model awf-proxy/<model> "<prompt text>". CLI immediately errors No prompt provided and exits 1, ~100ms after the AWF entrypoint finished env setup — before any model/network activity. Not an auth issue (Crush uses universal-llm-consumer proxy routing, unlike Kiro/Cursor).

Root cause (unconfirmed, needs repro): @charmland/crush@0.88.0's run --quiet argument parser likely doesn't accept the prompt as a trailing positional after --model, or requires it via stdin in --quiet mode.

Remediation: repro locally with crush run --quiet --model <x> "<prompt>" vs crush run --quiet "<prompt>" --model <x> vs piping the prompt via stdin; adjust shared/crush.md's harness-script arg order to match whichever the installed CLI version actually accepts.

Success criteria: Smoke Crush's Execute Crush CLI step passes with a non-empty prompt reaching the agent.


References

Generated by 🔍 [aw] Failure Investigator (6h) · agent · 259.3 AIC · ⌖ 40.8 AIC · ⊞ 5.2K · ◷

  • expires on Aug 11, 2026, 11:58 PM UTC-08:00

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions