Patch the Kiro network allow-list first — it's the confirmed P0 root cause
audit-diff on 30974811087 vs 30977137021 proves shared/kiro.md's allow-list is missing two live domains that kiro-cli calls on every run — this is a one-line config fix, do it before anything else below.
None of the 4 untracked P0/P1 clusters below have existing coverage among the 20 open agentic-workflows issues, so this is a new parent report. One sub-issue (Cluster B) is filed as a tracked child; Clusters C/D/E are documented inline only — create_issue is capped at 2 calls per run, so they're not yet separate issues (flagging so they aren't silently dropped).
Failure clusters (last 6h)
| Cluster |
Severity |
Signature |
Runs |
Status |
| A |
— |
Kiro/Cursor secret validation + E009 lock mismatch |
4 runs, 02:33–04:13Z |
Stale — fixed by 80b45ff, no open issue references it |
| B |
P0 |
Kiro/Cursor CLI auth fails intermittently |
4 runs |
New — sub-issue filed |
| C |
P1 |
check_token_telemetry false-fails on successful Kiro/Cursor agent runs |
2 runs |
New — documented inline |
| D |
P1 |
Daily Container Image Scan: Critical CVEs across 9 third-party MCP images |
1 run |
New — documented inline |
| E |
P2 |
Smoke Crush: crush run rejects its own prompt arg ("No prompt provided") |
1 run |
New — documented inline |
Cluster B (P0) — Kiro/Cursor auth breaks because the firewall allow-list and secret handling don't match reality
Fix the allow-list in shared/kiro.md today — one sentence why: audit-diff shows kiro-cli making live calls to q.us-east-1.amazonaws.com (42 requests, 100% denied) and client-telemetry.us-east-1.amazonaws.com (36 requests, 100% denied) in the run where the agent step still passed, neither of which is in network.defaults.
Evidence
network.defaults in shared/kiro.md currently lists only:
codewhisperer.us-east-1.amazonaws.com
cognito-identity.us-east-1.amazonaws.com
prod.us-east-1.telemetry.kiro.aws.dev
prod.assets.shortbread.aws.dev
*.kiro.dev (wildcard)
audit-diff (base=30974811087, compare=30977137021) output:
{"domain":"q.us-east-1.amazonaws.com:443","status":"new","is_anomaly":true,"anomaly_note":"new denied domain","run2_blocked":42,"run2_status":"denied"}
{"domain":"client-telemetry.us-east-1.amazonaws.com:443","status":"volume_changed","run1_blocked":2,"run2_blocked":36,"volume_change":"+1700%"}
Separately, both engines' harness-scripts (shared/kiro.md, shared/cursor.md) read process.env.KIRO_API_KEY / process.env.CURSOR_API_KEY directly, but the AWF entrypoint step "Unsetting sensitive tokens from parent shell environment" renames these to SECRET_KIRO_API_KEY / SECRET_CURSOR_API_KEY before the harness runs, with no fallback read of the SECRET_-prefixed name in either script. Confirmed in job logs for 30975982770 (Kiro: error: Failed to open URL, falls back to kiro-cli login --use-device-flow) and 30974539538 (Cursor: Error: Authentication required... or set CURSOR_API_KEY environment variable.).
Both mechanisms are intermittent-failure-shaped (some runs of the same workflow succeed, e.g. 30977137021, 30975573886), consistent with a race/retry-dependent network call plus an auth fallback path that only triggers when that call fails.
Root cause: (1) shared/kiro.md network allow-list is stale relative to kiro-cli's actual runtime call graph; (2) neither Kiro nor Cursor harness-script falls back to the SECRET_-prefixed env var name the AWF entrypoint renames auth secrets to.
Remediation:
- Add
q.us-east-1.amazonaws.com and client-telemetry.us-east-1.amazonaws.com to network.defaults in shared/kiro.md.
- In both
shared/kiro.md and shared/cursor.md harness-scripts, read process.env.KIRO_API_KEY || process.env.SECRET_KIRO_API_KEY (and the Cursor equivalent) before invoking the CLI.
Success criteria: 10 consecutive Smoke Kiro + Smoke Cursor runs with zero auth failures and zero denied entries for the two domains above in audit-diff/firewall logs.
Cluster C (P1) — check_token_telemetry is structurally wrong for provider-native engines
Stop importing shared/token-telemetry-check.md into Kiro/Cursor smoke workflows — one sentence why: the check asserts AWF's api-proxy token-usage.jsonl is non-empty, but Kiro and Cursor never route through the api-proxy sidecar, so the file is empty by design, not by failure.
Evidence
Runs 30977137021 (Kiro) and 30975573886 (Cursor): agent job = success, check_token_telemetry job = failure in both. Log line confirms design intent: [INFO] API proxy enabled: OpenAI=false, Anthropic=false, Copilot=false, Gemini=false, Vertex=false / [WARN] API proxy enabled but no API keys found in environment. The check has no branch for provider-native engines.
Root cause: shared/token-telemetry-check.md hardcodes an AWF-proxy-routing assumption that doesn't hold for provider-native engines (Kiro, Cursor; likely also any future non-proxied engine).
Remediation: Either drop the import in smoke-kiro.md/smoke-cursor.md, or make the check provider-aware (skip the token-usage.jsonl assertion when engine.provider.name is a non-proxied provider, keep the agent_usage.json non-zero-token check).
Success criteria: Smoke Kiro/Cursor runs stop failing check_token_telemetry while agent succeeds; provider-proxied engines (Claude/Codex/Copilot/Gemini) keep the existing check unchanged.
Cluster D (P1) — Daily Container Image Scan is failing on real Critical CVEs across 9 pinned MCP server images
Bump or re-pin the flagged images now — one sentence why: the scan gate (: error: [Critical]) is working correctly; it found genuine Critical-severity CVEs in github-mcp-server, serena-mcp-server, grafana/mcp-grafana, node:lts-alpine, mcp/ast-grep, mcp/arxiv-mcp-server, mcp/context7, mcp/memory, and ghcr.io/fabio-rovai/open-ontologies.
Evidence (representative lines, run (a href="https://github.com/github/gh-aw/actions/runs/30980674254")30980674254(/a))
serena-mcp-server:sha-891c160: [Critical] CVE-2025-68121 golang-1.24-go@1.24.4-1 (no fix yet, GO-2026-4337 stdlib advisory)
serena-mcp-server:sha-891c160: [Critical] CVE-2026-42010/CVE-2026-33845 libgnutls30t64 (fix: 3.8.9-3+deb13u4)
serena-mcp-server:sha-891c160: [Critical] GHSA-2w6w-674q-4c4q handlebars@4.7.7 (fix: 4.7.9)
node:lts-alpine: [Critical] GHSA-23hp-3jrh-7fpw tar@7.5.16 (fix: 7.5.19)
mcp/context7, mcp/memory: [Critical] GHSA-23hp-3jrh-7fpw tar@6.2.1/7.4.3 (fix: 7.5.19), CVE-2025-55130 node (fix: 20.20.0+/22.22.0+/24.13.0+/25.3.0+)
mcp/ast-grep, mcp/arxiv-mcp-server: [Critical] CVE-2025-15467/CVE-2026-34182/CVE-2026-31789 libcrypto3/libssl3
grafana/mcp-grafana, mcp/arxiv-mcp-server: [Critical] CVE-2026-8376/CVE-2026-13221/CVE-2026-42496/CVE-2026-12087/CVE-2026-57433 perl-base
50+ distinct Critical findings total across these 9 images — full list in job logs.
Root cause: these are third-party/upstream MCP server images pinned to versions (or floating tags like latest/lts-alpine) that have accumulated new Critical CVEs since last scan; not a gh-aw code defect.
Remediation: bump each pinned image to its latest patched tag where a fix exists (tar, libssl3/libcrypto3, perl-base, libgnutls30t64, handlebars all have fixed versions available); for golang-1.24-go CVE-2025-68121 (no fix yet) and any other no-fix CVEs, add a time-boxed scan exception with a tracking date rather than leaving the gate red indefinitely.
Success criteria: Daily Container Image Security Scan passes with zero [Critical] findings, or documented time-boxed exceptions for no-fix-available CVEs only.
Cluster E (P2) — Smoke Crush: harness passes a prompt the CLI refuses to see
Deprioritize but track — one sentence why: single occurrence, experimental engine, but the harness-script's own arg construction looks correct, so this is worth a focused repro before assuming it's transient.
Evidence
Run 30967390017: shared/crush.md harness-script runs spawnSync(command, [...commandArgs, "--model", "awf-proxy/<model>", prompt], ...) i.e. crush run --quiet --model awf-proxy/<model> "<prompt text>". CLI immediately errors No prompt provided and exits 1, ~100ms after the AWF entrypoint finished env setup — before any model/network activity. Not an auth issue (Crush uses universal-llm-consumer proxy routing, unlike Kiro/Cursor).
Root cause (unconfirmed, needs repro): @charmland/crush@0.88.0's run --quiet argument parser likely doesn't accept the prompt as a trailing positional after --model, or requires it via stdin in --quiet mode.
Remediation: repro locally with crush run --quiet --model <x> "<prompt>" vs crush run --quiet "<prompt>" --model <x> vs piping the prompt via stdin; adjust shared/crush.md's harness-script arg order to match whichever the installed CLI version actually accepts.
Success criteria: Smoke Crush's Execute Crush CLI step passes with a non-empty prompt reaching the agent.
References
Generated by 🔍 [aw] Failure Investigator (6h) · agent · 259.3 AIC · ⌖ 40.8 AIC · ⊞ 5.2K · ◷
Patch the Kiro network allow-list first — it's the confirmed P0 root cause
audit-diffon 30974811087 vs 30977137021 provesshared/kiro.md's allow-list is missing two live domains that kiro-cli calls on every run — this is a one-line config fix, do it before anything else below.None of the 4 untracked P0/P1 clusters below have existing coverage among the 20 open
agentic-workflowsissues, so this is a new parent report. One sub-issue (Cluster B) is filed as a tracked child; Clusters C/D/E are documented inline only —create_issueis capped at 2 calls per run, so they're not yet separate issues (flagging so they aren't silently dropped).Failure clusters (last 6h)
check_token_telemetryfalse-fails on successful Kiro/Cursor agent runscrush runrejects its own prompt arg ("No prompt provided")Cluster B (P0) — Kiro/Cursor auth breaks because the firewall allow-list and secret handling don't match reality
Fix the allow-list in
shared/kiro.mdtoday — one sentence why:audit-diffshows kiro-cli making live calls toq.us-east-1.amazonaws.com(42 requests, 100% denied) andclient-telemetry.us-east-1.amazonaws.com(36 requests, 100% denied) in the run where the agent step still passed, neither of which is innetwork.defaults.Evidence
network.defaultsinshared/kiro.mdcurrently lists only:audit-diff(base=30974811087, compare=30977137021) output:{"domain":"q.us-east-1.amazonaws.com:443","status":"new","is_anomaly":true,"anomaly_note":"new denied domain","run2_blocked":42,"run2_status":"denied"} {"domain":"client-telemetry.us-east-1.amazonaws.com:443","status":"volume_changed","run1_blocked":2,"run2_blocked":36,"volume_change":"+1700%"}Separately, both engines' harness-scripts (
shared/kiro.md,shared/cursor.md) readprocess.env.KIRO_API_KEY/process.env.CURSOR_API_KEYdirectly, but the AWF entrypoint step "Unsetting sensitive tokens from parent shell environment" renames these toSECRET_KIRO_API_KEY/SECRET_CURSOR_API_KEYbefore the harness runs, with no fallback read of theSECRET_-prefixed name in either script. Confirmed in job logs for 30975982770 (Kiro:error: Failed to open URL, falls back tokiro-cli login --use-device-flow) and 30974539538 (Cursor:Error: Authentication required... or set CURSOR_API_KEY environment variable.).Both mechanisms are intermittent-failure-shaped (some runs of the same workflow succeed, e.g. 30977137021, 30975573886), consistent with a race/retry-dependent network call plus an auth fallback path that only triggers when that call fails.
Root cause: (1)
shared/kiro.mdnetwork allow-list is stale relative to kiro-cli's actual runtime call graph; (2) neither Kiro nor Cursor harness-script falls back to theSECRET_-prefixed env var name the AWF entrypoint renames auth secrets to.Remediation:
q.us-east-1.amazonaws.comandclient-telemetry.us-east-1.amazonaws.comtonetwork.defaultsinshared/kiro.md.shared/kiro.mdandshared/cursor.mdharness-scripts, readprocess.env.KIRO_API_KEY || process.env.SECRET_KIRO_API_KEY(and the Cursor equivalent) before invoking the CLI.Success criteria: 10 consecutive Smoke Kiro + Smoke Cursor runs with zero auth failures and zero
deniedentries for the two domains above inaudit-diff/firewall logs.Cluster C (P1) —
check_token_telemetryis structurally wrong for provider-native enginesStop importing
shared/token-telemetry-check.mdinto Kiro/Cursor smoke workflows — one sentence why: the check asserts AWF's api-proxytoken-usage.jsonlis non-empty, but Kiro and Cursor never route through the api-proxy sidecar, so the file is empty by design, not by failure.Evidence
Runs 30977137021 (Kiro) and 30975573886 (Cursor):
agentjob = success,check_token_telemetryjob = failure in both. Log line confirms design intent:[INFO] API proxy enabled: OpenAI=false, Anthropic=false, Copilot=false, Gemini=false, Vertex=false/[WARN] API proxy enabled but no API keys found in environment. The check has no branch for provider-native engines.Root cause:
shared/token-telemetry-check.mdhardcodes an AWF-proxy-routing assumption that doesn't hold for provider-native engines (Kiro, Cursor; likely also any future non-proxied engine).Remediation: Either drop the import in
smoke-kiro.md/smoke-cursor.md, or make the check provider-aware (skip thetoken-usage.jsonlassertion whenengine.provider.nameis a non-proxied provider, keep theagent_usage.jsonnon-zero-token check).Success criteria: Smoke Kiro/Cursor runs stop failing
check_token_telemetrywhileagentsucceeds; provider-proxied engines (Claude/Codex/Copilot/Gemini) keep the existing check unchanged.Cluster D (P1) — Daily Container Image Scan is failing on real Critical CVEs across 9 pinned MCP server images
Bump or re-pin the flagged images now — one sentence why: the scan gate (
: error: [Critical]) is working correctly; it found genuine Critical-severity CVEs ingithub-mcp-server,serena-mcp-server,grafana/mcp-grafana,node:lts-alpine,mcp/ast-grep,mcp/arxiv-mcp-server,mcp/context7,mcp/memory, andghcr.io/fabio-rovai/open-ontologies.Evidence (representative lines, run (a href="https://github.com/github/gh-aw/actions/runs/30980674254")30980674254(/a))
50+ distinct Critical findings total across these 9 images — full list in job logs.
Root cause: these are third-party/upstream MCP server images pinned to versions (or floating tags like
latest/lts-alpine) that have accumulated new Critical CVEs since last scan; not a gh-aw code defect.Remediation: bump each pinned image to its latest patched tag where a fix exists (
tar,libssl3/libcrypto3,perl-base,libgnutls30t64,handlebarsall have fixed versions available); forgolang-1.24-goCVE-2025-68121 (no fix yet) and any other no-fix CVEs, add a time-boxed scan exception with a tracking date rather than leaving the gate red indefinitely.Success criteria: Daily Container Image Security Scan passes with zero
[Critical]findings, or documented time-boxed exceptions for no-fix-available CVEs only.Cluster E (P2) — Smoke Crush: harness passes a prompt the CLI refuses to see
Deprioritize but track — one sentence why: single occurrence, experimental engine, but the harness-script's own arg construction looks correct, so this is worth a focused repro before assuming it's transient.
Evidence
Run 30967390017:
shared/crush.mdharness-script runsspawnSync(command, [...commandArgs, "--model", "awf-proxy/<model>", prompt], ...)i.e.crush run --quiet --model awf-proxy/<model> "<prompt text>". CLI immediately errorsNo prompt providedand exits 1, ~100ms after the AWF entrypoint finished env setup — before any model/network activity. Not an auth issue (Crush usesuniversal-llm-consumerproxy routing, unlike Kiro/Cursor).Root cause (unconfirmed, needs repro):
@charmland/crush@0.88.0'srun --quietargument parser likely doesn't accept the prompt as a trailing positional after--model, or requires it via stdin in--quietmode.Remediation: repro locally with
crush run --quiet --model <x> "<prompt>"vscrush run --quiet "<prompt>" --model <x>vs piping the prompt via stdin; adjustshared/crush.md's harness-script arg order to match whichever the installed CLI version actually accepts.Success criteria: Smoke Crush's
Execute Crush CLIstep passes with a non-empty prompt reaching the agent.References