Guard against repeated identical Sentry 400s — cap the retry, don't burn the full invocation budget on one dead-end query.
Workflow Portfolio Analyst has failed every scheduled run since 2026-06-22 (7+ consecutive failures, last success 2026-06-15 run 27574552812) — this is the single most persistent untracked failure in the repo right now.
What's happening
In run §30833881004, the agent repeatedly calls the Sentry MCP list_events tool using sum() aggregation on fields that are string-typed, not numeric:
gh-aw.aic
gh-aw.action_minutes
Sentry returns HTTP 400 for each call, since sum() is invalid on a string field. The agent doesn't treat this as a terminal/unrecoverable error — it retries the same query shape with small variations, over and over, until it hits the harness's 20/20 LLM invocation cap and the run is killed with zero useful output produced.
Why it's getting worse, not just persisting
audit-diff between this run and the prior failure (§30285971411) shows:
- GitHub API core-rate-limit consumption: +75% run-over-run
- Total API call count: +18% run-over-run
That's consistent with the agent doing more wasted retry work each cycle rather than converging on a workaround — the failure mode is stable but the blast radius is increasing.
Suggested fix (either or both)
- Fix the query itself in the Workflow Portfolio Analyst prompt/tooling: don't
sum() string fields (gh-aw.aic, gh-aw.action_minutes) — use count() or whatever the correct numeric field/aggregation is, or cast/parse before aggregating.
- Harden the retry loop: if the same Sentry query returns the same 400 error class more than 1-2 times in a run, fail fast with a clear error instead of continuing to spend LLM invocations on it. This protects every workflow that talks to Sentry MCP, not just this one.
References
Generated by 🔍 [aw] Failure Investigator (6h) · agent · 224.7 AIC · ⌖ 41.3 AIC · ⊞ 6.8K · ◷
Guard against repeated identical Sentry 400s — cap the retry, don't burn the full invocation budget on one dead-end query.
Workflow Portfolio Analyst has failed every scheduled run since 2026-06-22 (7+ consecutive failures, last success 2026-06-15 run 27574552812) — this is the single most persistent untracked failure in the repo right now.
What's happening
In run §30833881004, the agent repeatedly calls the Sentry MCP
list_eventstool usingsum()aggregation on fields that are string-typed, not numeric:gh-aw.aicgh-aw.action_minutesSentry returns HTTP 400 for each call, since
sum()is invalid on a string field. The agent doesn't treat this as a terminal/unrecoverable error — it retries the same query shape with small variations, over and over, until it hits the harness's 20/20 LLM invocation cap and the run is killed with zero useful output produced.Why it's getting worse, not just persisting
audit-diffbetween this run and the prior failure (§30285971411) shows:That's consistent with the agent doing more wasted retry work each cycle rather than converging on a workaround — the failure mode is stable but the blast radius is increasing.
Suggested fix (either or both)
sum()string fields (gh-aw.aic,gh-aw.action_minutes) — usecount()or whatever the correct numeric field/aggregation is, or cast/parse before aggregating.References
gh api "/repos/github/gh-aw/actions/workflows/portfolio-analyst.lock.yml/runs?per_page=10"for the full consecutive-failure historyRelated to [aw-failures] [aw] Failure Investigator Report — 2026-08-03 (6h) #50077