1. Executive Summary
Status: WARNING
Period: 2026-05-20 12:40Z – 2026-05-21 03:01Z
Result: 3 of 100 runs failed (3%) — 70 succeeded, 11 cancelled, 16 skipped
Key findings:
- All 3 failure logs are unavailable (
log unavailable), preventing definitive root-cause identification
- Failures cluster tightly in the 01:10–03:01 UTC window on 2026-05-21 and are non-consecutive (each followed immediately by successful retriggers), strongly suggesting transient infrastructure or rate-limit events
- 11 cancelled runs warrant investigation —
cancel-in-progress: false means the concurrency slot is not auto-cancelling; manual cancellations or upstream queue flushes are likely
Action required: Investigate why GitHub is not retaining logs for runs #1527, #1531, #1537 — enable log artifact upload or step-level debug output so failures are diagnosable.
2. Failure Breakdown
| Failure Category |
Affected Runs |
Example Error Message |
| Log unavailable / retention gap |
#1527, #1531, #1537 |
(log unavailable for run 26199375580 / 26199692481 / 26202833228) |
| Transient infra / rate limit (suspected) |
#1527, #1531, #1537 |
Unknown — inferred from recovery pattern (each failure bracketed by successes) |
| Cancellation (origin unclear) |
#1440, #1441, #1450, #1470, #1490, #1491, #1492, #1502, #1503, #1511, #1521 |
(cancelled — no log) |
3. Error Patterns
Pattern A — Log retention gap (confirmed)
(log unavailable for run 26202833228)
(log unavailable for run 26199692481)
(log unavailable for run 26199375580)
- Step: Any — log was never captured or was purged before retrieval
- Root cause: GitHub Actions log retention is 90 days by default, so age is not the cause here. More likely: the runner was evicted mid-job before streaming finished, OR the job failed during setup (before any step output flushed), OR the log fetch API request raced with job cleanup. A runner OOM or spot-instance preemption during the
Install review engine CLIs step (which runs npm install -g commands) is a plausible trigger.
- Recovery signal: In all three cases, the immediately following run(s) at the same concurrency slot succeeded — consistent with transient infrastructure events (spot preemption, ephemeral network blip, npm registry hiccup).
Pattern B — Cancellations with cancel-in-progress: false
- Step: Pre-queue / runner allocation
- Root cause: With
cancel-in-progress: false, GitHub will not auto-cancel queued runs. The 11 cancellations are therefore either (a) manual workflow cancellations triggered by an operator or upstream automation, or (b) the concurrency group resolved to pr-review-batch for multiple simultaneous non-PR-specific triggers and an operator flushed the queue. No evidence of a code defect here.
4. Token Scope Analysis
| Credential |
Configured Via |
Scopes Needed |
Status |
DON_PETRY_BOT_GH_PAT → GH_TOKEN |
secrets.DON_PETRY_BOT_GH_PAT |
repo, read:org, pull-requests: write (workflow-level) |
Cannot verify — logs unavailable; no gh auth status output present |
CLAUDE_CODE_OAUTH_TOKEN |
secrets.CLAUDE_CODE_OAUTH_TOKEN |
Claude API access |
Cannot verify — no auth output in logs |
COPILOT_GITHUB_TOKEN |
secrets.GH_PAT |
read:user, Copilot subscription |
Cannot verify |
GOOGLE_API_KEY / GEMINI_API_KEY |
secrets.GOOGLE_API_KEY (shared) |
Gemini API enabled |
Cannot verify |
Scopes confirmed present: None — no gh auth status output is available in any retrieved log.
Scopes potentially missing (inferred from workflow risk surface):
read:org — required by list-prs.sh if it enumerates org PRs; missing scope would cause silent empty results or 403 on org API calls
workflow — required if the bot PAT ever needs to trigger other workflows via API; not currently used but noted as a common gap
- Gemini CLI may require project-level enablement beyond API key —
GEMINI_CLI_TRUST_WORKSPACE=true is set but no quota/billing errors are visible
Recommendations:
- Add
gh auth status output as a dedicated early step (or echo "::debug::$(gh auth status 2>&1)") so scope issues appear in every run log
- Validate
DON_PETRY_BOT_GH_PAT includes repo + read:org + read:discussion if cross-org PR listing is used
5. Recommendations
-
[CRITICAL] Capture logs before they go missing — add a failure artifact upload step
- What: Append to
.github/workflows/pr-review.yml after the Review each PR (cascade) step:
- name: Upload debug log on failure
if: failure()
uses: actions/upload-artifact@v4
with:
name: pr-review-debug-${{ github.run_number }}
path: |
prs.txt
/tmp/review-*.log
retention-days: 14
- Why: All 3 failures produced no diagnosable output. Without artifact capture, future failures will also be undiagnosable.
- Expected impact: 100% of future failures become diagnosable.
- Urgency: CRITICAL
-
[HIGH] Add gh auth status as an explicit early step
- What: Insert before
Enumerate candidate PRs:
- name: Verify auth scopes
run: gh auth status
- Why: Token scope errors are silent and hard to distinguish from logic failures without this output.
- Expected impact: Auth regressions (e.g. secret rotation breaking scope) become immediately visible.
- Urgency: HIGH
-
[HIGH] Investigate the 11 cancelled runs
- What: Cross-reference the cancellation timestamps against operator activity logs and any upstream automation that calls
gh workflow cancel. Determine if cancellations are intentional queue flushes or a bug in a caller.
- Why: 11% cancellation rate is non-trivial and may represent lost review coverage if the cancellations interrupt in-progress reviews (even with
cancel-in-progress: false, queued runs are cancellable).
- Expected impact: Confirm whether PRs are receiving reviews after a cancel event or falling through the gap.
- Urgency: HIGH
-
[MEDIUM] Add spot-instance resilience to the npm install step
- What: Wrap the
npm install -g calls in review-batch.sh or the workflow step with a retry loop:
for i in 1 2 3; do npm install -g "@anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}" && break || sleep 15; done
- Why: If failures are caused by transient npm registry or runner preemption events, a simple retry prevents full run failure.
- Expected impact: Reduces transient failure rate.
- Urgency: MEDIUM
-
[LOW] Scope the concurrency group more tightly for check_suite batch fallthrough
- What: The current logic falls back to
pr-review-batch when no PR URL is present. Multiple simultaneous check_suite events on non-PR branches all share this slot and queue. Consider adding a timestamp or run-id suffix to the batch group to allow parallelism.
- Why: The 11 cancellations may partly stem from batch-slot contention.
- Expected impact: Reduces unnecessary queuing and manual cancellation needs.
- Urgency: LOW
6. Health Score
Health: 7/10 — Failure rate is low (3%) and self-recovering, but all failures are completely undiagnosable due to missing logs, making it impossible to confirm the root cause or detect a silent worsening trend.
1. Executive Summary
Status: WARNING
Period: 2026-05-20 12:40Z – 2026-05-21 03:01Z
Result: 3 of 100 runs failed (3%) — 70 succeeded, 11 cancelled, 16 skipped
Key findings:
log unavailable), preventing definitive root-cause identificationcancel-in-progress: falsemeans the concurrency slot is not auto-cancelling; manual cancellations or upstream queue flushes are likelyAction required: Investigate why GitHub is not retaining logs for runs #1527, #1531, #1537 — enable log artifact upload or step-level debug output so failures are diagnosable.
2. Failure Breakdown
3. Error Patterns
Pattern A — Log retention gap (confirmed)
Install review engine CLIsstep (which runsnpm install -gcommands) is a plausible trigger.Pattern B — Cancellations with
cancel-in-progress: falsecancel-in-progress: false, GitHub will not auto-cancel queued runs. The 11 cancellations are therefore either (a) manual workflow cancellations triggered by an operator or upstream automation, or (b) the concurrency group resolved topr-review-batchfor multiple simultaneous non-PR-specific triggers and an operator flushed the queue. No evidence of a code defect here.4. Token Scope Analysis
DON_PETRY_BOT_GH_PAT→GH_TOKENsecrets.DON_PETRY_BOT_GH_PATrepo,read:org,pull-requests: write(workflow-level)gh auth statusoutput presentCLAUDE_CODE_OAUTH_TOKENsecrets.CLAUDE_CODE_OAUTH_TOKENCOPILOT_GITHUB_TOKENsecrets.GH_PATread:user, Copilot subscriptionGOOGLE_API_KEY/GEMINI_API_KEYsecrets.GOOGLE_API_KEY(shared)Scopes confirmed present: None — no
gh auth statusoutput is available in any retrieved log.Scopes potentially missing (inferred from workflow risk surface):
read:org— required bylist-prs.shif it enumerates org PRs; missing scope would cause silent empty results or403on org API callsworkflow— required if the bot PAT ever needs to trigger other workflows via API; not currently used but noted as a common gapGEMINI_CLI_TRUST_WORKSPACE=trueis set but no quota/billing errors are visibleRecommendations:
gh auth statusoutput as a dedicated early step (orecho "::debug::$(gh auth status 2>&1)") so scope issues appear in every run logDON_PETRY_BOT_GH_PATincludesrepo+read:org+read:discussionif cross-org PR listing is used5. Recommendations
[CRITICAL] Capture logs before they go missing — add a failure artifact upload step
.github/workflows/pr-review.ymlafter theReview each PR (cascade)step:[HIGH] Add
gh auth statusas an explicit early stepEnumerate candidate PRs:[HIGH] Investigate the 11 cancelled runs
gh workflow cancel. Determine if cancellations are intentional queue flushes or a bug in a caller.cancel-in-progress: false, queued runs are cancellable).[MEDIUM] Add spot-instance resilience to the npm install step
npm install -gcalls inreview-batch.shor the workflow step with a retry loop:[LOW] Scope the concurrency group more tightly for check_suite batch fallthrough
pr-review-batchwhen no PR URL is present. Multiple simultaneous check_suite events on non-PR branches all share this slot and queue. Consider adding a timestamp or run-id suffix to the batch group to allow parallelism.6. Health Score
Health: 7/10 — Failure rate is low (3%) and self-recovering, but all failures are completely undiagnosable due to missing logs, making it impossible to confirm the root cause or detect a silent worsening trend.