Skip to content

PR Review Agent — failures detected 2026-05-21 #335

Description

@github-actions

1. Executive Summary

Status: WARNING
Period: 2026-05-20 12:40Z – 2026-05-21 03:01Z
Result: 3 of 100 runs failed (3%) — 70 succeeded, 11 cancelled, 16 skipped

Key findings:

  • All 3 failure logs are unavailable (log unavailable), preventing definitive root-cause identification
  • Failures cluster tightly in the 01:10–03:01 UTC window on 2026-05-21 and are non-consecutive (each followed immediately by successful retriggers), strongly suggesting transient infrastructure or rate-limit events
  • 11 cancelled runs warrant investigation — cancel-in-progress: false means the concurrency slot is not auto-cancelling; manual cancellations or upstream queue flushes are likely

Action required: Investigate why GitHub is not retaining logs for runs #1527, #1531, #1537 — enable log artifact upload or step-level debug output so failures are diagnosable.


2. Failure Breakdown

Failure Category Affected Runs Example Error Message
Log unavailable / retention gap #1527, #1531, #1537 (log unavailable for run 26199375580 / 26199692481 / 26202833228)
Transient infra / rate limit (suspected) #1527, #1531, #1537 Unknown — inferred from recovery pattern (each failure bracketed by successes)
Cancellation (origin unclear) #1440, #1441, #1450, #1470, #1490, #1491, #1492, #1502, #1503, #1511, #1521 (cancelled — no log)

3. Error Patterns

Pattern A — Log retention gap (confirmed)

(log unavailable for run 26202833228)
(log unavailable for run 26199692481)
(log unavailable for run 26199375580)
  • Step: Any — log was never captured or was purged before retrieval
  • Root cause: GitHub Actions log retention is 90 days by default, so age is not the cause here. More likely: the runner was evicted mid-job before streaming finished, OR the job failed during setup (before any step output flushed), OR the log fetch API request raced with job cleanup. A runner OOM or spot-instance preemption during the Install review engine CLIs step (which runs npm install -g commands) is a plausible trigger.
  • Recovery signal: In all three cases, the immediately following run(s) at the same concurrency slot succeeded — consistent with transient infrastructure events (spot preemption, ephemeral network blip, npm registry hiccup).

Pattern B — Cancellations with cancel-in-progress: false

  • Step: Pre-queue / runner allocation
  • Root cause: With cancel-in-progress: false, GitHub will not auto-cancel queued runs. The 11 cancellations are therefore either (a) manual workflow cancellations triggered by an operator or upstream automation, or (b) the concurrency group resolved to pr-review-batch for multiple simultaneous non-PR-specific triggers and an operator flushed the queue. No evidence of a code defect here.

4. Token Scope Analysis

Credential Configured Via Scopes Needed Status
DON_PETRY_BOT_GH_PAT → GH_TOKEN secrets.DON_PETRY_BOT_GH_PAT repo, read:org, pull-requests: write (workflow-level) Cannot verify — logs unavailable; no gh auth status output present
CLAUDE_CODE_OAUTH_TOKEN secrets.CLAUDE_CODE_OAUTH_TOKEN Claude API access Cannot verify — no auth output in logs
COPILOT_GITHUB_TOKEN secrets.GH_PAT read:user, Copilot subscription Cannot verify
GOOGLE_API_KEY / GEMINI_API_KEY secrets.GOOGLE_API_KEY (shared) Gemini API enabled Cannot verify

Scopes confirmed present: None — no gh auth status output is available in any retrieved log.

Scopes potentially missing (inferred from workflow risk surface):

  • read:org — required by list-prs.sh if it enumerates org PRs; missing scope would cause silent empty results or 403 on org API calls
  • workflow — required if the bot PAT ever needs to trigger other workflows via API; not currently used but noted as a common gap
  • Gemini CLI may require project-level enablement beyond API key — GEMINI_CLI_TRUST_WORKSPACE=true is set but no quota/billing errors are visible

Recommendations:

  • Add gh auth status output as a dedicated early step (or echo "::debug::$(gh auth status 2>&1)") so scope issues appear in every run log
  • Validate DON_PETRY_BOT_GH_PAT includes repo + read:org + read:discussion if cross-org PR listing is used

5. Recommendations

  1. [CRITICAL] Capture logs before they go missing — add a failure artifact upload step

    • What: Append to .github/workflows/pr-review.yml after the Review each PR (cascade) step:
      - name: Upload debug log on failure
        if: failure()
        uses: actions/upload-artifact@v4
        with:
          name: pr-review-debug-${{ github.run_number }}
          path: |
            prs.txt
            /tmp/review-*.log
          retention-days: 14
    • Why: All 3 failures produced no diagnosable output. Without artifact capture, future failures will also be undiagnosable.
    • Expected impact: 100% of future failures become diagnosable.
    • Urgency: CRITICAL
  2. [HIGH] Add gh auth status as an explicit early step

    • What: Insert before Enumerate candidate PRs:
      - name: Verify auth scopes
        run: gh auth status
    • Why: Token scope errors are silent and hard to distinguish from logic failures without this output.
    • Expected impact: Auth regressions (e.g. secret rotation breaking scope) become immediately visible.
    • Urgency: HIGH
  3. [HIGH] Investigate the 11 cancelled runs

    • What: Cross-reference the cancellation timestamps against operator activity logs and any upstream automation that calls gh workflow cancel. Determine if cancellations are intentional queue flushes or a bug in a caller.
    • Why: 11% cancellation rate is non-trivial and may represent lost review coverage if the cancellations interrupt in-progress reviews (even with cancel-in-progress: false, queued runs are cancellable).
    • Expected impact: Confirm whether PRs are receiving reviews after a cancel event or falling through the gap.
    • Urgency: HIGH
  4. [MEDIUM] Add spot-instance resilience to the npm install step

    • What: Wrap the npm install -g calls in review-batch.sh or the workflow step with a retry loop:
      for i in 1 2 3; do npm install -g "@anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}" && break || sleep 15; done
    • Why: If failures are caused by transient npm registry or runner preemption events, a simple retry prevents full run failure.
    • Expected impact: Reduces transient failure rate.
    • Urgency: MEDIUM
  5. [LOW] Scope the concurrency group more tightly for check_suite batch fallthrough

    • What: The current logic falls back to pr-review-batch when no PR URL is present. Multiple simultaneous check_suite events on non-PR branches all share this slot and queue. Consider adding a timestamp or run-id suffix to the batch group to allow parallelism.
    • Why: The 11 cancellations may partly stem from batch-slot contention.
    • Expected impact: Reduces unnecessary queuing and manual cancellation needs.
    • Urgency: LOW

6. Health Score

Health: 7/10 — Failure rate is low (3%) and self-recovering, but all failures are completely undiagnosable due to missing logs, making it impossible to confirm the root cause or detect a silent worsening trend.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

automated-reportCreated by automated workflowhealth-checkAutomated health check report

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions