Skip to content

PR Review Agent — failures detected 2026-09-07 #1682

Description

@github-actions

Convergence latency (deterministic)

Workflow duration across 88 completed run(s) in the window — computed by pr_review_health.sh, not the model.

Workflow Runs min p50 p95 max
pr-review-trigger.yml 88 5s 30s 5m2s 7m23s

Run-outcome mix by triggering event

pr-review run conclusions over the window, split by the event that triggered each run — computed deterministically by pr-review-outcomes.sh, not the model. A high cancelled share means runs were evicted while PENDING (concurrency admits one running + one pending per group); a cancelled run is non-blocking but historically rendered as a failing check (#1421).

Event success cancelled skipped failure total cancel %
repository_dispatch 46 8 0 1 55 14%
workflow_dispatch 32 0 0 1 33 0%
TOTAL 78 8 0 2 88 9%

Sweep hit-rate (deterministic)

How often the pr-review-sweep scheduled backstop actually had to re-review a PR — computed deterministically by pr-review-sweep-metrics.sh, not the model. The scheduled (cron) path re-dispatches only for genuinely un-eventable cases (GitHub-App / cross-repo checks like SonarCloud that emit no workflow_run); the workflow_run fast path handles every eventable case, so a scheduled re-dispatch is always a case the fast path did not cover.

Metric Value
Scheduled sweep ticks 8
Ticks that re-dispatched a review (hits) 8
Sweep hit-rate 100%

Interpretation: a low rate confirms the timer stays exception-only (the event fast path is carrying the load); a rising rate is a signal that an eventable case is leaking onto the timer — a hole in the narrowing worth investigating.

### 1. Executive Summary

**Status:** WARNING
**Period:** 2026-09-06T12:49:27Z – 2026-09-07T10:33:27Z
**Result:** 2 of 88 runs failed (2.3%) — 8 additional runs cancelled (9.1%)

Key findings:
- Failed run logs were not captured or are unavailable — root cause of #72333 and #72347 is undiagnosable from this report
- 8 cancellations cluster in two tight bursts (10:22–10:29 UTC on 09-07), consistent with a concurrency-group cancel-in-progress strategy firing on rapid push events
- Both failures are isolated (surrounded by successes), suggesting intermittent rather than systemic breakage

Action required: Retrieve and attach logs for runs #72333 and #72347 to determine failure root cause before the next analysis window.

---

### 2. Failure Breakdown

| Failure Category | Affected Runs | Example Error Message |
|---|---|---|
| Unknown (no logs captured) | #72333, #72347 | *(log content not provided)* |
| Concurrency cancellation (expected) | #72381–#72383, #72389–#72390, #72393–#72395 | *(cancelled by a newer run in the same concurrency group)* |

---

### 3. Error Patterns

**Failures (#72333, #72347):**
No log content was provided for either failed run. Without step-level output, failure category, error message, and root cause cannot be determined. The two failures occurred ~2.5 hours apart (18:21 and 20:44 UTC on 09-06), which rules out a single event as the trigger. The pattern is consistent with transient external dependency errors (API rate limits, token expiry, ephemeral network issues), but this cannot be confirmed without logs.

**Cancellations (#72381, #72382, #72383, #72389, #72390, #72393, #72394, #72395):**
All 8 cancellations fall within a 7-minute window (10:22–10:29 UTC on 09-07), interleaved with successful runs at nearly the same timestamps. This is the expected behavior of a `concurrency: cancel-in-progress: true` group: rapid `pull_request` or `check_suite` events queue many runs, and each new run cancels its predecessor. The presence of successes immediately after each cancellation burst confirms the workflow ultimately completes. These are not actionable failures.

---

### 4. Token Scope Analysis

**Scopes granted by workflow source (`pr-review-trigger.yml`):**

| Scope | Level | Purpose |
|---|---|---|
| `contents: read` | job | Checkout, read repo files |
| `pull-requests: write` | job | Post review comments, approve/request changes |
| `checks: read` | job | Read check suite status (trigger source) |

**Scopes that appear missing or insufficient:**
Cannot be determined — no `gh auth status` output or permission-denied errors appear in the provided logs (logs for failed runs were not captured).

**Recommendations (based on workflow source alone):**
- `pull-requests: write` is correctly present for a review agent.
- If the reusable `pr-review.yml` dispatches further workflows or creates issues, `issues: write` or `actions: write` may be needed at the caller level — verify against the reusable's documented required permissions.
- If failures in #72333/#72347 turn out to be 403s, the first suspect should be whether `secrets: inherit` is correctly forwarding a PAT or App token with the org-level scopes the reusable expects.

---

### 5. Recommendations

1. **Retrieve logs for failed runs #72333 and #72347**
   - What: Run `gh run view 72333 --log-failed` and `gh run view 72347 --log-failed` in `petry-projects/.github-private`.
   - Why: The entire failure analysis section is blocked without this data. Every other recommendation below is conditional on what these logs show.
   - Expected impact: Unblocks root cause identification and targeted fixes.
   - Urgency: **CRITICAL** (report is incomplete without it)

2. **Confirm concurrency-group cancel behaviour is intentional**
   - What: Check `pr-review.yml@pr-review/v1-next` for a `concurrency:` block; if absent, check `pr-review-trigger.yml` (none visible in the provided source).
   - Why: 9% cancellation rate is acceptable if cancel-in-progress is deliberate, but if it is caused by an external workflow cancellation signal, it may represent lost reviews.
   - Expected impact: Either confirms the cancellations are benign or surfaces a mis-configured concurrency scope.
   - Urgency: **MEDIUM**

3. **Add structured log export to failed runs**
   - What: In `pr-review.yml` (the reusable), add a step that uploads the review engine log as an artifact on failure (`if: failure()`).
   - Why: The current setup produces no surfaceable log content when the called workflow fails, as seen in this report — the trigger stub has no logging of its own and the reusable's logs are not forwarded.
   - Expected impact: Future health reports will have actionable error messages rather than empty log sections.
   - Urgency: **HIGH**

4. **Add a failure notification or alerting threshold**
   - What: If failure rate exceeds 5% in any rolling 24-hour window, trigger a Slack or GitHub issue alert (e.g., via a scheduled health-check workflow already in this repo).
   - Why: Two failures in 24 hours is below the alert threshold today, but both occurred within a 2.5-hour window — a rising trend would be invisible until the next manual report.
   - Expected impact: Faster mean-time-to-detect for regressions introduced by new `pr-review/v1-next` channel releases (this repo is ring-0 canary).
   - Urgency: **LOW**

---

### 6. Health Score

`Health: 8/10 — Failure rate is low (2.3%) and cancellations appear benign, but missing logs make the two failures undiagnosable and leave root cause unknown.`

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    automated-reportCreated by automated workflowhealth-checkAutomated health check report

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions