Convergence latency (deterministic)
Workflow duration across 88 completed run(s) in the window — computed by pr_review_health.sh, not the model.
| Workflow |
Runs |
min |
p50 |
p95 |
max |
pr-review-trigger.yml |
88 |
5s |
30s |
5m2s |
7m23s |
Run-outcome mix by triggering event
pr-review run conclusions over the window, split by the event that triggered each run — computed deterministically by pr-review-outcomes.sh, not the model. A high cancelled share means runs were evicted while PENDING (concurrency admits one running + one pending per group); a cancelled run is non-blocking but historically rendered as a failing check (#1421).
| Event |
success |
cancelled |
skipped |
failure |
total |
cancel % |
repository_dispatch |
46 |
8 |
0 |
1 |
55 |
14% |
workflow_dispatch |
32 |
0 |
0 |
1 |
33 |
0% |
| TOTAL |
78 |
8 |
0 |
2 |
88 |
9% |
Sweep hit-rate (deterministic)
How often the pr-review-sweep scheduled backstop actually had to re-review a PR — computed deterministically by pr-review-sweep-metrics.sh, not the model. The scheduled (cron) path re-dispatches only for genuinely un-eventable cases (GitHub-App / cross-repo checks like SonarCloud that emit no workflow_run); the workflow_run fast path handles every eventable case, so a scheduled re-dispatch is always a case the fast path did not cover.
| Metric |
Value |
| Scheduled sweep ticks |
8 |
| Ticks that re-dispatched a review (hits) |
8 |
| Sweep hit-rate |
100% |
Interpretation: a low rate confirms the timer stays exception-only (the event fast path is carrying the load); a rising rate is a signal that an eventable case is leaking onto the timer — a hole in the narrowing worth investigating.
### 1. Executive Summary
**Status:** WARNING
**Period:** 2026-09-06T12:49:27Z – 2026-09-07T10:33:27Z
**Result:** 2 of 88 runs failed (2.3%) — 8 additional runs cancelled (9.1%)
Key findings:
- Failed run logs were not captured or are unavailable — root cause of #72333 and #72347 is undiagnosable from this report
- 8 cancellations cluster in two tight bursts (10:22–10:29 UTC on 09-07), consistent with a concurrency-group cancel-in-progress strategy firing on rapid push events
- Both failures are isolated (surrounded by successes), suggesting intermittent rather than systemic breakage
Action required: Retrieve and attach logs for runs #72333 and #72347 to determine failure root cause before the next analysis window.
---
### 2. Failure Breakdown
| Failure Category | Affected Runs | Example Error Message |
|---|---|---|
| Unknown (no logs captured) | #72333, #72347 | *(log content not provided)* |
| Concurrency cancellation (expected) | #72381–#72383, #72389–#72390, #72393–#72395 | *(cancelled by a newer run in the same concurrency group)* |
---
### 3. Error Patterns
**Failures (#72333, #72347):**
No log content was provided for either failed run. Without step-level output, failure category, error message, and root cause cannot be determined. The two failures occurred ~2.5 hours apart (18:21 and 20:44 UTC on 09-06), which rules out a single event as the trigger. The pattern is consistent with transient external dependency errors (API rate limits, token expiry, ephemeral network issues), but this cannot be confirmed without logs.
**Cancellations (#72381, #72382, #72383, #72389, #72390, #72393, #72394, #72395):**
All 8 cancellations fall within a 7-minute window (10:22–10:29 UTC on 09-07), interleaved with successful runs at nearly the same timestamps. This is the expected behavior of a `concurrency: cancel-in-progress: true` group: rapid `pull_request` or `check_suite` events queue many runs, and each new run cancels its predecessor. The presence of successes immediately after each cancellation burst confirms the workflow ultimately completes. These are not actionable failures.
---
### 4. Token Scope Analysis
**Scopes granted by workflow source (`pr-review-trigger.yml`):**
| Scope | Level | Purpose |
|---|---|---|
| `contents: read` | job | Checkout, read repo files |
| `pull-requests: write` | job | Post review comments, approve/request changes |
| `checks: read` | job | Read check suite status (trigger source) |
**Scopes that appear missing or insufficient:**
Cannot be determined — no `gh auth status` output or permission-denied errors appear in the provided logs (logs for failed runs were not captured).
**Recommendations (based on workflow source alone):**
- `pull-requests: write` is correctly present for a review agent.
- If the reusable `pr-review.yml` dispatches further workflows or creates issues, `issues: write` or `actions: write` may be needed at the caller level — verify against the reusable's documented required permissions.
- If failures in #72333/#72347 turn out to be 403s, the first suspect should be whether `secrets: inherit` is correctly forwarding a PAT or App token with the org-level scopes the reusable expects.
---
### 5. Recommendations
1. **Retrieve logs for failed runs #72333 and #72347**
- What: Run `gh run view 72333 --log-failed` and `gh run view 72347 --log-failed` in `petry-projects/.github-private`.
- Why: The entire failure analysis section is blocked without this data. Every other recommendation below is conditional on what these logs show.
- Expected impact: Unblocks root cause identification and targeted fixes.
- Urgency: **CRITICAL** (report is incomplete without it)
2. **Confirm concurrency-group cancel behaviour is intentional**
- What: Check `pr-review.yml@pr-review/v1-next` for a `concurrency:` block; if absent, check `pr-review-trigger.yml` (none visible in the provided source).
- Why: 9% cancellation rate is acceptable if cancel-in-progress is deliberate, but if it is caused by an external workflow cancellation signal, it may represent lost reviews.
- Expected impact: Either confirms the cancellations are benign or surfaces a mis-configured concurrency scope.
- Urgency: **MEDIUM**
3. **Add structured log export to failed runs**
- What: In `pr-review.yml` (the reusable), add a step that uploads the review engine log as an artifact on failure (`if: failure()`).
- Why: The current setup produces no surfaceable log content when the called workflow fails, as seen in this report — the trigger stub has no logging of its own and the reusable's logs are not forwarded.
- Expected impact: Future health reports will have actionable error messages rather than empty log sections.
- Urgency: **HIGH**
4. **Add a failure notification or alerting threshold**
- What: If failure rate exceeds 5% in any rolling 24-hour window, trigger a Slack or GitHub issue alert (e.g., via a scheduled health-check workflow already in this repo).
- Why: Two failures in 24 hours is below the alert threshold today, but both occurred within a 2.5-hour window — a rising trend would be invisible until the next manual report.
- Expected impact: Faster mean-time-to-detect for regressions introduced by new `pr-review/v1-next` channel releases (this repo is ring-0 canary).
- Urgency: **LOW**
---
### 6. Health Score
`Health: 8/10 — Failure rate is low (2.3%) and cancellations appear benign, but missing logs make the two failures undiagnosable and leave root cause unknown.`
Convergence latency (deterministic)
Workflow duration across 88 completed run(s) in the window — computed by
pr_review_health.sh, not the model.pr-review-trigger.ymlRun-outcome mix by triggering event
pr-review run conclusions over the window, split by the event that triggered each run — computed deterministically by
pr-review-outcomes.sh, not the model. A highcancelledshare means runs were evicted while PENDING (concurrency admits one running + one pending per group); a cancelled run is non-blocking but historically rendered as a failing check (#1421).repository_dispatchworkflow_dispatchSweep hit-rate (deterministic)
How often the pr-review-sweep scheduled backstop actually had to re-review a PR — computed deterministically by
pr-review-sweep-metrics.sh, not the model. The scheduled (cron) path re-dispatches only for genuinely un-eventable cases (GitHub-App / cross-repo checks like SonarCloud that emit noworkflow_run); theworkflow_runfast path handles every eventable case, so a scheduled re-dispatch is always a case the fast path did not cover.Interpretation: a low rate confirms the timer stays exception-only (the event fast path is carrying the load); a rising rate is a signal that an eventable case is leaking onto the timer — a hole in the narrowing worth investigating.