feat: implement issue #1019 — [dev-lead timeout C/3] Fleet-monitor observability: surface timeout count/rate - #1024
Conversation
…servability: surface timeout count/rate
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reached
Next review available in: 44 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (3)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Dev-Lead — review-changes (no-changes)No changes were needed for this PR. |
|
Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-07-02T23:44:00Z. |
Dev-Lead — fix-bot-comment (no-changes)Agent reasoning |
There was a problem hiding this comment.
Code Review
This pull request introduces Dev-Lead timeout observability by parsing and tallying status markers from issue comments to report timeout rates across the fleet. The changes include adding parsing and reporting logic in scripts/fleet_report.sh, integrating it into scripts/fleet_monitor.sh, and adding comprehensive BATS tests. The reviewer feedback focuses on improving robustness, specifically by refining the jq regular expression to handle tight HTML comments, adding default values to Bash arithmetic expansions to prevent syntax errors, and utilizing $BATS_TEST_TMPDIR in BATS tests to ensure proper cleanup of temporary files.
There was a problem hiding this comment.
Pull request overview
Implements Story C for #1019 by adding Dev-Lead timeout (“wall”) observability to the Actions Fleet Monitor: it now counts <!-- dev-lead-issue ... reason=timeout ... --> status markers over the lookback window (per repo) and surfaces timeout counts + overall timeout rate distinctly from generic reason=engine-error in both the step summary and the tracked issue body.
Changes:
- Add
summarize_dev_lead_timeoutsandgenerate_dev_lead_timeout_reporthelpers to parse Dev-Lead status markers and render a per-repo + fleet-total timeout section. - Extend
fleet_monitor.shto fetch windowed issue comments per repo (issues/comments?since=), aggregate timeout/engine-error totals, and append the new section to both report outputs. - Add Bats coverage for the new summarization and reporting helpers.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
| tests/fleet_report.bats | Adds unit tests validating timeout/engine-error/total tallying and the rendered timeout-rate report output. |
| scripts/fleet_report.sh | Introduces Dev-Lead marker prefix constant and new helper functions to compute and render timeout observability. |
| scripts/fleet_monitor.sh | Collects per-repo issue-comment markers over the window and appends the Dev-Lead timeout section to the generated reports. |
…pdirs Address Gemini review on #1024: - reason regex: capture reason=(?<r>[a-z]+(?:-[a-z]+)*) instead of [^ ]+, so a marker without a space before --> (e.g. reason=timeout-->) yields 'timeout', not a token with trailing punctuation, while still matching engine-error / rate-limited / missing-binary. - fleet-total arithmetic: default the TSV fields (${t:-0} etc.) so a malformed row can't break the sum under set -e. - tests: create temp files under $BATS_TEST_TMPDIR (BATS auto-cleans) instead of bare mktemp, so a failing assertion can't orphan temp files. shellcheck clean; 8/8 dev-lead-timeout bats pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HUhVsBbihbTjVvbcVf1WuU
|
Dev-Lead — fix-bot-comment (no-changes)Agent reasoning |
donpetry-bot
left a comment
There was a problem hiding this comment.
Automated review — APPROVED ✓
Risk: MEDIUM
Reviewed commit: ad25af2da52368e032acfae9aaf8fe9ac5edd520
Review mode: triage-approved (single reviewer)
Summary
Adds dev-lead timeout observability (#1019, story C of #901) to the fleet monitor: summarize_dev_lead_timeouts tallies dev-lead status markers by reason= from windowed per-repo issue comments, and generate_dev_lead_timeout_report renders a per-repo table with fleet totals and a timeout rate in both the Step Summary and the report/issue body. 8 new BATS tests cover the tally and rendering paths. Confirms the triage assessment; all 5 bot review findings were fixed in ad25af2 and their threads resolved.
Linked issue analysis
Issue #1019 acceptance criteria are all met: (1) reason=timeout count/rate is reported per repo, broken out from engine-error; (2) the section appears in $GITHUB_STEP_SUMMARY and the report file reused as the tracked issue body; (3) existing token-usage / high-failure tracking is untouched (purely additive section + temp-file cleanup); (4) shellcheck and bats are green and new tests are registered in tests/fleet_report.bats. The issue scope mentions the workflow file, but no workflow change was needed since it already invokes fleet_monitor.sh.
Findings
No blocking findings. Verified: the jq capture regex ([a-z]+(?:-[a-z]+)*) cannot absorb the trailing HTML-comment '-->' (addresses the gemini finding); arithmetic uses ${t:-0}-style defaults, safe under set -e; untrusted comment bodies are only pattern-matched/counted via jq, never executed or echoed into the report; BATS temp files now live in $BATS_TEST_TMPDIR. Minor non-blocking note: the per-repo comments fetch (gh api --paginate per repo) adds API cost proportional to fleet size, mitigated by since= filtering and best-effort per-repo warnings. Secret scan: run_secret_scanning MCP tool unavailable this run; gitleaks CI check is green and the diff contains no credential-like content.
CI status
All substantive checks green: shellcheck, bats, unit-tests, Lint, CodeQL (actions+python), gitleaks secret scan, SonarCloud, AgentShield, Agent Security Scan, guard, holdout-guard, template-drift, gh-aw-compile, validate-agent-profiles. Cancelled entries are superseded dev-lead/review orchestration runs, not code checks. mergeStateStatus=BLOCKED only pending review approval.
Reviewed automatically by the PR-review agent (single-reviewer mode: fable 5). Reply if you need a human review.
donpetry-bot
left a comment
There was a problem hiding this comment.
Automated review — APPROVED ✓
Risk: MEDIUM
Reviewed commit: ad25af2da52368e032acfae9aaf8fe9ac5edd520
Review mode: triage-approved (single reviewer)
Summary
Adds Dev-Lead timeout (wall) observability to the fleet monitor: a new section in fleet_monitor.sh pulls windowed issue comments per repo, and two new fleet_report.sh helpers (summarize_dev_lead_timeouts, generate_dev_lead_timeout_report) tally status markers by reason= and render a per-repo table with fleet totals and a timeout rate. 8 new bats tests cover the tallying and rendering paths. Triage assessment confirmed — the change is well-scoped, defensive (best-effort per-repo comment reads, empty-section omission, arithmetic defaults), and free of injection or secret-handling concerns.
Linked issue analysis
Closes #1019 (Story C of #901). Acceptance criteria met: (1) reason=timeout count/rate is reported distinct from engine-error via summarize_dev_lead_timeouts + the rendered rate line; (2) the section is appended to both $GITHUB_STEP_SUMMARY and the report file used as the tracked-issue body, reusing the existing aggregation/issue steps; (3) existing token-usage/high-failure tracking is untouched (section is additive and self-omitting when empty); (4) shellcheck and bats are green with 8 new registered tests. The issue mentioned actions-fleet-monitor.yml in scope, but no workflow change was needed since the existing summary/issue steps carry the new section.
Findings
No blocking findings.
- Secret scan: the run_secret_scanning MCP tool is not available in this session; relying on the passing gitleaks CI check. No credentials or tokens appear in the diff.
- All 5 prior review threads (gemini-code-assist: jq capture regex over-matching, arithmetic defaults under set -e, bats mktemp hygiene) are resolved; the final code uses an even stricter reason regex ([a-z]+(?:-[a-z]+)*) than the suggested fix.
- Nit (non-blocking): the thread replies say the redundant rm -f calls were removed, but the new tests still carry rm -f "$f" after moving to $BATS_TEST_TMPDIR — harmless redundancy. Similarly, printf in generate_dev_lead_timeout_report uses bare $t/$e/$tot (empty cells render blank) while only the arithmetic uses :-0 defaults; the arithmetic guard is what matters for set -e.
CI status
All code checks green: shellcheck, ShellCheck, bats, unit-tests, Lint, CodeQL (actions + python), SonarCloud, gitleaks, agent-shield, Agent Security Scan, holdout-guard, template-drift, gh-aw-compile, validate-agent-profiles all SUCCESS. CANCELLED entries are superseded dev-lead/review orchestration dispatch runs, not code checks; dependency-audit ecosystem jobs SKIPPED (no matching ecosystems). CodeRabbit approved.
Reviewed automatically by the PR-review agent (single-reviewer mode: fable 5). Reply if you need a human review.



Closes #1019
Implemented by dev-lead agent. Please review.