feat: implement issue #233 — Phase 2: CI Failure Analyst gh-aw workflow - #307
Conversation
|
Warning Rate limit exceeded
You’ve run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (3)
📝 WalkthroughWalkthroughThis PR implements Phase 2 of the GitHub Agentic Workflows rollout by introducing the CI Failure Analyst workflow. When a CI check fails on a PR, the workflow automatically fetches run logs, identifies the failing step, classifies the root cause into one of six categories, and posts a diagnostic comment with remediation guidance. The implementation includes workflow specification, compiled pipeline, infrastructure setup, and complete test and user documentation. ChangesCI Failure Analyst Agentic Workflow
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related issues
Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Dev-Lead — fix-bot-comment (no-changes)Engine ran but made no changes. |
There was a problem hiding this comment.
Code Review
This pull request introduces the CI Failure Analyst agentic workflow, including its documentation, test scenarios, and necessary configuration updates for Dependabot and Git attributes. The review feedback identifies several critical issues: the documentation is missing the required YAML frontmatter needed for the compilation tool, the compiled lock file for the workflow is missing from the PR, and a non-standard wildcard in the Dependabot configuration should be corrected to ensure it is interpreted correctly.
|
@dev-lead please fix the YAML lint failures in
The lock file was auto-generated by |
Dev-Lead — fix-reviews (applied)Changes committed and pushed. |
There was a problem hiding this comment.
Pull request overview
Implements Issue #233 (Phase 2) by adding a new CI Failure Analyst agentic workflow (gh-aw) that triggers on check_run.completed and posts (staged) diagnostic PR comments for CI failures, along with supporting documentation and CI validation.
Changes:
- Added scenario spec + documentation for the CI Failure Analyst workflow.
- Added
ci-failure-analystgh-aw source workflow (.md) and its compiled lock workflow (.lock.yml). - Added CI job to compile agentic workflows, plus supporting repo config (
dependabotignore, actions lock,.gitattributesfor generated locks).
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
tests/aw/ci-failure-analyst/scenarios.md |
Scenario/spec-first test cases for expected bot behavior and edge cases. |
docs/aw/ci-failure-analyst.md |
User-facing documentation for trigger, categories, comment format, and setup. |
.github/workflows/ci.yml |
Adds CI job to run gh aw compile --no-emit to validate workflows compile. |
.github/workflows/ci-failure-analyst.md |
New gh-aw workflow definition/instructions for diagnosing failed check runs. |
.github/workflows/ci-failure-analyst.lock.yml |
Compiled gh-aw workflow generated from the .md source. |
.github/dependabot.yml |
Adjusts dependabot config and ignores gh-aw-managed actions. |
.github/aw/actions-lock.json |
Introduces an actions lock file for gh-aw-managed action pinning. |
.gitattributes |
Marks *.lock.yml as generated and forces merge=ours. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Dev-Lead — fix-bot-comment (no-changes)Engine ran but made no changes. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b07087ce54
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
.github/workflows/ci.yml (1)
33-47:⚠️ Potential issue | 🟠 Major | ⚡ Quick winExclude generated lock files from yamllint to prevent CI breakage.
The
.github/workflows/linting scope includesci-failure-analyst.lock.yml, a generated file with lines significantly exceeding the 250-character limit (e.g., line 2: 1458 chars). This causes persistent pipeline failure. The ignore rule prevents linting generated artifacts while maintaining validation of authored workflows.Suggested change
- name: Lint YAML run: | pip install --quiet --only-binary :all: "yamllint==1.38.0" # v1.38.0 yamllint -c <(cat <<'YAMLLINTRC' extends: default + ignore: | + .github/workflows/*.lock.yml rules: line-length: max: 250 truthy: check-keys: false document-start: disable comments: min-spaces-from-content: 1 YAMLLINTRC ) .github/workflows/🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/ci.yml around lines 33 - 47, The CI yamllint invocation is linting generated lock files (e.g., ci-failure-analyst.lock.yml) which exceed the line-length rule; update the inline yamllint config passed to yamllint in the workflow run block to add an ignore entry (for example ignore: '*.lock.yml' or ignore: 'ci-failure-analyst.lock.yml') so those generated files in .github/workflows/ are skipped while keeping the existing rules (line-length, truthy, document-start, comments) intact.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/aw/ci-failure-analyst.md`:
- Line 65: Docs claim the safe-outputs job has pull-requests: write but the
workflow sets permissions: {}; fix by either updating the documentation line to
state that comment posting is done via the GH_AW_GITHUB_TOKEN token (token-based
write) or change the workflow job safe-outputs to grant pull-requests: write so
the docs match; locate the safe-outputs job and the sentence "Write access
(posting the comment) is handled by the `safe-outputs` job with its own scoped
`pull-requests: write` permission" and update one side (docs or workflow) so
both reference the same auth path (either GH_AW_GITHUB_TOKEN token-based write
or explicit pull-requests: write permission).
---
Outside diff comments:
In @.github/workflows/ci.yml:
- Around line 33-47: The CI yamllint invocation is linting generated lock files
(e.g., ci-failure-analyst.lock.yml) which exceed the line-length rule; update
the inline yamllint config passed to yamllint in the workflow run block to add
an ignore entry (for example ignore: '*.lock.yml' or ignore:
'ci-failure-analyst.lock.yml') so those generated files in .github/workflows/
are skipped while keeping the existing rules (line-length, truthy,
document-start, comments) intact.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: ASSERTIVE
Plan: Pro
Run ID: 93184206-3757-4a97-b5b8-b2416288593e
📒 Files selected for processing (8)
.gitattributes.github/aw/actions-lock.json.github/dependabot.yml.github/workflows/ci-failure-analyst.lock.yml.github/workflows/ci-failure-analyst.md.github/workflows/ci.ymldocs/aw/ci-failure-analyst.mdtests/aw/ci-failure-analyst/scenarios.md
|
@dev-lead the YAML lint is still failing on
Please fix the lock file directly:
Do NOT regenerate with |
|
@coderabbitai resolve |
Dev-Lead — fix-reviews (applied)Changes committed and pushed. |
✅ Actions performedComments resolved and changes approved. |
Dev-Lead — human-pr (no-changes)No changes were needed for this PR. |
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fae5127896
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| check_run: | ||
| types: | ||
| - completed |
There was a problem hiding this comment.
Recompile lock file with
workflow_run trigger
The executable workflow still listens on check_run, while the prompt logic in ci-failure-analyst.md reads github.event.workflow_run.*; with this mismatch, the runtime guard treats the expected workflow_run.conclusion as missing and the analyst will no-op instead of diagnosing CI failures. Fresh evidence in this commit is that the compiled .lock.yml (the file GitHub actually runs) still declares on: check_run, so the trigger/payload contract is inconsistent at runtime.
Useful? React with 👍 / 👎.
| ```json | ||
| { | ||
| "action": "completed", | ||
| "check_run": { |
There was a problem hiding this comment.
Align scenario inputs with the
workflow_run event payload
The scenario spec still models check_run payload fields, but the workflow source now triggers on workflow_run and reads github.event.workflow_run.*; this means the documented test cases validate a different event shape than production, so passing the spec would not verify the real execution path for PR association, run URL, or conclusion guards.
Useful? React with 👍 / 👎.



Closes #233
Implemented by dev-lead agent. Please review.
Summary by CodeRabbit
New Features
Documentation
Tests