Skip to content

feat: implement issue #1727 — [Phase 2] Retarget run attribution to job/role level so the collapse doesn't blind the fleet monitor - #1789

Merged
don-petry merged 14 commits into
mainfrom
dev-lead/issue-1727-20260913-0408
Sep 21, 2026
Merged

don-petry merged 14 commits into
mainfrom
dev-lead/issue-1727-20260913-0408

Conversation

@don-petry

@don-petry don-petry commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

User description

Closes #1727

Implemented by dev-lead agent. Please review.

Summary by CodeRabbit

  • New Features

    • Fleet monitoring now reports completed agent-ingress runs by role and outcome.
    • Pull request review health reports can include attribution for the configured review role.
    • Runs with missing role information are clearly identified as unattributed.
  • Tests

    • Added comprehensive coverage for run attribution, outcome handling, aggregation, and report generation.
    • Expanded lint workflow validation to include the attribution test suite.

CodeAnt-AI Description

Preserve per-role run health after workflows move to shared agent ingress

What Changed

  • Fleet monitoring attributes shared-ingress runs to their role-bearing jobs and combines them with legacy per-role runs
  • PR-review health reports retain the pr-review-mention signal after the workflow collapse
  • Missing or unnamed roles appear as UNATTRIBUTED instead of disappearing, while skipped roles are excluded
  • Reports include per-role totals, successes, failures, and cancellations, with warnings when run data cannot be read
  • Added automated coverage for attribution, outcome aggregation, migration parity, missing-role handling, and report output

Impact

✅ Accurate per-role fleet failure counts after workflow collapse
✅ Preserved PR-review health signals across workflow migration
✅ Visible missing-run attribution instead of silent metric loss

💡 Usage Guide

Checking Your Pull Request

Every time you make a pull request, our system automatically looks through it. We check for security issues, mistakes in how you're setting up your infrastructure, and common code problems. We do this to make sure your changes are solid and won't cause any trouble later.

Talking to CodeAnt AI

Got a question or need a hand with something in your pull request? You can easily get in touch with CodeAnt AI right here. Just type the following in a comment on your pull request, and replace "Your question here" with whatever you want to ask:

@codeant-ai ask: Your question here

This lets you have a chat with CodeAnt AI about your pull request, making it easier to understand and improve your code.

Example

@codeant-ai ask: Can you suggest a safer alternative to storing this secret?

Preserve Org Learnings with CodeAnt

You can record team preferences so CodeAnt AI applies them in future reviews. Reply directly to the specific CodeAnt AI suggestion (in the same thread) and replace "Your feedback here" with your input:

@codeant-ai: Your feedback here

This helps CodeAnt AI learn and adapt to your team's coding style and standards.

Example

@codeant-ai: Do not flag unused imports.

Retrigger review

Ask CodeAnt AI to review the PR again, by typing:

@codeant-ai: review

Check Your Repository Health

To analyze the health of your code repository, visit our dashboard at https://app.codeant.ai. This tool helps you identify potential issues and areas for improvement in your codebase, ensuring your repository maintains high standards of code health.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@qodo-code-review

Copy link
Copy Markdown

ⓘ Qodo reviews are paused because your trial has ended. Ask your workspace admin to add credits to resume reviews. Manage billing

@codeant-ai

codeant-ai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

🤖 CodeAnt AI — Review Status

Status Commit Started (UTC) Finished (UTC)
✅ Incremental review completed 862a92e Sep 17, 2026 · 19:44 19:45
✅ Incremental review completed e2f0eef Sep 16, 2026 · 12:25 12:25
✅ Incremental review completed 1348211 Sep 13, 2026 · 12:09 12:09
✅ Incremental review completed 05d3264 Sep 13, 2026 · 08:29 08:29
✅ Reviewed your PR 7194d63 Sep 13, 2026 · 04:23 04:25

@codeant-ai

codeant-ai Bot commented Sep 13, 2026

Copy link
Copy Markdown

Thanks for using CodeAnt! 🎉

We're free for open-source projects. if you're enjoying it, help us grow by sharing.

Share on X ·
Reddit ·
LinkedIn

@coderabbitai

coderabbitai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Warning

Review limit reached

Next included review available in 19 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Advanced

Run ID: af70d404-acba-4fc8-bb0d-cc2b191ee462

📥 Commits

Reviewing files that changed from the base of the PR and between 05d3264 and 35810b9.

📒 Files selected for processing (3)
  • .github/workflows/lint.yml
  • scripts/fleet_monitor.sh
  • scripts/pr_review_health.sh
📝 Walkthrough

Walkthrough

The change adds pure run-attribution helpers for legacy workflows and collapsed agent-ingress.yml jobs. Fleet monitoring and PR-review health scans use job-level attribution, render role reports, and retain legacy compatibility. Bats coverage validates normalization, parity, bucketing, and reporting.

Changes

Run attribution

Layer / File(s) Summary
Attribution helpers and validation
scripts/lib/run-attribution.sh, tests/run_attribution.bats, .github/workflows/lint.yml
Adds role extraction, legacy and ingress run normalization, outcome precedence, bucket aggregation, report rendering, parity checks, and CI execution for the new Bats suite.
Fleet monitor ingress sampling
scripts/fleet_monitor.sh
Adds capped agent-ingress run and jobs API sampling, role attribution output, report wiring, and temporary-file cleanup.
PR-review health attribution
scripts/pr_review_health.sh
Adds configurable pr-review role selection, agent-ingress job sampling, normalized report output, warnings for API failures, and temporary-file cleanup.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Feature · Severity of issue fixed: Medium

Sequence Diagram(s)

sequenceDiagram
  participant FleetMonitor
  participant GitHubActionsAPI
  participant RunAttribution
  participant Report
  FleetMonitor->>GitHubActionsAPI: Find agent-ingress workflow and completed runs
  FleetMonitor->>GitHubActionsAPI: Fetch jobs for each run
  FleetMonitor->>RunAttribution: Normalize jobs into role outcomes
  RunAttribution->>Report: Aggregate buckets and render attribution table
  Report-->>FleetMonitor: Add attribution section to monitor output
Loading

Merge Risk: 🟠 High · up to 05d32

Fleet and PR-review reports can present incomplete or misleading role health, so the attribution behavior should be corrected before merge.

🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description includes a useful summary, change list, and impact statement. It omits the required Interaction contract section and Checklist, including the N/A selection and shellcheck and documenta… Add the required Interaction contract section. Select N/A if this PR does not add or change an agentic role. Add the Checklist section and record the shellcheck and documentation statuses. Remove sections that do not apply.
Linked Issues check ⚠️ Warning The pull request implements the main [#1727] attribution objectives. scripts/fleet_monitor.sh and scripts/pr_review_health.sh query agent-ingress jobs and use role-bearing job names. `scripts/lib/… Remove INGRESS_ATTR_MAX_RUNS and the associated early-stop and warning logic from scripts/fleet_monitor.sh. Sample all completed ingress runs in the lookback window, subject only to the existing API pagination and stop conditions. Add o…
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the issue, phase, and primary change: retargeting run attribution to job and role level for the fleet monitor. It is specific and related to the main changeset.
Out of Scope Changes check ✅ Passed The changed files support [#1727]. They implement attribution, fleet and PR-review reporting, pure attribution tests, or CI execution of those tests. No unrelated change is demonstrated. The numeric c…
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 11 functions across 4 files. (1 skipped: 1…
Full details: Description check

Explanation

The description includes a useful summary, change list, and impact statement. It omits the required Interaction contract section and Checklist, including the N/A selection and shellcheck and documentation checks.

Full details: Linked Issues check

Explanation

The pull request implements the main [#1727] attribution objectives. scripts/fleet_monitor.sh and scripts/pr_review_health.sh query agent-ingress jobs and use role-bearing job names. scripts/lib/run-attribution.sh preserves legacy workflow roles, deduplicates roles per run, applies outcome precedence, combines pre- and post-collapse records, and emits __unattributed__. tests/run_attribution.bats covers parity and missing-attribution cases, and .github/workflows/lint.yml runs the suite. The pull request also adds INGRESS_ATTR_MAX_RUNS, defaults it to 200, and stops sampling after the cap. [#1727] explicitly prohibits a new numeric cap for this observability change.

Resolution

Remove INGRESS_ATTR_MAX_RUNS and the associated early-stop and warning logic from scripts/fleet_monitor.sh. Sample all completed ingress runs in the lookback window, subject only to the existing API pagination and stop conditions. Add or update a test that proves runs beyond the former cap remain attributable.

✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch dev-lead/issue-1727-20260913-0408

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codeant-ai codeant-ai Bot added the size:XL This PR changes 500-999 lines, ignoring generated files label Sep 13, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new run-to-role attribution mechanism for the collapsed agent-ingress workflow by adding scripts/lib/run-attribution.sh and integrating it into scripts/fleet_monitor.sh and scripts/pr_review_health.sh, along with a BATS test suite. The review feedback suggests avoiding potential SIGPIPE failures in shell scripts by replacing pipes to head -1 with jq's built-in first function, and recommends asserting specific non-zero exit codes in tests instead of using generic [ "$status" -ne 0 ] checks.

Comment thread scripts/fleet_monitor.sh Outdated
Comment thread scripts/pr_review_health.sh Outdated
Comment thread tests/run_attribution.bats Outdated
Comment thread tests/run_attribution.bats Outdated
Comment thread tests/run_attribution.bats Outdated
Comment thread tests/run_attribution.bats Outdated
Comment thread scripts/fleet_monitor.sh
Comment thread scripts/lib/run-attribution.sh Outdated
Comment thread scripts/pr_review_health.sh Outdated
Comment thread scripts/pr_review_health.sh Outdated
@don-petry

Copy link
Copy Markdown
Collaborator Author

Auto-rebase failed — merge conflict — this branch has conflicts with main that must be resolved.

dev-lead will attempt to resolve this automatically. If it cannot, a follow-up comment will explain what needs manual attention.

To resolve manually instead:

git fetch origin
git merge origin/main
# resolve conflicts, then:
git add .
git commit
git push

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — review-changes (applied)

Changes committed and pushed. Requested items addressed:

  • scripts/fleet_monitor.sh:528 — applied
  • scripts/pr_review_health.sh:151 — applied
  • tests/run_attribution.bats:34 — applied
  • tests/run_attribution.bats:39 — applied
  • tests/run_attribution.bats:176 — applied
  • tests/run_attribution.bats:229 — applied
  • scripts/fleet_monitor.sh:551 — applied
  • scripts/lib/run-attribution.sh:97 — applied
  • scripts/pr_review_health.sh:151 — applied
  • scripts/pr_review_health.sh:156 — applied

@don-petry
don-petry enabled auto-merge (squash) September 13, 2026 04:33
@donpetry-bot

Copy link
Copy Markdown
Contributor

CI checks on this PR are still running. The PR-review sweep re-reviews this PR automatically once the checks complete — no action is needed.

Posted by the donpetry-bot PR-review cascade.

@don-petry

Copy link
Copy Markdown
Collaborator Author

Auto-rebase failed — merge conflict — this branch has conflicts with main that must be resolved.

dev-lead will attempt to resolve this automatically. If it cannot, a follow-up comment will explain what needs manual attention.

To resolve manually instead:

git fetch origin
git merge origin/main
# resolve conflicts, then:
git add .
git commit
git push

@donpetry-bot

Copy link
Copy Markdown
Contributor

pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's head-age-timeout fallback proceeded. Recorded for the miss-rate metric (#1596).

@don-petry
don-petry disabled auto-merge September 13, 2026 05:38
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — rate-limited (intent: rebase)

PR: #1789
The retry cron will re-attempt automatically.

@don-petry
don-petry enabled auto-merge (squash) September 13, 2026 05:46
@don-petry

Copy link
Copy Markdown
Collaborator Author

No description provided.

@don-petry
don-petry disabled auto-merge September 13, 2026 08:26
@don-petry
don-petry force-pushed the dev-lead/issue-1727-20260913-0408 branch from 473d257 to 05d3264 Compare September 13, 2026 08:29
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — rebase (applied)

Rebase completed and pushed.

@don-petry
don-petry enabled auto-merge (squash) September 13, 2026 08:29
@don-petry
don-petry disabled auto-merge September 13, 2026 08:31
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — fix-bot-comment (no-changes)

Agent reasoning
- 0 Security Hotspots ✓
- 0.0% Coverage on New Code
- 0.0% Duplication on New Code
Per the task instructions: "A neutral overview is not an actionable finding. If the bot comment merely *describes* or summarizes the diff without reporting a specific, actionable defect tied to a file/line, there is nothing to fix." The SonarCloud comment is a **passing status summary**, not a defect report.
**Tier 1 blocker (template-drift):** The `template-drift` check failure is about `petry-projects/repo-template` being out of sync with standards. This is:
- A **known pre-existing issue (#1448)** documented in lint.yml as "known-red"
- **Unrelated to this PR** — the PR doesn't modify any of the drifted files or standards/
- **Outside PR scope** — fixing it would require reseeding repo-template separately
## Conclusion
**No changes needed.** The SonarCloud bot has no findings to address (quality gate passed with 0 issues), and the template-drift failure is a pre-existing environmental issue outside the scope of this PR implementation (issue #1727: run attribution).

@donpetry-bot

Copy link
Copy Markdown
Contributor

pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's head-age-timeout fallback proceeded. Recorded for the miss-rate metric (#1596).

@donpetry-bot

Copy link
Copy Markdown
Contributor

pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's quiescence-timeout fallback proceeded. Recorded for the miss-rate metric (#1596).

@donpetry-bot

Copy link
Copy Markdown
Contributor

pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's head-age-timeout fallback proceeded. Recorded for the miss-rate metric (#1596).

@don-petry don-petry removed the needs-human-review Flagged by automated PR review agent label Sep 20, 2026
@donpetry-bot

Copy link
Copy Markdown
Contributor

Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-09-20T21:29:44Z.

@donpetry-bot donpetry-bot added the needs-human-review Flagged by automated PR review agent label Sep 20, 2026
@donpetry-bot

Copy link
Copy Markdown
Contributor

pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's head-age-timeout fallback proceeded. Recorded for the miss-rate metric (#1596).

@donpetry-bot

Copy link
Copy Markdown
Contributor

Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-09-21T01:04:32Z.

@donpetry-bot

Copy link
Copy Markdown
Contributor

pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's quiescence-timeout fallback proceeded. Recorded for the miss-rate metric (#1596).

@donpetry-bot

Copy link
Copy Markdown
Contributor

CI checks on this PR are still running. The PR-review sweep re-reviews this PR automatically once the checks complete — no action is needed.

Posted by the donpetry-bot PR-review cascade.

@donpetry-bot

Copy link
Copy Markdown
Contributor

Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-09-21T01:27:41Z.

@sonarqubecloud

Copy link
Copy Markdown

@donpetry-bot

Copy link
Copy Markdown
Contributor

Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-09-21T01:43:10Z.

@donpetry-bot

Copy link
Copy Markdown
Contributor

pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's head-age-timeout fallback proceeded. Recorded for the miss-rate metric (#1596).

@don-petry don-petry removed the needs-human-review Flagged by automated PR review agent label Sep 21, 2026
@don-petry
don-petry merged commit a1ccf3b into main Sep 21, 2026
66 of 67 checks passed
@don-petry
don-petry deleted the dev-lead/issue-1727-20260913-0408 branch September 21, 2026 02:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Phase 2] Retarget run attribution to job/role level so the collapse doesn't blind the fleet monitor

2 participants