feat(compliance): add retrigger for stale issues + dev-lead workflow health enforcement - #326
Conversation
|
Warning Rate limit exceeded
You’ve run out of usage credits. Purchase more in the billing tab. ⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Pull request overview
This PR introduces an automated remediation “nudge” for compliance-audit findings by (1) enforcing that dev-lead.yml is enabled across fleet repos and (2) re-triggering stale open compliance issues by cycling the claude label so issues:labeled fires again.
Changes:
- Add
scripts/compliance-retrigger.shto scan org repos, enable disableddev-lead.ymlworkflows, and re-trigger stale compliance issues. - Add a scheduled + manually-dispatchable workflow to run the retrigger script daily.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 6 comments.
| File | Description |
|---|---|
scripts/compliance-retrigger.sh |
New fleet-wide script to re-enable disabled dev-lead workflows and cycle labels on stale compliance issues. |
.github/workflows/compliance-retrigger.yml |
New scheduled/manual workflow to run the retrigger script with configurable inputs. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| repos=$(gh api "orgs/$ORG/repos?per_page=100" \ | ||
| --jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \ |
| has_open_pr() { | ||
| local repo="$1" issue="$2" | ||
| local count | ||
| count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \ |
| # Search across the whole org | ||
| local issues | ||
| issues=$(gh api \ | ||
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ |
| stale_cutoff() { | ||
| date -u -d "${STALE_DAYS} days ago" '+%Y-%m-%dT%H:%M:%SZ' 2>/dev/null \ | ||
| || python3 -c "from datetime import datetime,timedelta,timezone; \ | ||
| print((datetime.now(timezone.utc)-timedelta(days=${STALE_DAYS})).strftime('%Y-%m-%dT%H:%M:%SZ'))" | ||
| } |
| disabled_manually|disabled_inactivity) | ||
| warn "dev-lead workflow is '$state' in $repo — enabling it" | ||
| if [ "$DRY_RUN" != "true" ]; then | ||
| local wf_id | ||
| wf_id=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \ | ||
| --jq '.id' 2>/dev/null || echo "") | ||
| if [ -n "$wf_id" ] && [ "$wf_id" != "null" ]; then | ||
| gh api -X PUT "repos/$ORG/$repo/actions/workflows/$wf_id/enable" 2>/dev/null \ | ||
| && info "Enabled dev-lead in $repo" \ | ||
| || warn "Failed to enable dev-lead in $repo" | ||
| fi | ||
| fi | ||
| WORKFLOWS_DISABLED=$((WORKFLOWS_DISABLED + 1)) | ||
| ;; |
| # is re-applied. This script recovers those lost events automatically. | ||
| # | ||
| # Environment: | ||
| # GH_TOKEN — must have issues:write and contents:read across the org |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8d03996235
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \ | ||
| --jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \ |
There was a problem hiding this comment.
Paginate PR lookup before deciding to retrigger
has_open_pr queries repos/$ORG/$repo/pulls?state=open once, but this endpoint is paginated and gh api only fetches one page unless --paginate is used (per gh api manual). In repositories with more than the first page of open PRs, an existing dev-lead/issue-* PR can be missed, so the script will incorrectly cycle labels and retrigger work that already has an open fix PR.
Useful? React with 👍 / 👎.
| repos=$(gh api "orgs/$ORG/repos?per_page=100" \ | ||
| --jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \ | ||
| 2>/dev/null || echo "") |
There was a problem hiding this comment.
Paginate org repo listing in workflow health check
The fleet scan uses orgs/$ORG/repos?per_page=100 without pagination, so only the first page of repositories is inspected. Once the organization has more than 100 active repos, any disabled dev-lead.yml workflows beyond page 1 will never be re-enabled, which undermines the script’s stated org-wide enforcement behavior.
Useful? React with 👍 / 👎.
| issues=$(gh api \ | ||
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ | ||
| --jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \ |
There was a problem hiding this comment.
Paginate search results for stale compliance issues
The stale-issue query requests per_page=100 but does not paginate, so the retrigger pass only processes the first page of matching issues. If more than 100 open compliance-audit issues match the query, older stale issues on later pages are silently skipped and never retriggered.
Useful? React with 👍 / 👎.
| local repo="$1" issue="$2" | ||
| local count | ||
| count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \ | ||
| --jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \ |
There was a problem hiding this comment.
Match exact issue branch prefix when detecting open PRs
The PR detector uses startswith("dev-lead/issue-${issue}"), which also matches other issue numbers that share the same prefix (for example, issue 12 matches dev-lead/issue-123...). That can cause false positives and skip retriggering the actual stale issue because an unrelated PR is mistaken as its fix.
Useful? React with 👍 / 👎.
| # Search across the whole org | ||
| local issues | ||
| issues=$(gh api \ | ||
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ |
There was a problem hiding this comment.
Search all audit issues instead of requiring trigger label
The query requires both label:compliance-audit and label:claude, but the script’s goal is to recover stale compliance issues regardless of current trigger-label state. Any stale compliance issue that lost the claude label (manual edits, prior automation drift) is excluded from processing and will never be retriggered, even though cycle_label already handles absent labels safely.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Code Review
This pull request introduces a script to automate the re-triggering of stale compliance-audit issues and ensure the health of the dev-lead workflow across an organization's repositories. The review feedback focuses on improving the script's scalability and efficiency by suggesting the use of pagination for GitHub API calls when fetching repositories, issues, and pull requests. Additionally, it recommends consolidating redundant API requests when checking workflow states to optimize performance.
| repos=$(gh api "orgs/$ORG/repos?per_page=100" \ | ||
| --jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \ | ||
| 2>/dev/null || echo "") |
There was a problem hiding this comment.
The repository list is fetched without pagination, limiting the check to the first 100 repositories in the organization. For organizations with more than 100 repositories, workflow health enforcement will be incomplete. Adding --paginate ensures all repositories are processed.
| repos=$(gh api "orgs/$ORG/repos?per_page=100" \ | |
| --jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \ | |
| 2>/dev/null || echo "") | |
| repos=$(gh api "orgs/$ORG/repos?per_page=100" --paginate \ | |
| --jq '.[] | select(.archived == false and .disabled == false) | .name' \ | |
| 2>/dev/null || echo "") |
| issues=$(gh api \ | ||
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ | ||
| --jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \ | ||
| 2>/dev/null || echo "") |
There was a problem hiding this comment.
The search for stale issues is limited to the first 100 results. If the organization has a large number of open compliance issues, some will be skipped during the re-trigger process. Adding --paginate ensures all matching issues are processed.
| issues=$(gh api \ | |
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ | |
| --jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \ | |
| 2>/dev/null || echo "") | |
| issues=$(gh api \ | |
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ | |
| --paginate \ | |
| --jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \ | |
| 2>/dev/null || echo "") |
| count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \ | ||
| --jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \ | ||
| 2>/dev/null || echo "0") |
There was a problem hiding this comment.
The check for existing pull requests does not account for pagination. If a repository has more than 30 open pull requests (the default page size), a dev-lead PR might exist on a subsequent page and go undetected, leading to redundant re-triggers. Using --paginate and counting the results ensures all pages are checked.
| count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \ | |
| --jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \ | |
| 2>/dev/null || echo "0") | |
| count=$(gh_api "repos/$ORG/$repo/pulls?state=open" --paginate \ | |
| --jq ".[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\")) | .number" \ | |
| 2>/dev/null | wc -l) |
| local state | ||
| state=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \ | ||
| --jq '.state' 2>/dev/null || echo "missing") |
There was a problem hiding this comment.
The script makes redundant API calls to fetch the workflow state and then the workflow ID. These can be combined into a single call to the same endpoint to improve efficiency and reduce API usage.
| local state | |
| state=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \ | |
| --jq '.state' 2>/dev/null || echo "missing") | |
| local state wf_info | |
| wf_info=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \ | |
| --jq '{state: .state, id: .id}' 2>/dev/null || echo '{"state":"missing"}') | |
| state=$(echo "$wf_info" | jq -r '.state') |
| wf_id=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \ | ||
| --jq '.id' 2>/dev/null || echo "") |
There was a problem hiding this comment.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 071e9c2c54
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| issues=$(gh api \ | ||
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ | ||
| --jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \ | ||
| 2>/dev/null || echo "") |
There was a problem hiding this comment.
Fail fast when stale-issue search API call fails
The stale-issue fetch masks all gh api failures with || echo "", and the next branch treats an empty result as “no open compliance-audit issues.” If the token is missing/expired, rate-limited, or the API returns an error, the job will succeed while doing no retrigger work, which silently disables the recovery mechanism this workflow is meant to enforce.
Useful? React with 👍 / 👎.
| && info "Enabled dev-lead in $repo" \ | ||
| || warn "Failed to enable dev-lead in $repo" | ||
| fi | ||
| fi | ||
| WORKFLOWS_DISABLED=$((WORKFLOWS_DISABLED + 1)) |
There was a problem hiding this comment.
Count only successful workflow re-enables
The script increments WORKFLOWS_DISABLED unconditionally for every disabled workflow, but later reports that value as “re-enabled.” This overstates remediation in both dry-run mode and real runs where enable calls fail, so the summary can claim successful fixes even when nothing was actually re-enabled.
Useful? React with 👍 / 👎.
| # Search across the whole org | ||
| local issues | ||
| issues=$(gh api \ | ||
| "search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \ |
There was a problem hiding this comment.
Limit org search results to issues, not pull requests
The search/issues query does not include is:issue, so matching pull requests can be returned and treated as stale compliance issues. In that case the script will cycle labels on PRs and count them as retriggered even though they are not target issues, creating false positives in the retrigger metrics and unnecessary automation churn.
Useful? React with 👍 / 👎.
| count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \ | ||
| --jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \ | ||
| 2>/dev/null || echo "0") |
There was a problem hiding this comment.
Stop treating PR lookup failures as "no open PR"
The open-PR check converts any gh api failure into 0, so authentication/rate-limit/network errors are interpreted as “no matching PR exists.” That causes stale issues to be retriggered even when an open fix PR is already present, creating duplicate label churn and incorrect automation decisions.
Useful? React with 👍 / 👎.
| repos=$(gh api "orgs/$ORG/repos?per_page=100" \ | ||
| --jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \ | ||
| 2>/dev/null || echo "") |
There was a problem hiding this comment.
Fail fast when repo inventory API call fails
The fleet repo listing falls back to an empty string on any API error, and the function then reports the check as complete with zero repos inspected. If org listing fails (token scope, transient API failure), workflow-health enforcement is silently skipped for the entire run.
Useful? React with 👍 / 👎.
| state=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \ | ||
| --jq '.state' 2>/dev/null || echo "missing") | ||
|
|
There was a problem hiding this comment.
Distinguish workflow fetch errors from missing workflows
Per-repo workflow state retrieval maps every API failure to missing, which is then treated as “not a fleet repo” and skipped. A 403/429/transient error on a real fleet repo therefore hides disabled workflows from remediation and makes the enforcement summary unreliable.
Useful? React with 👍 / 👎.
| gh api -X DELETE "repos/$ORG/$repo/issues/$issue/labels/$TRIGGER_LABEL" 2>/dev/null || true | ||
| gh api -X POST "repos/$ORG/$repo/issues/$issue/labels" \ | ||
| --field "labels[]=$TRIGGER_LABEL" >/dev/null |
There was a problem hiding this comment.
Avoid dropping trigger label when re-add fails
Label cycling deletes the trigger label before re-adding it, but does not guard against POST failure. If deletion succeeds and re-add fails (for example transient API or permission errors), the issue is left without the trigger label, and because stale-issue discovery filters on that label, future runs can stop seeing and retriggering that issue.
Useful? React with 👍 / 👎.
|
…health enforcement (#326) * feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues * feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
…health enforcement (#326) * feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues * feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
…health enforcement (#326) * feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues * feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
…health enforcement (#326) * feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues * feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
…health enforcement (#326) * feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues * feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)



Problem
The 2026-05-15 compliance audit created 29 issues across 8 repos. None were fixed by 2026-05-20 (5 days later). Root-cause analysis identified a three-failure-cascade:
Failure 1 — dev-lead template error (May 15)
dev-lead-reusable.ymlhad a JSON template error that caused everyissues:labeledrun to fail. Both label events fired correctly (compliance-audit→ skip,claude→ issue intent) but the issue-intent run failed at the reusable workflow level. Fixed by PR #196 on 2026-05-16.Failure 2 — dev-lead workflows manually disabled (persisted through May 20)
After the May 16 fix, the
claude-label events for the May 15 issues had already fired and would not re-fire. Meanwhile, dev-lead was manually disabled in all fleet repos (broodly, ContentTwin, markets, TalkTerm, bmad-bgreat-suite, google-app-scripts). New label events would silently no-op with no workflow run created.Failure 3 — git identity not configured in
dev-lead-fix-issue.shWhen labels were finally cycled to re-trigger (May 20) and workflows re-enabled, every
issueintent run failed withfatal: empty ident name. Thefix-issue.shscript commits on behalf of Claude (the prompt instructs Claude NOT to commit), but never setsgit config user.name/emailbefore committing. Fix: petry-projects/.github-private#326.Solution (this PR)
1.
scripts/compliance-retrigger.shNew daily script that:
dev-lead.ymlthat was disabled manually or by inactivity.compliance-auditissues older thanSTALE_DAYS(default: 2) that have no opendev-lead/issue-*PR and cycles theclaudelabel to re-fireissues:labeled.2.
.github/workflows/compliance-retrigger.ymlRuns the script daily at 14:00 UTC. Also has a manual trigger with
stale_daysanddry_runinputs for ad-hoc use.Why label cycling (not
workflow_dispatch)?The
dev-lead.ymlin fleet repos has noworkflow_dispatchtrigger, so the only way to re-invoke theissueintent is viaissues:labeled. Label cycling is the correct and supported mechanism.What this does NOT fix
dev-lead-fix-issue.shgit identity bug → fixed separately in.github-privatePR feat(compliance): add retrigger for stale issues + dev-lead workflow health enforcement #326.Test plan
bash scripts/compliance-retrigger.shlocally withDRY_RUN=trueand verify it lists all open compliance issues + disabled workflows without making changes.compliance-retrigger.ymlmanually withdry_run: trueand confirm the step summary shows correct counts..github-privatemerges, trigger withdry_run: falseand verify dev-lead PRs are created for remaining open compliance issues.🤖 Generated with Claude Code