Skip to content

feat(compliance): add retrigger for stale issues + dev-lead workflow health enforcement - #326

Merged
don-petry merged 3 commits into
mainfrom
feat/compliance-retrigger-stale-issues
May 20, 2026
Merged

don-petry merged 3 commits into
mainfrom
feat/compliance-retrigger-stale-issues

Conversation

@don-petry

Copy link
Copy Markdown
Contributor

Problem

The 2026-05-15 compliance audit created 29 issues across 8 repos. None were fixed by 2026-05-20 (5 days later). Root-cause analysis identified a three-failure-cascade:

Failure 1 — dev-lead template error (May 15)

dev-lead-reusable.yml had a JSON template error that caused every issues:labeled run to fail. Both label events fired correctly (compliance-audit → skip, claude → issue intent) but the issue-intent run failed at the reusable workflow level. Fixed by PR #196 on 2026-05-16.

Failure 2 — dev-lead workflows manually disabled (persisted through May 20)

After the May 16 fix, the claude-label events for the May 15 issues had already fired and would not re-fire. Meanwhile, dev-lead was manually disabled in all fleet repos (broodly, ContentTwin, markets, TalkTerm, bmad-bgreat-suite, google-app-scripts). New label events would silently no-op with no workflow run created.

Failure 3 — git identity not configured in dev-lead-fix-issue.sh

When labels were finally cycled to re-trigger (May 20) and workflows re-enabled, every issue intent run failed with fatal: empty ident name. The fix-issue.sh script commits on behalf of Claude (the prompt instructs Claude NOT to commit), but never sets git config user.name/email before committing. Fix: petry-projects/.github-private#326.

Solution (this PR)

1. scripts/compliance-retrigger.sh

New daily script that:

  • Enforces dev-lead workflow state — scans every active repo, auto-re-enables any dev-lead.yml that was disabled manually or by inactivity.
  • Re-triggers stale issues — finds all open compliance-audit issues older than STALE_DAYS (default: 2) that have no open dev-lead/issue-* PR and cycles the claude label to re-fire issues:labeled.

2. .github/workflows/compliance-retrigger.yml

Runs the script daily at 14:00 UTC. Also has a manual trigger with stale_days and dry_run inputs for ad-hoc use.

Why label cycling (not workflow_dispatch)?

The dev-lead.yml in fleet repos has no workflow_dispatch trigger, so the only way to re-invoke the issue intent is via issues:labeled. Label cycling is the correct and supported mechanism.

What this does NOT fix

Test plan

  • Run bash scripts/compliance-retrigger.sh locally with DRY_RUN=true and verify it lists all open compliance issues + disabled workflows without making changes.
  • Trigger compliance-retrigger.yml manually with dry_run: true and confirm the step summary shows correct counts.
  • After PR feat(compliance): add retrigger for stale issues + dev-lead workflow health enforcement #326 in .github-private merges, trigger with dry_run: false and verify dev-lead PRs are created for remaining open compliance issues.

🤖 Generated with Claude Code

@don-petry
don-petry requested a review from a team as a code owner May 20, 2026 12:41
Copilot AI review requested due to automatic review settings May 20, 2026 12:41
@don-petry
don-petry enabled auto-merge (squash) May 20, 2026 12:41
@coderabbitai

coderabbitai Bot commented May 20, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@don-petry has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 37 minutes and 26 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 19dfe7fe-9871-4650-8919-2ce94fe1f0cc

📥 Commits

Reviewing files that changed from the base of the PR and between e292471 and 071e9c2.

📒 Files selected for processing (2)
  • .github/workflows/compliance-retrigger.yml
  • scripts/compliance-retrigger.sh
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/compliance-retrigger-stale-issues

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR introduces an automated remediation “nudge” for compliance-audit findings by (1) enforcing that dev-lead.yml is enabled across fleet repos and (2) re-triggering stale open compliance issues by cycling the claude label so issues:labeled fires again.

Changes:

  • Add scripts/compliance-retrigger.sh to scan org repos, enable disabled dev-lead.yml workflows, and re-trigger stale compliance issues.
  • Add a scheduled + manually-dispatchable workflow to run the retrigger script daily.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 6 comments.

File Description
scripts/compliance-retrigger.sh New fleet-wide script to re-enable disabled dev-lead workflows and cycle labels on stale compliance issues.
.github/workflows/compliance-retrigger.yml New scheduled/manual workflow to run the retrigger script with configurable inputs.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +90 to +91
repos=$(gh api "orgs/$ORG/repos?per_page=100" \
--jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \
has_open_pr() {
local repo="$1" issue="$2"
local count
count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \
Comment on lines +140 to +143
# Search across the whole org
local issues
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \
Comment on lines +53 to +57
stale_cutoff() {
date -u -d "${STALE_DAYS} days ago" '+%Y-%m-%dT%H:%M:%SZ' 2>/dev/null \
|| python3 -c "from datetime import datetime,timedelta,timezone; \
print((datetime.now(timezone.utc)-timedelta(days=${STALE_DAYS})).strftime('%Y-%m-%dT%H:%M:%SZ'))"
}
Comment on lines +104 to +117
disabled_manually|disabled_inactivity)
warn "dev-lead workflow is '$state' in $repo — enabling it"
if [ "$DRY_RUN" != "true" ]; then
local wf_id
wf_id=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \
--jq '.id' 2>/dev/null || echo "")
if [ -n "$wf_id" ] && [ "$wf_id" != "null" ]; then
gh api -X PUT "repos/$ORG/$repo/actions/workflows/$wf_id/enable" 2>/dev/null \
&& info "Enabled dev-lead in $repo" \
|| warn "Failed to enable dev-lead in $repo"
fi
fi
WORKFLOWS_DISABLED=$((WORKFLOWS_DISABLED + 1))
;;
# is re-applied. This script recovers those lost events automatically.
#
# Environment:
# GH_TOKEN — must have issues:write and contents:read across the org

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8d03996235

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +64 to +65
count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \
--jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Paginate PR lookup before deciding to retrigger

has_open_pr queries repos/$ORG/$repo/pulls?state=open once, but this endpoint is paginated and gh api only fetches one page unless --paginate is used (per gh api manual). In repositories with more than the first page of open PRs, an existing dev-lead/issue-* PR can be missed, so the script will incorrectly cycle labels and retrigger work that already has an open fix PR.

Useful? React with 👍 / 👎.

Comment on lines +90 to +92
repos=$(gh api "orgs/$ORG/repos?per_page=100" \
--jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \
2>/dev/null || echo "")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Paginate org repo listing in workflow health check

The fleet scan uses orgs/$ORG/repos?per_page=100 without pagination, so only the first page of repositories is inspected. Once the organization has more than 100 active repos, any disabled dev-lead.yml workflows beyond page 1 will never be re-enabled, which undermines the script’s stated org-wide enforcement behavior.

Useful? React with 👍 / 👎.

Comment on lines +142 to +144
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \
--jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Paginate search results for stale compliance issues

The stale-issue query requests per_page=100 but does not paginate, so the retrigger pass only processes the first page of matching issues. If more than 100 open compliance-audit issues match the query, older stale issues on later pages are silently skipped and never retriggered.

Useful? React with 👍 / 👎.

local repo="$1" issue="$2"
local count
count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \
--jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Match exact issue branch prefix when detecting open PRs

The PR detector uses startswith("dev-lead/issue-${issue}"), which also matches other issue numbers that share the same prefix (for example, issue 12 matches dev-lead/issue-123...). That can cause false positives and skip retriggering the actual stale issue because an unrelated PR is mistaken as its fix.

Useful? React with 👍 / 👎.

# Search across the whole org
local issues
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Search all audit issues instead of requiring trigger label

The query requires both label:compliance-audit and label:claude, but the script’s goal is to recover stale compliance issues regardless of current trigger-label state. Any stale compliance issue that lost the claude label (manual edits, prior automation drift) is excluded from processing and will never be retriggered, even though cycle_label already handles absent labels safely.

Useful? React with 👍 / 👎.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a script to automate the re-triggering of stale compliance-audit issues and ensure the health of the dev-lead workflow across an organization's repositories. The review feedback focuses on improving the script's scalability and efficiency by suggesting the use of pagination for GitHub API calls when fetching repositories, issues, and pull requests. Additionally, it recommends consolidating redundant API requests when checking workflow states to optimize performance.

Comment on lines +90 to +92
repos=$(gh api "orgs/$ORG/repos?per_page=100" \
--jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \
2>/dev/null || echo "")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The repository list is fetched without pagination, limiting the check to the first 100 repositories in the organization. For organizations with more than 100 repositories, workflow health enforcement will be incomplete. Adding --paginate ensures all repositories are processed.

Suggested change
repos=$(gh api "orgs/$ORG/repos?per_page=100" \
--jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \
2>/dev/null || echo "")
repos=$(gh api "orgs/$ORG/repos?per_page=100" --paginate \
--jq '.[] | select(.archived == false and .disabled == false) | .name' \
2>/dev/null || echo "")

Comment on lines +142 to +145
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \
--jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \
2>/dev/null || echo "")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The search for stale issues is limited to the first 100 results. If the organization has a large number of open compliance issues, some will be skipped during the re-trigger process. Adding --paginate ensures all matching issues are processed.

Suggested change
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \
--jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \
2>/dev/null || echo "")
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \
--paginate \
--jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \
2>/dev/null || echo "")

Comment on lines +64 to +66
count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \
--jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \
2>/dev/null || echo "0")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The check for existing pull requests does not account for pagination. If a repository has more than 30 open pull requests (the default page size), a dev-lead PR might exist on a subsequent page and go undetected, leading to redundant re-triggers. Using --paginate and counting the results ensures all pages are checked.

Suggested change
count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \
--jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \
2>/dev/null || echo "0")
count=$(gh_api "repos/$ORG/$repo/pulls?state=open" --paginate \
--jq ".[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\")) | .number" \
2>/dev/null | wc -l)

Comment on lines +96 to +98
local state
state=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \
--jq '.state' 2>/dev/null || echo "missing")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The script makes redundant API calls to fetch the workflow state and then the workflow ID. These can be combined into a single call to the same endpoint to improve efficiency and reduce API usage.

Suggested change
local state
state=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \
--jq '.state' 2>/dev/null || echo "missing")
local state wf_info
wf_info=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \
--jq '{state: .state, id: .id}' 2>/dev/null || echo '{"state":"missing"}')
state=$(echo "$wf_info" | jq -r '.state')

Comment on lines +108 to +109
wf_id=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \
--jq '.id' 2>/dev/null || echo "")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The workflow ID can be retrieved from the cached response from the previous API call, avoiding an unnecessary network request.

Suggested change
wf_id=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \
--jq '.id' 2>/dev/null || echo "")
wf_id=$(echo "$wf_info" | jq -r '.id')

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 071e9c2c54

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +142 to +145
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \
--jq '.items[] | {number: .number, repo: (.repository_url | split("/") | last), created_at: .created_at, title: .title}' \
2>/dev/null || echo "")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Fail fast when stale-issue search API call fails

The stale-issue fetch masks all gh api failures with || echo "", and the next branch treats an empty result as “no open compliance-audit issues.” If the token is missing/expired, rate-limited, or the API returns an error, the job will succeed while doing no retrigger work, which silently disables the recovery mechanism this workflow is meant to enforce.

Useful? React with 👍 / 👎.

Comment on lines +112 to +116
&& info "Enabled dev-lead in $repo" \
|| warn "Failed to enable dev-lead in $repo"
fi
fi
WORKFLOWS_DISABLED=$((WORKFLOWS_DISABLED + 1))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Count only successful workflow re-enables

The script increments WORKFLOWS_DISABLED unconditionally for every disabled workflow, but later reports that value as “re-enabled.” This overstates remediation in both dry-run mode and real runs where enable calls fail, so the summary can claim successful fixes even when nothing was actually re-enabled.

Useful? React with 👍 / 👎.

# Search across the whole org
local issues
issues=$(gh api \
"search/issues?q=org:${ORG}+label:${AUDIT_LABEL}+label:${TRIGGER_LABEL}+state:open&per_page=100" \

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Limit org search results to issues, not pull requests

The search/issues query does not include is:issue, so matching pull requests can be returned and treated as stale compliance issues. In that case the script will cycle labels on PRs and count them as retriggered even though they are not target issues, creating false positives in the retrigger metrics and unnecessary automation churn.

Useful? React with 👍 / 👎.

Comment on lines +64 to +66
count=$(gh_api "repos/$ORG/$repo/pulls?state=open" \
--jq "[.[] | select(.head.ref | startswith(\"dev-lead/issue-${issue}\"))] | length" \
2>/dev/null || echo "0")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Stop treating PR lookup failures as "no open PR"

The open-PR check converts any gh api failure into 0, so authentication/rate-limit/network errors are interpreted as “no matching PR exists.” That causes stale issues to be retriggered even when an open fix PR is already present, creating duplicate label churn and incorrect automation decisions.

Useful? React with 👍 / 👎.

Comment on lines +90 to +92
repos=$(gh api "orgs/$ORG/repos?per_page=100" \
--jq '[.[] | select(.archived == false and .disabled == false) | .name][]' \
2>/dev/null || echo "")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Fail fast when repo inventory API call fails

The fleet repo listing falls back to an empty string on any API error, and the function then reports the check as complete with zero repos inspected. If org listing fails (token scope, transient API failure), workflow-health enforcement is silently skipped for the entire run.

Useful? React with 👍 / 👎.

Comment on lines +97 to +99
state=$(gh api "repos/$ORG/$repo/actions/workflows/dev-lead.yml" \
--jq '.state' 2>/dev/null || echo "missing")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Distinguish workflow fetch errors from missing workflows

Per-repo workflow state retrieval maps every API failure to missing, which is then treated as “not a fleet repo” and skipped. A 403/429/transient error on a real fleet repo therefore hides disabled workflows from remediation and makes the enforcement summary unreliable.

Useful? React with 👍 / 👎.

Comment on lines +78 to +80
gh api -X DELETE "repos/$ORG/$repo/issues/$issue/labels/$TRIGGER_LABEL" 2>/dev/null || true
gh api -X POST "repos/$ORG/$repo/issues/$issue/labels" \
--field "labels[]=$TRIGGER_LABEL" >/dev/null

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Avoid dropping trigger label when re-add fails

Label cycling deletes the trigger label before re-adding it, but does not guard against POST failure. If deletion succeeds and re-add fails (for example transient API or permission errors), the issue is left without the trigger label, and because stale-issue discovery filters on that label, future runs can stop seeing and retriggering that issue.

Useful? React with 👍 / 👎.

@sonarqubecloud

Copy link
Copy Markdown

@don-petry
don-petry merged commit cb9fac2 into main May 20, 2026
28 of 29 checks passed
@don-petry
don-petry deleted the feat/compliance-retrigger-stale-issues branch May 20, 2026 23:55
don-petry added a commit that referenced this pull request Jun 8, 2026
…health enforcement (#326)

* feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues

* feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
don-petry added a commit that referenced this pull request Jun 10, 2026
…health enforcement (#326)

* feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues

* feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
don-petry added a commit that referenced this pull request Jun 11, 2026
…health enforcement (#326)

* feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues

* feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
don-petry added a commit that referenced this pull request Jun 11, 2026
…health enforcement (#326)

* feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues

* feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
don-petry added a commit that referenced this pull request Jun 11, 2026
…health enforcement (#326)

* feat(compliance): add compliance-retrigger.sh to re-dispatch stale issues

* feat(compliance): add compliance-retrigger.yml workflow (daily at 14:00 UTC)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants