Skip to content

fix(ci): Automatically rerun failed main rolling builds - #20345

Open
Ankit Jain (radical) wants to merge 16 commits into
microsoft:mainfrom
radical:auto-rerun-current-main-ci
Open

Ankit Jain (radical) wants to merge 16 commits into
microsoft:mainfrom
radical:auto-rerun-current-main-ci

Conversation

@radical

@radical Ankit Jain (radical) commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Description

The existing automatic failed-job rerun system now also runs for CI push builds on main. Previously, the system retried pull request CI, but a transient failure on main left the branch red until another commit triggered a new run.

A failed current-main run follows the same shared retry policy:

  • Same rerun implementation: it reuses the existing failed-job rerun helper rather than introducing a separate main retry system.
  • Three automatic reruns: failed attempts 1–3 can each request a rerun, for a maximum of four total attempts.
  • Same terminal analysis: if attempt 4 fails, the existing Analyze CI Failure workflow collects, validates, classifies, and publishes the failure as before.
  • Additional main freshness checks: immediately before a rerun or analysis, the workflow verifies the exact run and attempt, confirms the failed SHA is still the current main SHA, and rejects work superseded by a newer main run.

If a rerun request does not complete successfully, fallback analysis is pinned to that exact failed attempt and SHA. If GitHub started the rerun despite losing the response, the analyzer detects the advanced attempt and skips the stale fallback.

Rerun decisions and skip reasons remain in workflow logs and job summaries. Durable retry-decision history is tracked separately in #20409; this PR preserves the analyzer's existing failure/cause history.

Validation

The live matrix used a trial tree built from exact PR head 8607eaa355f. The production JavaScript helper and terminal guard were byte-identical; trial-only changes were limited to fork activation gates, deterministic failure/timing controls, and API fault-injection wrappers. After validation, radical/main was restored, the temporary issue was closed, the memory branch was deleted, and the trial worktrees were removed.

Fixes #20340

Checklist

  • Is this feature complete?
    • Yes. Ready to ship.
    • No. Follow-up changes expected.
  • Are you including unit tests for the changes and scenario tests if relevant?
    • Yes
    • No
  • Did you add public API?
    • Yes
      • If yes, did you have an API Review for it?
        • Yes
        • No
      • Did you add <remarks /> and <code /> elements on your triple slash comments?
        • Yes
        • No
    • No
  • Does the change make any security assumptions or guarantees?
    • Yes
      • If yes, have you done a threat model and had a security review?
        • Yes
        • No
    • No

@github-actions

Copy link
Copy Markdown
Contributor

🚀 Dogfood this PR with:

⚠️ WARNING: Do not do this without first carefully reviewing the code of this PR to satisfy yourself it is safe.

curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 20345

Or

  • Run remotely in PowerShell:
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 20345"

@aspire-repo-bot
aspire-repo-bot Bot requested a balanced review from Copilot September 23, 2026 01:56
@github-actions github-actions Bot added the area-engineering-systems infrastructure helix infra engineering repo stuff label Sep 23, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Rerun validation and durable attempt history have unresolved reliability issues.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Adds guarded automatic reruns for failed current-main CI runs, with outcome persistence and expanded validation.

Changes:

  • Adds live-state, attempt-cap, SHA, and supersession checks.
  • Integrates reruns with CI failure analysis and persistence.
  • Expands tests and documentation.
File Description
tests/​Infrastructure.Tests/​WorkflowScripts/​AutoRerunTransientCiFailuresTests.cs Tests rerun safeguards and outcomes.
tests/​Infrastructure.Tests/​WorkflowScripts/​auto-rerun-transient-ci-failures.harness.js Extends the GitHub API test harness.
tests/​Infrastructure.Tests/​WorkflowScripts/​AnalyzeCiFailureWorkflowTests.cs Tests workflow orchestration and persistence.
docs/​ci/​auto-rerun-transient-ci-failures.md Documents current-main retry behavior.
docs/​ci/​analyze-ci-failure.md Documents analysis and rerun integration.
.github/​workflows/​auto-rerun-transient-ci-failures.js Implements rerun validation and requests.
.github/​workflows/​analyze-ci-failure.md Orchestrates reruns and analysis.
.github/​workflows/​analyze-ci-failure.lock.yml Updates the generated workflow.
.github/​workflows/​analyze-ci-failure-persistence.sh Persists rerun decisions and outcomes.

Comment thread .github/workflows/analyze-ci-failure.md Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

Note

This error may be related to your runner configuration. You can now configure runners for Copilot code review separately from Copilot cloud agent by creating a copilot-code-review.yml file with your setup steps. Read the docs for details.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Moderate retry, process-environment, and resource-disposal issues remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 Medium severity · 2 Low severity

Open (4)
Files not reviewed (1)
  • src/Aspire.Dashboard/Resources/KnownUrlsDisplay.Designer.cs: Generated file

Comment thread .github/workflows/analyze-ci-failure-persistence.sh Outdated
Comment thread eng/Versions.props
Comment thread src/Aspire.Dashboard/Components/Controls/TerminalView.razor
Failed main pushes had no automatic retry path, leaving transient
failures red until another push. Reuse the existing failed-job rerun
workflow with live attempt, main SHA, and supersession checks, a cap
of three retries, and persisted decisions and recovery.

Keep main reruns out of failure analysis and make memory-branch
publication resilient to concurrent updates.

Refs microsoft#20340

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Failure analysis could file a main-breakage issue on attempts 1–3
while retries were pending, but excluded the fourth failed attempt.

Analyze only that terminal attempt. Recheck the live run, main SHA,
and newer CI runs before collecting evidence and publishing issues,
so superseded failures do not create or update issues.

Fixes microsoft#20340

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved critical fail-open validation issues and moderate workflow regressions must be addressed before approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 High severity · 2 Medium severity · 2 Low severity

Open (6)

Comment thread .github/workflows/analyze-ci-failure-terminal.sh Outdated
Comment thread .github/workflows/auto-rerun-transient-ci-failures.js Outdated
Synchronize with upstream main while retaining the CI retry changes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Current-main retry decisions were written to the analyzer's memory
branch even though nothing consumes them. A failed write could make
the rerun workflow fail after GitHub accepted a retry, while competing
writes required the analyzer to rebase its analysis records.

Keep the live-run guards and three retry attempts, but report decisions
in the workflow summary and logs. Remove the audit-specific writer,
recovery-only path, conflict handling, and tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Durable rerun-state persistence and concurrent publication retry handling remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 High severity · 3 Medium severity · 2 Low severity

Open (7)

Comment thread .github/workflows/auto-rerun-transient-ci-failures.yml
Document when the terminal helper runs and how it handles stale or
unverifiable main runs. Remove the obsolete audit-storage claim and
clarify that the generic CI failure tracker can still report an early
failed push.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The fail-closed run-list validation accepts missing source runs, and stale analysis is checked only after evidence collection begins.

Review effort: Balanced
Findings: 1 High severity · 3 Medium severity · 2 Low severity

Open (6)
Resolved since last review (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Current-run validation occurs after expensive evidence queries

.github/​workflows/​analyze-ci-failure.md:252

This current-run guard executes only after the main path has already queried the triggering merge, last successful run, and candidate merge history (lines 197–228). That contradicts the stated guarantee that stale attempt-4 analysis is rejected before evidence collection and lets superseded runs perform the expensive attribution calls first. Move this verification to immediately after the run metadata/scope is validated, before the main-specific evidence queries.

A failed main CI attempt could be left without Copilot analysis when
GitHub rejected the request for the next automatic rerun. The analyzer
only admitted the configured final attempt, so a failed request stopped
the sequence before an analyzable attempt existed.

Dispatch fallback analysis when an otherwise eligible rerun request
fails, while retaining current-main, live-attempt, and superseding-run
checks. Keep the retry count in defaultMaxRunAttempt so PR reruns, main
reruns, and final analysis derive from one versioned policy value.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@radical
Ankit Jain (radical) marked this pull request as ready for review September 24, 2026 17:18
@radical Ankit Jain (radical) changed the title fix(ci): Automatically rerun failed current-main jobs fix(ci): Automatically rerun failed main rolling builds Sep 24, 2026
Synchronize with upstream main while preserving the current-main retry
and terminal-analysis guards.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Resolve the main merge using the repository's v0.89.17 compiler so the
generated workflow keeps the existing approved action and container pins.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@radical

Copy link
Copy Markdown
Member Author

[automated] Extended CI confidence requires action. The focused final-head tests are strong, but the PR is now conflicting with main, the current base adds a required agentic compile/lint gate, and every cited live workflow scenario predates two load-bearing fail-closed fixes.

Reviewer actions

  • Required before merge: resolve the main conflicts and run Validate Agentic Workflows on the resolved head.
    Done when: fix(ci): Automatically rerun failed main rolling builds #20345 is mergeable and the dedicated run passes compilation, generated-file drift detection, gh aw lint --shellcheck, and the agentic contract tests on the resolved head.
  • Required before merge: repeat the fork/trial workflow scenarios with the resolved final workflow sources.
    Done when: linked runs use the resolved source, lock, JavaScript helper, and terminal guard—apart from documented fork-only gates—and demonstrate automatic recovery, retry-cap/final analysis, stale-run rejection, and pinned fallback analysis.
  • Recommended: record a zizmor disposition for the resolved executable workflows.
    Done when: the changed rerun workflow and generated analyzer workflow have a recorded scan result with findings fixed or explicitly reviewed and suppressed.
Ready-to-use pr-testing prompt

Use the pr-testing skill to test PR #20345 after its conflicts with current main are resolved. Verify the tested source, generated lock, JavaScript helper, and terminal guard match the resolved PR head except for documented fork-only trigger gates and deterministic failure fixtures. In a fork/trial repository, exercise: (1) a current-main failure that automatically reruns and succeeds, (2) failures through source attempts 1–3 followed by attempt 4 stopping at the cap and running final analysis, (3) main advancing before rerun and before publication so both paths stop stale work, and (4) a rejected or ambiguous rerun request that dispatches fallback analysis pinned to the original run ID, attempt, and SHA. Include a malformed or incomplete workflow-run-list case proving rerun and publication fail closed. Also confirm the current Validate Agentic Workflows job passes compilation, zero generated drift, gh aw lint --shellcheck, and agentic contracts, and record a zizmor disposition.

Change classification

  • Changes the non-PR workflow_run workload that can rerun failed CI jobs on current main, including actions: write, retry caps, stale-run checks, and fallback dispatch.
  • Changes the agentic failure analyzer source and generated workflow, including final-attempt gating, immutable fallback identity, and repeated guards before issue or memory mutations.
  • Adds focused JavaScript/shell fixtures and Infrastructure.Tests coverage; it does not change internal Azure DevOps, outerloop, quarantine, or deployment scenarios.

Confidence at a glance

🔴 Final-tree workflow execution — Load-bearing behavior was changed after the cited live trials
  • Why it applies: The workflows mutate CI runs and can publish analyzer results or issues. The final two commits added fail-closed source-run-list validation and pinned fallback identity.
  • Evidence: The cited fork runs completed between 16:18 and 16:32 UTC. Commit d0e1350900e added trusted-source presence checks at 16:57 UTC, and final head 816dd08307b added immutable fallback run-attempt/SHA handling at 17:06 UTC.
  • Gap or disposition: Evidence state: Missing. The live automatic-recovery, retry-cap, stale-run, and fallback scenarios did not execute the final load-bearing behavior. Source-reading and fixture tests do not prove the agentic workflow's final decisions and mutations.
  • Action: Repeat the representative fork/trial scenarios after resolving the current base conflicts, and link runs whose tested blobs match the resolved head.
🔴 Merge-tree compilation and lint — Current-base validation has not run on a resolved tree
  • Why it applies: The PR changes .github/workflows/analyze-ci-failure.md and its generated lock. Current main introduced the path-triggered Validate Agentic Workflows gate and also changed those same analyzer files.
  • Evidence: The PR was tested at head 816dd08307b against base bcdcd4b05ed; current main is 5b70140de58, 41 commits ahead. GitHub reports fix(ci): Automatically rerun failed main rolling builds #20345 as CONFLICTING/DIRTY. The dedicated validator was added by 9272b28a8a7 after the PR's recorded base, so no final-head run exists for it.
  • Gap or disposition: Evidence state: Missing. The PR body reports local recompilation, actionlint, and ShellCheck success, but that evidence is not tied to the unresolved final merge tree. The current base uses a newer canonical compiler/lock-validation path.
  • Action: Resolve conflicts, regenerate through the current canonical path, and require a green Validate Agentic Workflows run with an empty generated-file diff.
🟡 Workflow security lint — No zizmor result is recorded
  • Why it applies: The changed workload uses workflow_run, actions: write, workflow_dispatch, and issue/publication paths across trusted and queued run state.
  • Evidence: The PR records actionlint and ShellCheck success. The repository's gh aw lint --shellcheck command explicitly runs actionlint with ShellCheck and does not run zizmor.
  • Gap or disposition: Evidence state: Advisory gap. No zizmor result or explicit reviewed disposition is attached to the final executable workflows.
  • Action: Scan the resolved workflow files and record the result, resolving or documenting any findings.
🟢 Focused contract coverage — Final-head fixtures exercise the critical guards
  • Why it applies: Most deterministic retry, stale-state, fallback, and publication behavior can be exercised hermetically.
  • Evidence: Final-head CI run 36031960714 used head 816dd08307b; its Infrastructure.Tests job passed 1,367 / 1,367 tests. The PR also records 405 / 405 focused tests. Assertions cover current-main rerun success, attempt-cap rejection, advanced attempt/SHA rejection, missing or malformed source-run lists, pinned fallback inputs, early stop before data collection, and repeated publication guards.
  • Gap or disposition: Evidence state: Verified for the PR head against its recorded base. Reverting the final guards would fail the missing-source-run, mismatched-identity, stale-main, and no-side-effect assertions.
🟢 Action pin provenance — Added action references resolve to approved immutable releases
  • Why it applies: The new current-main rerun job adds two uses: references.
  • Evidence: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd resolves from v6.0.2; actions/github-script@3a2844b7e9c422d3c10d287c895573f7108da1b3 resolves from v9.0.0. Both are full 40-character SHAs from approved actions/* publishers, and the displayed versions match.
  • Gap or disposition: Evidence state: Verified. These are added invocations, not unexplained downgrades or mismatched repins.
⚪ Internal pipelines and specialized test lanes — Not applicable to this diff
  • Why it applies: These categories were checked because the report targets extended CI coverage.
  • Evidence: The final diff is limited to GitHub workflow/analyzer behavior, focused infrastructure tests, and documentation. It does not change Azure DevOps pipeline inputs, signing, packaging, publishing artifacts, outerloop or quarantined tests/runners, or deployment E2E scenarios and their runtime dependencies.
  • Gap or disposition: Evidence state: Not applicable.

Workflow-run jobs persisted checkout credentials even though they never use
authenticated Git operations, and the intentional workflow_run trigger lacked
a recorded security-lint disposition.

Disable persisted credentials, document the trusted-default-branch and live
identity validation model, and expand fail-closed coverage for malformed or
incomplete main run lists.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Main-scope analyzer runs have no subject PR, but the publish_data schema
required a non-empty pr_numbers string. gh-aw rejected the empty value, so
final-attempt analysis reported incomplete instead of publishing the result.

Make pr_numbers optional only for publish-data while retaining the required PR
identity for agent-requested reruns. Add a contract test that preserves that
distinction.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
A new main commit can land after the initial ref check while the workflow-run
list is loading, before the corresponding CI run appears. The stale source run
could then pass the run-list checks and request another attempt.

Re-read refs/heads/main after run-list validation and immediately before the
rerun POST. Add a sequenced harness test that advances the modeled ref during
validation and verifies that no rerun is requested.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

Note

This error may be related to your runner configuration. You can now configure runners for Copilot code review separately from Copilot cloud agent by creating a copilot-code-review.yml file with your setup steps. Read the docs for details.

A workflow_run event can start before the workflow-runs endpoint includes the
completed source run. The strict source-presence check then treated an eligible
current-main failure as an invalid list and skipped its rerun.

Retry structurally valid source-missing lists up to four times with bounded
delays. Malformed responses and superseding runs still fail closed immediately,
and the final main-ref check remains immediately before the rerun POST.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Run-attempt freshness and workflow-list propagation races can cause stale rerun requests or lost terminal analysis.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)

Comment thread .github/workflows/analyze-ci-failure-terminal.sh Outdated
Comment thread .github/workflows/auto-rerun-transient-ci-failures.js
Assert that persistent source-run absence performs exactly four workflow-run-list
requests, while a later superseding or malformed response stops the retry loop
immediately. These tests preserve the bounded and fail-closed guarantees of the
propagation retry.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The rerun path has a stale-attempt race, and transient run-list propagation can permanently suppress final analysis.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)

@radical

Copy link
Copy Markdown
Member Author

[automated] All requested confidence gaps are resolved on final head 8607eaa355f.

The fork was restored after validation: radical/main is back at its baseline, the temporary issue is closed, the memory branch is deleted, and the trial worktrees are removed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left some comments. Thanks for working on this Ankit.

Comment thread .github/workflows/auto-rerun-transient-ci-failures.yml Outdated
Comment thread .github/workflows/auto-rerun-transient-ci-failures.yml Outdated
Comment thread .github/workflows/analyze-ci-failure.md Outdated
Failed rolling main CI runs could request up to three retries by sharing the
pull request policy, and the rerun path dispatched analyzer fallback work when
the GitHub rerun request failed. The run could also advance while the workflow
waited for run-list propagation, allowing a stale attempt to reach the write.

Give rolling main a separate maximum eligible source attempt of 1, remove
analyzer coupling, and fail the rerun workflow directly when the request is
rejected. Revalidate the completed failed run and current main SHA immediately
before requesting the rerun, minimizing the API's unavoidable check-to-write
race.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The implementation provides one retry and lacks the claimed terminal and fallback analysis paths.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)
Resolved since last review (2)

const ignoredJobs = new Set(['Final Results', 'Tests / Final Test Results']);
const defaultMaxRetryableJobs = 5;
const defaultMaxRunAttempt = 3;
const mainMaxRunAttempt = 1;
Comment on lines +1162 to +1166
return {
...mainRerunState,
outcome: 'failed',
reason: 'request-failed',
};
Comment on lines +514 to +518
github.repository == 'microsoft/aspire' &&
github.event_name == 'workflow_run' &&
github.event.workflow_run.event == 'push' &&
github.event.workflow_run.head_branch == 'main' &&
github.event.workflow_run.conclusion == 'failure'

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-engineering-systems infrastructure helix infra engineering repo stuff

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[automated] Add automatic reruns for current main CI failures

3 participants