Fix main CI failure attribution and classification - #19455
Conversation
There was a problem hiding this comment.
Copilot wasn't able to review any files in this pull request.
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 3 out of 3 changed files in this pull request and generated no new comments.
Suppressed comments (3)
.github/workflows/analyze-ci-failure.md:169
- The main-run context records only each candidate PR's title/URL and then deletes the comparison payload. Because the agent is explicitly barred from querying GitHub, it receives no changed-file or patch data for these merges and cannot verify the semantic conflicts this change is intended to diagnose. Preserve and render the comparison's changed files/patches (or collect equivalent per-PR change data) so attribution can be based on code evidence rather than titles.
'. + [$commit + {pull_request: {
number: $pr.number,
title: $pr.title,
url: $pr.html_url,
merged_at: $pr.merged_at
}}]' ci-failure-data/candidate-merges.json \
.github/workflows/analyze-ci-failure.md:746
- For a main run this row contains
main, but reused cause issues created by the old workflow still have aPRcolumn header. The existing-issue path only appends this row and never migrates that header, so stable infra/flaky issues will misleadingly displaymainas a PR value. Rewrite the legacy header toContextbefore appending a main occurrence.
if [ "$RUN_SCOPE" = "main" ]; then
OCCURRENCE_CONTEXT="main"
else
OCCURRENCE_CONTEXT="#${PR_NUMBER}"
fi
NEW_OCCURRENCE_ROW="| ${OCC_DATE} | [${RUN_ID}](${RUN_URL}) | ${FIRST_JOB} | ${OCCURRENCE_CONTEXT} |"
.github/workflows/analyze-ci-failure.md:1189
- The schema now permits
main-repository-breakagefor jobs, but deterministic failed tests on main still have onlyflaky | code-issueavailable in both the JSON example andfailed_tests[].classificationdetails. That forces a main test regression to be mislabeled as a PR code issue or omitted. Addmain-repository-breakageto the failed-test classification contract and example as well.
- `failed_jobs[].classification`: Per-job classification — one of `"transient-infra"`, `"flaky-test"`, `"code-issue"`, or `"main-repository-breakage"`.
|
🚀 Dogfood this PR with:
curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 19455Or
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 19455" |
This comment has been minimized.
This comment has been minimized.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This comment has been minimized.
This comment has been minimized.
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
fab07dc to
b2e5215
Compare
This comment has been minimized.
This comment has been minimized.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 3 out of 3 changed files in this pull request and generated 4 comments.
Suppressed comments (2)
.github/workflows/analyze-ci-failure.md:783
TRUSTED_PR_NUMBERSis used for the final PR comment, but the cause persistence and issue-occurrence paths below still read.pr.numberfrom the agent-authored analysis (lines 836 and 878). A mismatched agent PR number can therefore still be written to the memory branch and attributed in a cause issue, which leaves the trust-boundary problem this change is intended to close. Derive all occurrence PR metadata from this trusted value instead.
TRUSTED_PR_NUMBERS=$(jq -r '.pr_numbers' "$RUN_CONTEXT_FILE")
.github/workflows/analyze-ci-failure.md:990
- The main-breakage issue still publishes agent-authored
main_contextandtriggering_merge_prvalues without comparing them to the downloaded trusted artifacts. An incorrect analysis can therefore put the wrong failed SHA or triggering PR into the privileged issue side effect. Read these fields directly fromrun-context.json,last-successful-main-run.json, andtriggering-merge-pr.json.
LAST_SUCCESSFUL_SHA=$(jq -r '.main_context.last_successful_main_sha // "unknown"' "$ANALYSIS_FILE")
FAILED_SHA=$(jq -r '.main_context.failed_sha // "unknown"' "$ANALYSIS_FILE")
TRIGGERING_MERGE=$(jq -r 'if .triggering_merge_pr then "#\(.triggering_merge_pr.number) \(.triggering_merge_pr.title)" else "Not found" end' "$ANALYSIS_FILE")
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
There was a problem hiding this comment.
🟡 Changes recommended
Unresolved attribution and validation flaws can blame unrelated PRs and omit flaky failures.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 10/12 changed files
- Comments generated: 3
- Review effort level: Balanced
This comment has been minimized.
This comment has been minimized.
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
This comment has been minimized.
This comment has been minimized.
There was a problem hiding this comment.
🔵 Needs a closer look
A moderate labeling defect remains, and the workflow changes affect security-sensitive CI attribution and side effects.
Review details
- Files reviewed: 10/12 changed files
- Comments generated: 0 new
- Review effort level: Balanced
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
There was a problem hiding this comment.
🟡 Changes recommended
Candidate commit and PR metadata can bypass redaction when persisted to failure history.
Get a fresh assessment by requesting another Copilot review.
Review details
- Files reviewed: 11/13 changed files
- Comments generated: 1
- Review effort level: Balanced
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
There was a problem hiding this comment.
🟡 Changes recommended
Rerun logging can bypass credential redaction, and PR mutations lack a final state check.
Get a fresh assessment by requesting another Copilot review.
Review details
Suppressed comments (1)
.github/workflows/analyze-ci-failure.lock.yml:2986
- After this PR-state check, another API request is made to inspect the workflow run before the rerun mutation. The PR can close or become locked during that interval, yet the rerun still proceeds. Move or repeat the PR check after the run-attempt check so the open/unlocked gate is the final check before
reRunWorkflowFailedJobs.
const trustedPrNumber = Number(trustedPrNumberText);
try {
const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: trustedPrNumber });
if (pr.state !== 'open') {
core.info('The subject PR is closed. Skipping rerun.');
return;
}
if (pr.locked) {
core.info('The subject PR is locked. Skipping rerun.');
return;
}
} catch (e) {
core.warning(`Failed to check PR #${trustedPrNumber}: ${e.message}`);
return;
- Files reviewed: 11/13 changed files
- Comments generated: 2
- Review effort level: Balanced
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
Failed `push` runs on `main` could be blamed on whichever pull request
happened to trigger the build, even when the break came from an earlier
merge or from the combined repository state. The motivating case was a
semantic merge conflict on `main`:
error CS0117: 'AzureBicepResourceScope' does not contain a definition for 'ForSubscription'
PR #19090 triggered that build, but the failure came from the interaction
between the earlier merges of #19084 and #18976.
The workflow inferred attribution from associated pull requests rather
than the failed run's immutable metadata, and let non-linear history,
ambiguous associations, partial test evidence, and incomplete cause
coverage reach downstream actions.
Run identity and candidate attribution now derive from trusted GitHub
metadata. Pull-request runs require one unambiguous subject PR, and
failed `main` runs consider the complete merge range since the last
successful run, authorizing merge attribution only for a coherent
`ahead` comparison.
Agent output stays a proposal. Deterministic validation runs before any
side effect and requires exact run and failed-job identity, complete
structured test evidence, causes covering every flaky identity, and a
valid rerun request. Issue titles, diagnostics, and labels render from
trusted metadata, including safe migration of existing matching issues.
Analysis logic moves out of the inline `analyze-ci-failure.js` script
into reviewable shell helpers so each stage is independently testable.
Connection-string and primary/secondary key fields are redacted before
diagnostics are rendered or persisted, and PR mutations recheck that the
pull request is open and unlocked immediately beforehand.
Fixes #19454
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
PR comments and automatic reruns currently look for analysis files in artifacts that do not contain them.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
This comment has been minimized.
This comment has been minimized.
The "Comment on PR" step and the rerun job derived the analysis file
location from the agent output path:
ANALYSIS_FILE="$(dirname "$OUTPUT_FILE")/agent/analysis-result.json"
That path never exists on the runner. `analysis-result.json` and
`causes/*.json` are uploaded as the `ci-analysis-output` artifact, whose
root is `/tmp/gh-aw/agent`, so they unpack to the download directory
root. The `agent` artifact roots at `/tmp/gh-aw` and contains only the
agent output and logs; its `sandbox/agent/logs/` entry does not create a
sibling `agent/` directory.
Both consumers therefore read a missing file. Under `set -euo pipefail`
the comment step failed on `jq`, taking the whole publish job down after
it had already pushed to the memory branch and created cause issues, and
no analysis comment was ever posted. The rerun job hit `core.setFailed`
every time, so automatic reruns never ran.
Both now take the directory from the `download-analysis` step, and the
rerun job downloads `ci-analysis-output` itself because it is a separate
job. The `${ANALYSIS_DIR:-...}` fallbacks are gone: the fallback path is
never valid in these jobs, so a missing or renamed download step now
fails loudly instead of resolving to a bogus path.
The existing fixtures wrote the analysis file next to the fake agent
output, which matched the derived path and made the bug invisible; one
test asserted the incorrect expression verbatim. Fixtures now use the
real artifact layout and pass `ANALYSIS_DIR` explicitly, and the
compiled-workflow contract test covers all four consumers. Reverting
the workflow change alone fails 26 tests.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Tests selector1 / 99 PR test projects · 0 PR jobs, from 12 changed files. Selected PR test projects (1 / 99)
Selected PR jobs (0)none How these were chosen — grouped by what changed📄 📄 📄 📄 📄 📄 📄 📄 📄 🧪 🧪 🧪 Job reasonsnone Selection computed for commit |


What broke
Failed
pushruns onmaincould be blamed on the pull request that happened to trigger the build, even when the failure came from an earlier merge or the combined repository state.The motivating failure was a semantic merge conflict on
main:PR #19090 triggered the build, but the break came from the interaction between the earlier merges of #19084 and #18976.
Root cause
The workflow inferred attribution from associated pull requests instead of the failed run's immutable metadata. It also allowed non-linear or incomplete history, ambiguous associations, partial test evidence, and incomplete cause coverage to reach downstream actions.
How this fixes it
The analyzer now derives run identity and candidate attribution from trusted GitHub metadata. Pull-request runs require one unambiguous subject PR; failed
mainruns consider the complete merge range since the last successful run and authorize merge attribution only for a coherentaheadcomparison.Agent output remains a proposal. Before any side effect, deterministic validation requires exact run and failed-job identity, complete structured test evidence with exact unique
{test, job}pairs, compatible causes covering every flaky identity, and a valid rerun request. Main-breakage issue titles, diagnostics, and labels are rendered from trusted metadata, including safe migration of existing matching issues.Important callouts
tests.yml,run-tests.yml, andextension-e2e-tests.ymlnaming contracts for TRX and Mocha results. Required TRX evidence remains fail-closed; optional extension diagnostics can be absent, oversized, unavailable, or contain no Mocha result when setup or the browser harness fails, in which case the job remains classifiable from trusted steps and sanitized logs.Testing
AnalyzeCiFailureWorkflowTestscontains 277 passing tests covering attribution, evidence completeness, exact test/job binding, TRX and Mocha parsing, cause identity, issue migration and labeling, publication, persistence, redaction, and rerun gates.Checklist
Fixes #19454