Skip to content

Fix CI analysis misclassifying setup failures - #20126

Closed
Ella Hathaway (ellahathaway) wants to merge 1 commit into
mainfrom
ellahathaway-vs-code-e2e-stability
Closed

Ella Hathaway (ellahathaway) wants to merge 1 commit into
mainfrom
ellahathaway-vs-code-e2e-stability

Conversation

@ellahathaway

Copy link
Copy Markdown
Contributor

Description

CI analysis reopened #19639 after an Azure Functions Core Tools download failed with HTTP 504 / curl exit code 22, even though Run extension E2E tests was skipped. The analyzer reused the older flaky-test cause based on the shard name rather than the current failure.

This keeps the fix limited to that misclassification:

  • Preserve readable job logs and sanitized log-fetch errors, including support for GitHub CLI's terminal-escape protection.
  • Include skipped steps alongside failed steps in the analysis summary.
  • Determine the current failure phase before matching prior causes. A test recurrence must match both the actual failed test and its diagnostic; a prerequisite download failure must not reuse a test-failure cause.

The existing classifications, publishing flow, issue handling, and extension E2E tests are unchanged.

Validation

  • AnalyzeCiFailureWorkflowTests: 42 passed, including two small regression theories for log diagnostics and collector wiring.
  • Regenerated the workflow with gh-aw v0.86.2 in strict mode; existing action and engine pins preserved.
  • Checked the saved incident log: all four HTTP 504 errors remain readable, with prerequisites failed and the E2E step skipped.

Fixes #19639

Checklist

  • Is this feature complete?
    • Yes. Ready to ship.
    • No. Follow-up changes expected.
  • Are you including unit tests for the changes and scenario tests if relevant?
    • Yes
    • No
  • Did you add public API?
    • Yes
      • If yes, did you have an API Review for it?
        • Yes
        • No
      • Did you add <remarks /> and <code /> elements on your triple slash comments?
        • Yes
        • No
    • No
  • Does the change make any security assumptions or guarantees?
    • Yes
      • If yes, have you done a threat model and had a security review?
        • Yes
        • No
    • No

Preserve job-log diagnostics, show skipped steps, and require matching failure evidence before reusing a prior cause.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

🚀 Dogfood this PR with:

⚠️ WARNING: Do not do this without first carefully reviewing the code of this PR to satisfy yourself it is safe.

curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 20126

Or

  • Run remotely in PowerShell:
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 20126"

@github-actions

Copy link
Copy Markdown
Contributor

Tests selector

1 / 99 PR test projects · 0 PR jobs, from 4 changed files.

Selected PR test projects (1 / 99)

Infrastructure.Tests

Selected PR jobs (0)

none


How these were chosen — grouped by what changed

📄 .github/workflows/analyze-ci-failure.js (changed)
→ 1 directly: Infrastructure.Tests

📄 .github/workflows/analyze-ci-failure.lock.yml (changed)
→ 1 directly: Infrastructure.Tests

📄 .github/workflows/analyze-ci-failure.md (changed)
→ 1 directly: Infrastructure.Tests

🧪 tests/Infrastructure.Tests/WorkflowScripts/AnalyzeCiFailureWorkflowTests.cs (changed test)
→ 1 directly: Infrastructure.Tests

Job reasons

none


Selection computed for commit ffe6ead.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The implementation is focused, generated workflow stays synchronized, and regression coverage addresses the reported failure mode.

Pull request overview

Updates CI failure analysis to distinguish prerequisite/setup failures from recurring flaky tests.

Changes:

  • Preserves sanitized job-log diagnostics and strips terminal escapes.
  • Reports skipped steps and tightens prior-cause matching.
  • Adds focused workflow and formatter regression tests.
File summaries
File Description
.github/workflows/analyze-ci-failure.md Updates collection and classification logic.
.github/workflows/analyze-ci-failure.lock.yml Regenerates the compiled workflow.
.github/workflows/analyze-ci-failure.js Adds safe job-log normalization.
tests/Infrastructure.Tests/WorkflowScripts/AnalyzeCiFailureWorkflowTests.cs Covers diagnostics and workflow wiring.
Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 0
  • Review effort level: Balanced (auto)

Note

Copilot is running an experiment and ran this review at Balanced.


💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Ankit Jain (radical) added a commit that referenced this pull request Sep 15, 2026
Merge #19455 onto current main while preserving its trusted shell-based
validation and persistence flow.

Keep current TRX stdout/stderr diagnostics, and incorporate #20126's
setup-failure safeguards so failed downloads cannot be mistaken for recurring
test failures. Preserve job-log fetch diagnostics and require classifications
to match the current failed phase and evidence.

Regenerate and validate the gh-aw lock file.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@radical

Copy link
Copy Markdown
Member

[automated] The current diffs for this PR and #19455 now cover the same setup-failure scenario:

  • preserving GitHub CLI log-fetch errors and exit status;
  • retaining skipped-step and curl/HTTP failure evidence;
  • classifying from the current failed phase;
  • preventing an older flaky-test cause from matching a current prerequisite/setup failure;
  • adding focused regression coverage for this behavior.

#19455 also addresses the broader attribution, evidence-validation, persistence, and publication paths around this logic. I suggest landing #19455 first. If it lands in its current form, the remaining behavior in this PR appears to be superseded.

@ellahathaway

Copy link
Copy Markdown
Contributor Author

[automated] The current diffs for this PR and #19455 now cover the same setup-failure scenario:

  • preserving GitHub CLI log-fetch errors and exit status;
  • retaining skipped-step and curl/HTTP failure evidence;
  • classifying from the current failed phase;
  • preventing an older flaky-test cause from matching a current prerequisite/setup failure;
  • adding focused regression coverage for this behavior.

#19455 also addresses the broader attribution, evidence-validation, persistence, and publication paths around this logic. I suggest landing #19455 first. If it lands in its current form, the remaining behavior in this PR appears to be superseded.

Closing in favor of #19455

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-engineering-systems infrastructure helix infra engineering repo stuff

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[CI Failure] Flaky: VS Code extension E2E (Linux, azure-functions) shard fails with generic exit code 1, unrelated to PR changes

3 participants