Skip to content

Fix flaky detector: unblock dnceng feed, fail fast, and defer quarantine to a second run - #13936

Merged
ViktorHofer merged 5 commits into
mainfrom
flaky-detector-nuget-firewall
Jun 3, 2026
Merged

Fix flaky detector: unblock dnceng feed, fail fast, and defer quarantine to a second run#13936
ViktorHofer merged 5 commits into
mainfrom
flaky-detector-nuget-firewall

Conversation

@ViktorHofer

@ViktorHofer ViktorHofer commented Jun 3, 2026

Copy link
Copy Markdown
Member

Problem

Workflow run 26877237491 failed (agent job, exit 1). Investigation found three compounding causes — one network/infra and two behavioral design flaws.

1. Network restriction — legacy AzDO feed host was firewall-blocked

The AWF firewall denied all 432 requests to dnceng.pkgs.visualstudio.com:443 (the only blocked domain in the run):

[WARN] Firewall blocked domains:
[WARN]   - Blocked: dnceng.pkgs.visualstudio.com:443
| dnceng.pkgs.visualstudio.com | 0 allowed | 432 denied |

NuGet.config splits feeds across two AzDO hosts. pkgs.dev.azure.com is covered by the dotnet ecosystem allowlist, but the dotnet8/dotnet9/dotnet10 feeds live on the legacy dnceng.pkgs.visualstudio.com host, which was not allowlisted. With the runtime-pack feed blocked, NuGet restore (and therefore the whole-repo build) failed.

2. Token-budget exhaustion: chasing the number of an issue it just created

The deeper root cause was not the build loop. After fast-failing the build, the agent created 4 tracking issues via create_issue, then spent 20+ tool calls trying to discover their issue numbers so it could write [ActiveIssue(.../issues/NNNN)] in the same run. That is architecturally impossible: create_issue is a gh-aw safe output filed by a post-run job — the tool returns only {"result":"success"} with no number during the agent run. The flailing ballooned context into the 25M effective-token hard rail:

attempt 1 failed: ... isMaxEffectiveTokensExceededError=true
done: exitCode=1 totalDuration=22m 5s

It was hard-railed before writing any safe output → no PR, exit 1. This wall is latent even on a green build, so the fast-fail fix below is necessary but not sufficient on its own.

3. Dedup-marker drift

The agent wrote <!-- flaky-test: … --> markers while every dedup search looks for <!-- flaky-test-id: … -->, and it re-filed pre-existing issue #13762 (which predates the marker convention).

Fix

  • network.allowed: add dnceng.pkgs.visualstudio.com so restore can reach the dotnet8/9/10 feeds. Recompiled the lock (the firewall allowlist is frontmatter-derived).
  • Step 6 (fail fast): a single environmental/network build failure goes straight to the Step 7d quarantine fallback instead of retrying or attempting a fix.
  • Step 5 (two-run quarantine model): a flake is quarantine-eligible only if its tracking issue was already OPEN before this run with a readable number. A just-created issue does not count this run and becomes eligible next run. Explicitly forbids polling/guessing the number of an issue created this run — the exact behavior that exhausted the budget.
  • Step 4 (dedup hardening): note that create_issue returns no number during the run; stress the marker must be copied exactly as <!-- flaky-test-id: … -->; add a title-search fallback for legacy issues lacking the marker; note the benign Malformed version: gh stderr warning so it is not misread as a failure.

The workflow body is runtime-imported ({{#runtime-import …}}), so the Step 4/5 prose edits take effect without regenerating the lock; only the network allowlist required a recompile.


Also in this PR: merge the detector skill into the workflow

The separate .github/skills/flaky-test-detector/SKILL.md duplicated guidance the workflow already states in its own steps; its only unique content (the evidence-source model + detector JSON field glosses) is now a compact Background section in flaky-test-detector.agent.md. The detector script moved to .github/workflows/scripts/Get-FlakyTests.ps1 (next to the workflow that calls it), and SKILL.md was deleted. Net −241 lines, one fewer moving part. No other workflow or pipeline referenced the skill/script; body is runtime-imported so no lock recompile is needed.

…failures

Run 26877237491 failed: the AWF firewall blocked all 432 requests to
dnceng.pkgs.visualstudio.com (the legacy AzDO host backing the dotnet8/9/10
NuGet feeds in NuGet.config), so the whole-repo restore/build could not
complete. The 'defaults'/'dotnet' allowlists only cover pkgs.dev.azure.com.

Add dnceng.pkgs.visualstudio.com to network.allowed so restore can reach the
runtime-pack feeds.

The agent also burned the entire 25M effective-token budget (22 min) looping
on the failing build before being hard-railed, so it never reached the
quarantine fallback or opened a PR. Harden Step 6 to fail fast: a single
environmental/network build failure must go straight to the Step 7d quarantine
fallback instead of retrying or attempting a fix.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings June 3, 2026 10:27
@ViktorHofer
ViktorHofer requested a review from a team as a code owner June 3, 2026 10:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates the flaky-test-detector Agentic Workflow configuration to avoid restore/build failures caused by the legacy AzDO feed host being blocked by the firewall, and to prevent token-budget exhaustion by guiding the agent to quarantine rather than loop when the environment is broken.

Changes:

  • Allowlist dnceng.pkgs.visualstudio.com in the workflow’s network.allowed so NuGet restore can reach the dotnet8/9/10 feeds referenced by NuGet.config.
  • Add “fail fast” guidance to the build step so an initial environment/network build failure immediately triggers the Step 7d quarantine fallback instead of repeated rebuild attempts.
  • Recompile the generated .agent.lock.yml so the firewall/allowed-domain metadata matches the updated frontmatter.
Show a summary per file
File Description
.github/workflows/flaky-test-detector.agent.md Adds the legacy feed host to the allowlist and documents a fail-fast path to quarantine when the initial build failure is environmental.
.github/workflows/flaky-test-detector.agent.lock.yml Regenerates the lock so the runtime allowed-domain lists include dnceng.pkgs.visualstudio.com and prompt/config hashes reflect the updated source.

Copilot's findings

  • Files reviewed: 2/2 changed files
  • Comments generated: 0

Analysis of run 26877237491 showed the agent created 4 tracking issues via
create_issue, then exhausted the 25M effective-token budget trying to
discover their issue numbers to write [ActiveIssue(.../issues/NNNN)]
attributes in the same run. That is impossible: create_issue is a safe
output filed by a post-run job and returns no number during the run.

- Step 5: a flake is quarantine-eligible only if its tracking issue was
  already OPEN before this run (real number readable now). A just-created
  issue does not count and must not be quarantined this run; it becomes
  eligible next run. Explicitly forbid polling/guessing for a just-created
  issue's number (the behavior that exhausted the budget).
- Step 4: note create_issue returns no number; stress the dedup marker must
  be copied exactly as <!-- flaky-test-id: ... --> (the agent dropped the
  '-id', which would orphan the issue from future marker searches); add a
  title-search fallback for legacy issues lacking the marker; note the
  benign 'Malformed version:' gh stderr warning so it isn't misread as a
  command failure.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@ViktorHofer ViktorHofer changed the title Unblock dnceng.pkgs.visualstudio.com feed and fail fast on env build failures in flaky detector Fix flaky detector: unblock dnceng feed, fail fast, and defer quarantine to a second run Jun 3, 2026
Mirror the workflow behavioral fixes in the skill doc: explain that
create_issue is a deferred safe output returning no number during the run,
so a new flake is filed one run and quarantined the next (never poll for a
just-created issue's number); and stress the dedup marker must be copied
exactly with the -id segment, with a title-search fallback for legacy
issues.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

github-actions Bot commented Jun 3, 2026

Copy link
Copy Markdown
Contributor

🔍 Skill Validator Results

⚠️ Warnings or advisories found

Scope Checked
Skills 1
Agents 0
Total 1
Severity Count
--- ---:
❌ Errors 0
⚠️ Warnings 1
ℹ️ Advisories 0

Summary

Level Finding
ℹ️ Found 1 skill(s)
ℹ️ [flaky-test-detector] 📊 flaky-test-detector: 3,897 BPE tokens [chars/4: 3,794] (standard ~), 13 sections, 5 code blocks
ℹ️ [flaky-test-detector] ⚠ Skill is 3,897 BPE tokens (chars/4 estimate: 3,794) — approaching "comprehensive" range where gains diminish.
ℹ️ ✅ All checks passed (1 skill(s))
Full validator output ```text Found 1 skill(s) [flaky-test-detector] 📊 flaky-test-detector: 3,897 BPE tokens [chars/4: 3,794] (standard ~), 13 sections, 5 code blocks [flaky-test-detector] ⚠ Skill is 3,897 BPE tokens (chars/4 estimate: 3,794) — approaching "comprehensive" range where gains diminish. ✅ All checks passed (1 skill(s)) ```

ViktorHofer and others added 2 commits June 3, 2026 12:59
Run 26877237491 ran the detector twice with identical inputs: it was
launched backgrounded with heavy -MaxBuilds/-MaxArtifactDownloads, polled
after ~65s while the scan was still running, saw no -JsonOut file yet
(the script writes JSON only on completion), misread that as a failure, and
relaunched -- re-downloading every AzDO artifact.

Instruct the agent to run the detector in the foreground and wait for exit,
and note that a missing -JsonOut file mid-run means 'not finished', not
'failed'. Mirror a short note in the skill doc.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The flaky-test-detector SKILL.md was a separate discoverable skill whose
content the daily workflow already re-stated in its own steps (thresholds,
dedup marker, quarantine/un-quarantine conventions, flake-vs-regression
classification, assembly->project mapping). Its only unique content was the
evidence-source model and a gloss of the detector JSON fields.

Fold that essential background into a new 'Background' section of
flaky-test-detector.agent.md so the workflow is self-contained, move the
detector script to .github/workflows/scripts/Get-FlakyTests.ps1 next to the
workflow that calls it (a scripts/ dir with no SKILL.md would be a malformed
skill), update the two invocation paths and the assembly-mapping reference,
and delete SKILL.md. No other workflow or pipeline referenced the skill or
script; the workflow body is runtime-imported so no lock recompile is needed.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@ViktorHofer
ViktorHofer enabled auto-merge (squash) June 3, 2026 11:38
@ViktorHofer
ViktorHofer merged commit 10799a7 into main Jun 3, 2026
10 of 11 checks passed
@ViktorHofer
ViktorHofer deleted the flaky-detector-nuget-firewall branch June 3, 2026 11:53
This was referenced Aug 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants