Fix flaky detector: unblock dnceng feed, fail fast, and defer quarantine to a second run - #13936
Merged
Conversation
…failures Run 26877237491 failed: the AWF firewall blocked all 432 requests to dnceng.pkgs.visualstudio.com (the legacy AzDO host backing the dotnet8/9/10 NuGet feeds in NuGet.config), so the whole-repo restore/build could not complete. The 'defaults'/'dotnet' allowlists only cover pkgs.dev.azure.com. Add dnceng.pkgs.visualstudio.com to network.allowed so restore can reach the runtime-pack feeds. The agent also burned the entire 25M effective-token budget (22 min) looping on the failing build before being hard-railed, so it never reached the quarantine fallback or opened a PR. Harden Step 6 to fail fast: a single environmental/network build failure must go straight to the Step 7d quarantine fallback instead of retrying or attempting a fix. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Contributor
There was a problem hiding this comment.
Pull request overview
Updates the flaky-test-detector Agentic Workflow configuration to avoid restore/build failures caused by the legacy AzDO feed host being blocked by the firewall, and to prevent token-budget exhaustion by guiding the agent to quarantine rather than loop when the environment is broken.
Changes:
- Allowlist
dnceng.pkgs.visualstudio.comin the workflow’snetwork.allowedso NuGet restore can reach the dotnet8/9/10 feeds referenced byNuGet.config. - Add “fail fast” guidance to the build step so an initial environment/network build failure immediately triggers the Step 7d quarantine fallback instead of repeated rebuild attempts.
- Recompile the generated
.agent.lock.ymlso the firewall/allowed-domain metadata matches the updated frontmatter.
Show a summary per file
| File | Description |
|---|---|
.github/workflows/flaky-test-detector.agent.md |
Adds the legacy feed host to the allowlist and documents a fail-fast path to quarantine when the initial build failure is environmental. |
.github/workflows/flaky-test-detector.agent.lock.yml |
Regenerates the lock so the runtime allowed-domain lists include dnceng.pkgs.visualstudio.com and prompt/config hashes reflect the updated source. |
Copilot's findings
- Files reviewed: 2/2 changed files
- Comments generated: 0
Analysis of run 26877237491 showed the agent created 4 tracking issues via create_issue, then exhausted the 25M effective-token budget trying to discover their issue numbers to write [ActiveIssue(.../issues/NNNN)] attributes in the same run. That is impossible: create_issue is a safe output filed by a post-run job and returns no number during the run. - Step 5: a flake is quarantine-eligible only if its tracking issue was already OPEN before this run (real number readable now). A just-created issue does not count and must not be quarantined this run; it becomes eligible next run. Explicitly forbid polling/guessing for a just-created issue's number (the behavior that exhausted the budget). - Step 4: note create_issue returns no number; stress the dedup marker must be copied exactly as <!-- flaky-test-id: ... --> (the agent dropped the '-id', which would orphan the issue from future marker searches); add a title-search fallback for legacy issues lacking the marker; note the benign 'Malformed version:' gh stderr warning so it isn't misread as a command failure. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Mirror the workflow behavioral fixes in the skill doc: explain that create_issue is a deferred safe output returning no number during the run, so a new flake is filed one run and quarantined the next (never poll for a just-created issue's number); and stress the dedup marker must be copied exactly with the -id segment, with a title-search fallback for legacy issues. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Contributor
🔍 Skill Validator Results
Summary
Full validator output```text Found 1 skill(s) [flaky-test-detector] 📊 flaky-test-detector: 3,897 BPE tokens [chars/4: 3,794] (standard ~), 13 sections, 5 code blocks [flaky-test-detector] ⚠ Skill is 3,897 BPE tokens (chars/4 estimate: 3,794) — approaching "comprehensive" range where gains diminish. ✅ All checks passed (1 skill(s)) ``` |
Run 26877237491 ran the detector twice with identical inputs: it was launched backgrounded with heavy -MaxBuilds/-MaxArtifactDownloads, polled after ~65s while the scan was still running, saw no -JsonOut file yet (the script writes JSON only on completion), misread that as a failure, and relaunched -- re-downloading every AzDO artifact. Instruct the agent to run the detector in the foreground and wait for exit, and note that a missing -JsonOut file mid-run means 'not finished', not 'failed'. Mirror a short note in the skill doc. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The flaky-test-detector SKILL.md was a separate discoverable skill whose content the daily workflow already re-stated in its own steps (thresholds, dedup marker, quarantine/un-quarantine conventions, flake-vs-regression classification, assembly->project mapping). Its only unique content was the evidence-source model and a gloss of the detector JSON fields. Fold that essential background into a new 'Background' section of flaky-test-detector.agent.md so the workflow is self-contained, move the detector script to .github/workflows/scripts/Get-FlakyTests.ps1 next to the workflow that calls it (a scripts/ dir with no SKILL.md would be a malformed skill), update the two invocation paths and the assembly-mapping reference, and delete SKILL.md. No other workflow or pipeline referenced the skill or script; the workflow body is runtime-imported so no lock recompile is needed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
AlesProkop
approved these changes
Jun 3, 2026
ViktorHofer
enabled auto-merge (squash)
June 3, 2026 11:38
This was referenced Jun 4, 2026
This was referenced Aug 11, 2026
Bump Microsoft.Build from 18.4.0 to 18.9.6
SkylineCommunications/Skyline.DataMiner.CICD.Packages#172
Open
This was referenced Aug 19, 2026
Closed
Open
This was referenced Aug 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Workflow run 26877237491 failed (
agentjob, exit 1). Investigation found three compounding causes — one network/infra and two behavioral design flaws.1. Network restriction — legacy AzDO feed host was firewall-blocked
The AWF firewall denied all 432 requests to
dnceng.pkgs.visualstudio.com:443(the only blocked domain in the run):NuGet.configsplits feeds across two AzDO hosts.pkgs.dev.azure.comis covered by thedotnetecosystem allowlist, but thedotnet8/dotnet9/dotnet10feeds live on the legacydnceng.pkgs.visualstudio.comhost, which was not allowlisted. With the runtime-pack feed blocked, NuGet restore (and therefore the whole-repo build) failed.2. Token-budget exhaustion: chasing the number of an issue it just created
The deeper root cause was not the build loop. After fast-failing the build, the agent created 4 tracking issues via
create_issue, then spent 20+ tool calls trying to discover their issue numbers so it could write[ActiveIssue(.../issues/NNNN)]in the same run. That is architecturally impossible:create_issueis a gh-aw safe output filed by a post-run job — the tool returns only{"result":"success"}with no number during the agent run. The flailing ballooned context into the 25M effective-token hard rail:It was hard-railed before writing any safe output → no PR, exit 1. This wall is latent even on a green build, so the fast-fail fix below is necessary but not sufficient on its own.
3. Dedup-marker drift
The agent wrote
<!-- flaky-test: … -->markers while every dedup search looks for<!-- flaky-test-id: … -->, and it re-filed pre-existing issue #13762 (which predates the marker convention).Fix
network.allowed: adddnceng.pkgs.visualstudio.comso restore can reach the dotnet8/9/10 feeds. Recompiled the lock (the firewall allowlist is frontmatter-derived).create_issuereturns no number during the run; stress the marker must be copied exactly as<!-- flaky-test-id: … -->; add a title-search fallback for legacy issues lacking the marker; note the benignMalformed version:gh stderr warning so it is not misread as a failure.The workflow body is runtime-imported (
{{#runtime-import …}}), so the Step 4/5 prose edits take effect without regenerating the lock; only the network allowlist required a recompile.Also in this PR: merge the detector skill into the workflow
The separate
.github/skills/flaky-test-detector/SKILL.mdduplicated guidance the workflow already states in its own steps; its only unique content (the evidence-source model + detector JSON field glosses) is now a compact Background section inflaky-test-detector.agent.md. The detector script moved to.github/workflows/scripts/Get-FlakyTests.ps1(next to the workflow that calls it), andSKILL.mdwas deleted. Net −241 lines, one fewer moving part. No other workflow or pipeline referenced the skill/script; body is runtime-imported so no lock recompile is needed.