You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Aspire's PR CI uses conditional test selection to avoid running unrelated
tests and jobs. As the codebase evolves, the project graph and curated rules
for runtime-only dependencies can drift, causing CI either to run too much or
miss required coverage.
This workflow runs weekly to audit recent selection evidence for that drift.
When it finds a high-confidence correction, it files at most one actionable
issue and assigns it to the Copilot coding agent, which implements and
validates the update and opens a PR for human review.
When it runs
Every Monday, examining PRs from the previous 14 days by default.
Manually, with either a different lookback period or an explicit list of
PR numbers.
The 14-day window intentionally overlaps two weekly runs. The overlap gives
the audit another chance to observe CI that completed late, a rerun that
changed the selection, or a week when the audit did not complete. Previously
processed run attempts are reused rather than downloaded and counted again.
Raw PR-head observations are retained for a rolling 14 days. Before each
audit, deterministic compaction removes older rows and recomputes active
counts and examples. Settled decisions remain durable so known correct,
filed, in-flight, and fixed cases are not repeatedly rediscovered.
What it does
The workflow collects deterministic selection-time evidence from each PR's
latest completed CI run before the agent starts. It compares selected tests
and jobs with the repository's project graph, trigger map, runtime-only
dependencies, and relevant historical changes.
Only attributable same-repository evidence is credited. Fork-produced
selection artifacts are reported as data gaps because the fork controls the
workflow and selector code that produced them.
What it looks for
Over-selection: A changed file caused ALL tests and jobs to run even
though it has a smaller known set of consumers, or should not affect test
selection at all.
Under-selection: A narrow result omitted a test or job that consumes a
changed package, template, generated artifact, fixture, script, or other
runtime-only input.
High-impact gaps: Incorrect routing for Aspire.Hosting,
Aspire.TypeSystem and code generation, the dashboard, the VS Code
extension, the CLI, and CI workflow infrastructure.
Repeated problems: The same path or missing consumer appearing across
multiple PR heads in the rolling window, rather than a single ambiguous
example.
When the correct behavior is uncertain, the audit favors broader coverage
instead of reducing CI unsafely.
Safety boundaries
The collector writes immutable evidence and pre-agent ledger snapshots.
Only enforcing-mode selection artifacts are credited. Audit-mode artifacts
are data gaps because they describe advisory selection, not what CI ran.
Agent-edited memory is restricted to processed-runs.jsonl and watchlist.jsonl, then validated against the protected evidence before it
can be pushed.
Malformed schemas, unsupported lifecycle transitions, incorrect counters,
or observations without corresponding evidence fail the workflow.
Safe-output issue processing is blocked unless the agent job completes
successfully.
Findings are proposed through safe outputs rather than filed directly by
the agent.
Agent inference is capped at 300 AIC per run, with a rolling daily
guardrail of 600 AIC.
Validation
gh aw compile test-selection-audit --approve --validate --actionlint
Full repository build with native compilation skipped.
9 focused C# workflow, validator, compactor, layout, and credit-limit tests.
38 Python collector tests covering artifact validation, provenance,
pagination, retries, bounded scope, and deterministic evidence handling.
Repeating the same run recognized the existing record, left the memory
branch unchanged, and did not produce a duplicate row or issue: run 36027308544.
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🟡 Changes recommended
The workflow cannot reliably create labeled issues or inspect targeted fork PRs, lacks deduplication, and uses a compiler version that fails repository validation.
Get a fresh assessment by requesting another Copilot review.
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
The compiler version conflicts with repository validation, fork PRs are filtered by the ineffective integrity setting, and the configured ci label does not exist.
Remove incorrect permanent exemption for Aspire.slnx
.github/workflows/test-selection-audit.md:257
Aspire.slnx is not a correct-by-design broad input: the linked #20322 establishes that Layer 1 already roots its graph at the PR-head solution and narrows this 99/99 fallback. Hard-exempting it means the audit can never rediscover this over-selection if that PR closes unmerged or a future map change reintroduces it; step 8 already handles fixes that are in flight. Remove it from this permanent exemption list.
Remove Aspire.slnx from the correct-by-design exemption list
.github/workflows/test-selection-audit.md:319
Aspire.slnx is not a correct-by-design broad input: PR #20322 demonstrates that Layer 1 already roots its graph from the PR-head solution and narrows this case safely. Exempting it here prevents step 8 from recognizing that open fix and permanently records future occurrences as intended, contrary to this PR's stated audit behavior. Remove it from the exemption list and let the normal candidate/in-flight checks handle it.
Inspect PR file lists to detect in-flight fixes
.github/workflows/test-selection-audit.md:362
GitHub pull-request search does not index changed-file paths, so search_pull_requests cannot establish that an open PR touches test-trigger-map.yml. This can miss an in-flight fix and start a duplicate issue/agent session. Search can find PRs naming the rule, but file-based detection must inspect candidate PR file lists.
Retain correct-by-design rows to preserve settled decisions
.github/workflows/test-selection-audit.md:449
Dropping a correct-by-design row contradicts lines 179 and 198-201, which rely on that durable verdict to avoid re-reading the map and history. Once dropped, the next occurrence is treated as new, so the workflow repeats the expensive analysis and can change a settled decision. Retain settled rows; only remove entries whose escalation was eliminated by a merged fix.
This unconditional baseline check rejects the workflow's required removal of a zero-count watch row. If a newer trusted CI attempt replaces the last contribution for a path/edge, affectedWatchKeys contains that key and the recomputed count is zero, but deleting the row still fails as “missing baseline watch row,” so an ALL-to-narrow rerun cannot retract the final watch entry. Allow deletion only for affected baseline rows whose verdict is watch; the following overCounts/missCounts checks still reject deletion when any contribution remains.
Validate dates against contributing processed-run records
.github/workflows/test-selection-audit.md:424
The validator only checks that the dates are well-formed and ordered; it never derives them from the contributing processed-runs.jsonl rows. For any currently affected key, the agent can therefore write arbitrary historical dates and still pass validation, even though the workflow relies on these dates to characterize recurrence and explicitly claims they are independently recomputed. Compare nonzero-count rows against the minimum/maximum contributing seen dates (and preserve the prior dates only for settled zero-count rows).
Require examples for nonzero-count affected rows
.github/workflows/test-selection-audit.md:487
This is only a subset check, so example_prs: [] passes even when actualCount is nonzero. That lets an affected row discard every retained example while the counter remains valid, weakening the evidence used for repeat findings and contradicting the stated independent recomputation guarantee. Require exactly min(3, contributingPrs.size) distinct credited examples (or define and compare a deterministic expected list).
The automatic lookback branch is the workflow's default weekly path, but the added tests only exercise the explicit PR_NUMBERS branch. There is no regression coverage for cutoff inclusion, page-boundary stopping, or enumerationTruncated, so a pagination change could silently omit PRs while the audit still completes successfully. Add mocked multi-page tests for the automatic path, including a cutoff-boundary PR and the 10-page/MAX_PRS truncation cases.
Audit recent pull request test selections against changed files, selector
output, and repository history so repeated over-selection and missed
consumers can be surfaced from bounded evidence.
Collect GitHub evidence before agent execution, persist only validated
JSONL memory, and stage issue proposals through safe outputs. Cap each
agent run at 300 AIC and rolling daily usage at 600 AIC.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Audit-mode selector artifacts describe advisory selection while CI runs
all tests and jobs. Reject them as enforced evidence so they cannot
produce false under-selection findings.
Block safe-output processing unless the agent job succeeds. This keeps
failed memory validation or another agent failure from creating an issue.
Add regression coverage for both trust boundaries.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The weekly scheduled path always uses this paginated lookback branch, but PullRequestScopeTests only exercises explicit PR_NUMBERS. There is no regression coverage for cutoff filtering, open-versus-closed timestamps, page termination, or the truncation flag, so the workflow's default scope could silently omit PRs while all 34 collector tests pass. Add a mocked multi-page test with no explicit scope that covers in-window/out-of-window records and truncation.
Reset audit-date.txt read-only attribute after final compaction
The final compactor invocation recreates audit-date.txt with mode 0444, but this path is not reset afterward as it is before the earlier reruns. On Windows that leaves a read-only file, so TemporaryWorkspace.Dispose() cannot remove the workspace and silently leaves it behind (the cleanup helper catches UnauthorizedAccessException). Clear the attribute after the final assertions.
The weekly audit enumerates recent pull requests when no explicit scope is
provided, but that path had no regression coverage. A pagination or cutoff
change could silently omit evidence while the collector tests still passed.
Cover open and closed pull request timestamp semantics, cutoff and short-page
termination, and both the page-count and configured PR truncation limits.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
[automated] Addressed the latest Copilot review summary in eccb4edc090.
Added focused regression coverage for the scheduled, paginated PR scope: open and closed timestamp semantics, cutoff and short-page termination, the ten-page cap, and the configured PR limit. The collector suite now has 38 tests.
Declined the audit-date.txt cleanup suggestion because TemporaryWorkspace.DeleteDirectoryWithRetries already clears Windows ReadOnly attributes recursively, retries deletion, and uses an individual-file fallback. The file is intentionally read-only to match the workflow's immutable provenance; the existing explicit resets are required only before another compactor invocation rewrites it.
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
The autonomous issue-filing and persistent-memory workflow has broad operational and security implications requiring final human review.
Review effort: Balanced Findings: None
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Aspire's PR CI uses conditional test selection to avoid running unrelated
tests and jobs. As the codebase evolves, the project graph and curated rules
for runtime-only dependencies can drift, causing CI either to run too much or
miss required coverage.
This workflow runs weekly to audit recent selection evidence for that drift.
When it finds a high-confidence correction, it files at most one actionable
issue and assigns it to the Copilot coding agent, which implements and
validates the update and opens a PR for human review.
When it runs
PR numbers.
The 14-day window intentionally overlaps two weekly runs. The overlap gives
the audit another chance to observe CI that completed late, a rerun that
changed the selection, or a week when the audit did not complete. Previously
processed run attempts are reused rather than downloaded and counted again.
Raw PR-head observations are retained for a rolling 14 days. Before each
audit, deterministic compaction removes older rows and recomputes active
counts and examples. Settled decisions remain durable so known correct,
filed, in-flight, and fixed cases are not repeatedly rediscovered.
What it does
The workflow collects deterministic selection-time evidence from each PR's
latest completed CI run before the agent starts. It compares selected tests
and jobs with the repository's project graph, trigger map, runtime-only
dependencies, and relevant historical changes.
Only attributable same-repository evidence is credited. Fork-produced
selection artifacts are reported as data gaps because the fork controls the
workflow and selector code that produced them.
What it looks for
ALLtests and jobs to run eventhough it has a smaller known set of consumers, or should not affect test
selection at all.
changed package, template, generated artifact, fixture, script, or other
runtime-only input.
Aspire.TypeSystem and code generation, the dashboard, the VS Code
extension, the CLI, and CI workflow infrastructure.
multiple PR heads in the rolling window, rather than a single ambiguous
example.
When the correct behavior is uncertain, the audit favors broader coverage
instead of reducing CI unsafely.
Safety boundaries
are data gaps because they describe advisory selection, not what CI ran.
processed-runs.jsonlandwatchlist.jsonl, then validated against the protected evidence before itcan be pushed.
or observations without corresponding evidence fail the workflow.
successfully.
the agent.
guardrail of 600 AIC.
Validation
gh aw compile test-selection-audit --approve --validate --actionlintpagination, retries, bounded scope, and deterministic evidence handling.
treated it as an untrusted data gap and made no persistent change:
run 36025027688.
processed-head row:
run 36025925235.
branch unchanged, and did not produce a duplicate row or issue:
run 36027308544.