Skip to content

Add weekly CI test-selection audit workflow - #20324

Draft
Ankit Jain (radical) wants to merge 4 commits into
microsoft:mainfrom
radical:test-selection-audit-factory
Draft

Ankit Jain (radical) wants to merge 4 commits into
microsoft:mainfrom
radical:test-selection-audit-factory

Conversation

@radical

@radical Ankit Jain (radical) commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Aspire's PR CI uses conditional test selection to avoid running unrelated
tests and jobs. As the codebase evolves, the project graph and curated rules
for runtime-only dependencies can drift, causing CI either to run too much or
miss required coverage.

This workflow runs weekly to audit recent selection evidence for that drift.
When it finds a high-confidence correction, it files at most one actionable
issue and assigns it to the Copilot coding agent, which implements and
validates the update and opens a PR for human review.

When it runs

  • Every Monday, examining PRs from the previous 14 days by default.
  • Manually, with either a different lookback period or an explicit list of
    PR numbers.

The 14-day window intentionally overlaps two weekly runs. The overlap gives
the audit another chance to observe CI that completed late, a rerun that
changed the selection, or a week when the audit did not complete. Previously
processed run attempts are reused rather than downloaded and counted again.

Raw PR-head observations are retained for a rolling 14 days. Before each
audit, deterministic compaction removes older rows and recomputes active
counts and examples. Settled decisions remain durable so known correct,
filed, in-flight, and fixed cases are not repeatedly rediscovered.

What it does

The workflow collects deterministic selection-time evidence from each PR's
latest completed CI run before the agent starts. It compares selected tests
and jobs with the repository's project graph, trigger map, runtime-only
dependencies, and relevant historical changes.

Only attributable same-repository evidence is credited. Fork-produced
selection artifacts are reported as data gaps because the fork controls the
workflow and selector code that produced them.

What it looks for

  • Over-selection: A changed file caused ALL tests and jobs to run even
    though it has a smaller known set of consumers, or should not affect test
    selection at all.
  • Under-selection: A narrow result omitted a test or job that consumes a
    changed package, template, generated artifact, fixture, script, or other
    runtime-only input.
  • High-impact gaps: Incorrect routing for Aspire.Hosting,
    Aspire.TypeSystem and code generation, the dashboard, the VS Code
    extension, the CLI, and CI workflow infrastructure.
  • Repeated problems: The same path or missing consumer appearing across
    multiple PR heads in the rolling window, rather than a single ambiguous
    example.

When the correct behavior is uncertain, the audit favors broader coverage
instead of reducing CI unsafely.

Safety boundaries

  • The collector writes immutable evidence and pre-agent ledger snapshots.
  • Only enforcing-mode selection artifacts are credited. Audit-mode artifacts
    are data gaps because they describe advisory selection, not what CI ran.
  • Agent-edited memory is restricted to processed-runs.jsonl and
    watchlist.jsonl, then validated against the protected evidence before it
    can be pushed.
  • Malformed schemas, unsupported lifecycle transitions, incorrect counters,
    or observations without corresponding evidence fail the workflow.
  • Safe-output issue processing is blocked unless the agent job completes
    successfully.
  • Findings are proposed through safe outputs rather than filed directly by
    the agent.
  • Agent inference is capped at 300 AIC per run, with a rolling daily
    guardrail of 600 AIC.

Validation

  • gh aw compile test-selection-audit --approve --validate --actionlint
  • Full repository build with native compilation skipped.
  • 9 focused C# workflow, validator, compactor, layout, and credit-limit tests.
  • 38 Python collector tests covering artifact validation, provenance,
    pagination, retries, bounded scope, and deterministic evidence handling.
  • A GitHub-hosted run against fork-authored evidence from Add weekly CI test-selection audit workflow #20324 correctly
    treated it as an untrusted data gap and made no persistent change:
    run 36025027688.
  • A GitHub-hosted run against same-repository PR Fix truncation of selected dashboard dropdown values #20426 recorded exactly one
    processed-head row:
    run 36025925235.
  • Repeating the same run recognized the existing record, left the memory
    branch unchanged, and did not produce a duplicate row or issue:
    run 36027308544.

@aspire-repo-bot
aspire-repo-bot Bot requested a balanced review from Copilot September 22, 2026 21:30
@github-actions

Copy link
Copy Markdown
Contributor

🚀 Dogfood this PR with:

⚠️ WARNING: Do not do this without first carefully reviewing the code of this PR to satisfy yourself it is safe.

curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 20324

Or

  • Run remotely in PowerShell:
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 20324"

@github-actions github-actions Bot added the area-engineering-systems infrastructure helix infra engineering repo stuff label Sep 22, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new .gitattributes routing can under-select tests, and several audit safeguards are incomplete.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 4 Medium severity

Open (5)
What changed in this PR

Adds a weekly agentic workflow to audit excessive ALL test selections and applies several initial selector refinements.

Changes:

  • Adds and compiles the weekly audit workflow.
  • Narrows several CI trigger-map routes and documents them.
  • Adds regression guards and removes an orphaned shared source file.
File Description
.github/​workflows/​test-selection-audit.md Defines the weekly audit.
.github/​workflows/​test-selection-audit.lock.yml Compiled workflow.
.github/​aw/​actions-lock.json Updates workflow action locks.
eng/​github-ci/​test-trigger-map.yml Refines test routing.
eng/​github-ci/​ci-skip-entirely-patterns.txt Documents .gitattributes handling.
tests/​Infrastructure.Tests/​TestTriggerMap/​TestTriggerMapTests.cs Adds routing guards.
tests/​Infrastructure.Tests/​TestTriggerMap/​SelectTestsAcceptanceTests.cs Updates prefilter expectations.
tests/​Infrastructure.Tests/​GitAttributesTests.cs Validates critical line-ending rules.
src/​Shared/​Utf8JsonWriterExtensions.cs Removes orphaned shared code.
docs/​ci/​test-trigger-map.md Documents routing changes.
docs/​ci/​test-trigger-selector-design.md Documents fleet-wide auditing.

Comment thread eng/github-ci/test-trigger-map.yml Outdated
Comment thread .github/workflows/test-selection-audit.md Outdated
Comment thread tests/Infrastructure.Tests/TestTriggerMap/TestTriggerMapTests.cs Outdated
Comment thread tests/Infrastructure.Tests/TestTriggerMap/TestTriggerMapTests.cs Outdated
Comment thread tests/Infrastructure.Tests/TestTriggerMap/TestTriggerMapTests.cs Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The workflow cannot reliably create labeled issues or inspect targeted fork PRs, lacks deduplication, and uses a compiler version that fails repository validation.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 3 High severity · 4 Medium severity

Open (7)
Resolved since last review (1)

Comment thread .github/workflows/test-selection-audit.lock.yml Outdated
Comment thread .github/workflows/test-selection-audit.md Outdated
Comment thread .github/workflows/test-selection-audit.md Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The compiler version conflicts with repository validation, fork PRs are filtered by the ineffective integrity setting, and the configured ci label does not exist.

Review effort: Balanced
Findings: 3 High severity · 4 Medium severity

Open (7)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Fork PRs are filtered out, and the generated lock uses a compiler version inconsistent with the repository bootstrap.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 4 High severity · 4 Medium severity

Open (8)

Comment thread .github/workflows/test-selection-audit.lock.yml Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread .github/aw/actions-lock.json Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread .github/workflows/test-selection-audit.md Outdated
Comment thread .github/workflows/test-selection-audit.md Outdated
Comment thread .github/workflows/test-selection-audit.md Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The compiler mismatch breaks existing tests, and the audit has fork-filtering and persistence logic conflicts.

Review effort: Balanced
Findings: 4 High severity · 7 Medium severity · 1 Low severity

Open (12)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Remove incorrect permanent exemption for Aspire.slnx

.github/​workflows/​test-selection-audit.md:257

Aspire.slnx is not a correct-by-design broad input: the linked #20322 establishes that Layer 1 already roots its graph at the PR-head solution and narrows this 99/99 fallback. Hard-exempting it means the audit can never rediscover this over-selection if that PR closes unmerged or a future map change reintroduces it; step 8 already handles fixes that are in flight. Remove it from this permanent exemption list.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The lock uses an inconsistent compiler version, and the audit contains candidate-detection and state-retention errors.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 5 High severity · 7 Medium severity · 1 Low severity

Open (13)
Previously missed (3)

In code that hasn't changed since last review

Medium severity Remove Aspire.slnx from the correct-by-design exemption list

.github/​workflows/​test-selection-audit.md:319

Aspire.slnx is not a correct-by-design broad input: PR #20322 demonstrates that Layer 1 already roots its graph from the PR-head solution and narrows this case safely. Exempting it here prevents step 8 from recognizing that open fix and permanently records future occurrences as intended, contrary to this PR's stated audit behavior. Remove it from the exemption list and let the normal candidate/in-flight checks handle it.

Medium severity Inspect PR file lists to detect in-flight fixes

.github/​workflows/​test-selection-audit.md:362

GitHub pull-request search does not index changed-file paths, so search_pull_requests cannot establish that an open PR touches test-trigger-map.yml. This can miss an in-flight fix and start a duplicate issue/agent session. Search can find PRs naming the rule, but file-based detection must inspect candidate PR file lists.

Medium severity Retain correct-by-design rows to preserve settled decisions

.github/​workflows/​test-selection-audit.md:449

Dropping a correct-by-design row contradicts lines 179 and 198-201, which rely on that durable verdict to avoid re-reading the map and history. Once dropped, the next occurrence is treated as new, so the workflow repeats the expensive analysis and can change a settled decision. Retain settled rows; only remove entries whose escalation was eliminated by a merged fix.

Comment thread .github/workflows/test-selection-audit.lock.yml Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Valid UTF-8 repository paths are rejected, a test is midnight-sensitive, and artifact security checks lack behavioral coverage.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity · 1 Low severity

Open (3)
Resolved since last review (2)

Comment thread tests/Infrastructure.Tests/WorkflowScripts/AgenticWorkflowTests.cs Outdated
Comment thread .github/workflows/test-selection-audit/collect_evidence.py Outdated
Comment thread .github/workflows/test-selection-audit/collect_evidence.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Memory validation currently prevents removing the final zero-count watch entry after a trusted rerun retracts its contribution.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity · 1 Low severity

Open (3)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Allow deletion of zero-count baseline watch rows

.github/​workflows/​test-selection-audit.md:506

This unconditional baseline check rejects the workflow's required removal of a zero-count watch row. If a newer trusted CI attempt replaces the last contribution for a path/edge, affectedWatchKeys contains that key and the recomputed count is zero, but deleting the row still fails as “missing baseline watch row,” so an ALL-to-narrow rerun cannot retract the final watch entry. Allow deletion only for affected baseline rows whose verdict is watch; the following overCounts/missCounts checks still reject deletion when any contribution remains.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Memory validation incorrectly rejects removal of zero-count watch rows after trusted reruns retract their final contribution.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 High severity · 1 Medium severity · 1 Low severity

Open (4)

Comment thread .github/workflows/test-selection-audit.md

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

GitHub API rate-limit responses are handled incorrectly, making the recurring audit unreliable under throttling.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Resolved since last review (4)

Comment thread .github/workflows/test-selection-audit/collect_evidence.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Two compaction tests attempt to delete read-only files and will fail on Windows.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity

Open (2)

Comment thread tests/Infrastructure.Tests/WorkflowScripts/AgenticWorkflowTests.cs Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Memory validation does not recompute retained examples or date ranges, allowing inconsistent audit evidence to pass.

Review effort: Balanced
Findings: None

Resolved since last review (2)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Validate dates against contributing processed-run records

.github/​workflows/​test-selection-audit.md:424

The validator only checks that the dates are well-formed and ordered; it never derives them from the contributing processed-runs.jsonl rows. For any currently affected key, the agent can therefore write arbitrary historical dates and still pass validation, even though the workflow relies on these dates to characterize recurrence and explicitly claims they are independently recomputed. Compare nonzero-count rows against the minimum/maximum contributing seen dates (and preserve the prior dates only for settled zero-count rows).

Medium severity Require examples for nonzero-count affected rows

.github/​workflows/​test-selection-audit.md:487

This is only a subset check, so example_prs: [] passes even when actualCount is nonzero. That lets an affected row discard every retained example while the counter remains valid, weakening the evidence used for repeat findings and contradicting the stated independent recomputation guarantee. Require exactly min(3, contributingPrs.size) distinct credited examples (or define and compare a deterministic expected list).

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Artifact validation accepts noncanonical job names, which can produce false under-selection evidence.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)

Comment thread .github/workflows/test-selection-audit/collect_evidence.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The primary scheduled pagination and cutoff path lacks focused regression coverage and could silently omit PRs.

Review effort: Balanced
Findings: None

Resolved since last review (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Add pagination and cutoff regression tests for automatic lookback

.github/​workflows/​test-selection-audit/​collect_evidence.py:346

The automatic lookback branch is the workflow's default weekly path, but the added tests only exercise the explicit PR_NUMBERS branch. There is no regression coverage for cutoff inclusion, page-boundary stopping, or enumerationTruncated, so a pagination change could silently omit PRs while the audit still completes successfully. Add mocked multi-page tests for the automatic path, including a cutoff-boundary PR and the 10-page/MAX_PRS truncation cases.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The collector can misinterpret audit-mode artifacts as enforced narrow selections and persist false findings.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)

Comment thread .github/workflows/test-selection-audit/collect_evidence.py
Audit recent pull request test selections against changed files, selector
output, and repository history so repeated over-selection and missed
consumers can be surfaced from bounded evidence.

Collect GitHub evidence before agent execution, persist only validated
JSONL memory, and stage issue proposals through safe outputs. Cap each
agent run at 300 AIC and rolling daily usage at 600 AIC.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Exercise the authored workflow and generated lock together, including the
memory validator trust boundaries, compaction behavior, and AI-credit
limits.

Cover collector artifact validation, fork provenance, retries, bounded
scope, and deterministic evidence handling with focused Python tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Issue creation can still proceed when repository-memory validation fails.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)

Comment thread .github/workflows/test-selection-audit.md
Audit-mode selector artifacts describe advisory selection while CI runs
all tests and jobs. Reject them as enforced evidence so they cannot
produce false under-selection findings.

Block safe-output processing unless the agent job succeeds. This keeps
failed memory validation or another agent failure from creating an issue.

Add regression coverage for both trust boundaries.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The scheduled enumeration path lacks regression coverage, and a Windows test leaves a read-only temporary file behind.

Review effort: Balanced
Findings: None

Resolved since last review (2)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Add regression coverage for paginated default PR scope

.github/​workflows/​test-selection-audit/​collect_evidence.py:378

The weekly scheduled path always uses this paginated lookback branch, but PullRequestScopeTests only exercises explicit PR_NUMBERS. There is no regression coverage for cutoff filtering, open-versus-closed timestamps, page termination, or the truncation flag, so the workflow's default scope could silently omit PRs while all 34 collector tests pass. Add a mocked multi-page test with no explicit scope that covers in-window/out-of-window records and truncation.

Medium severity Reset audit-date.txt read-only attribute after final compaction

tests/​Infrastructure.Tests/​WorkflowScripts/​TestSelectionAuditWorkflowTests.cs:682

The final compactor invocation recreates audit-date.txt with mode 0444, but this path is not reset afterward as it is before the earlier reruns. On Windows that leaves a read-only file, so TemporaryWorkspace.Dispose() cannot remove the workspace and silently leaves it behind (the cleanup helper catches UnauthorizedAccessException). Clear the attribute after the final assertions.

The weekly audit enumerates recent pull requests when no explicit scope is
provided, but that path had no regression coverage. A pagination or cutoff
change could silently omit evidence while the collector tests still passed.

Cover open and closed pull request timestamp semantics, cutoff and short-page
termination, and both the page-count and configured PR truncation limits.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@radical

Copy link
Copy Markdown
Member Author

[automated] Addressed the latest Copilot review summary in eccb4edc090.

  • Added focused regression coverage for the scheduled, paginated PR scope: open and closed timestamp semantics, cutoff and short-page termination, the ten-page cap, and the configured PR limit. The collector suite now has 38 tests.
  • Declined the audit-date.txt cleanup suggestion because TemporaryWorkspace.DeleteDirectoryWithRetries already clears Windows ReadOnly attributes recursively, retries deletion, and uses an individual-file fallback. The file is intentionally read-only to match the workflow's immutable provenance; the existing explicit resets are required only before another compactor invocation rewrites it.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The autonomous issue-filing and persistent-memory workflow has broad operational and security implications requiring final human review.

Review effort: Balanced
Findings: None

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-engineering-systems infrastructure helix infra engineering repo stuff

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants