Skip to content

[Bug]: Workspace skill catalogs retain failed discovery and discard newer scan results #16866

Description

@Igloczek

Before submitting

  • Searched existing issues; related reports are linked below. This report adds deterministic failed-probe, overlapping-scan, and composer-retry reproductions.
  • Included detail to reproduce or investigate the problem.

Summary

Workspace skill discovery can record a failed probe as a successful scan and discard a newer scan when discovery requests overlap. The composer can consequently offer an incomplete or outdated skill inventory even though a fresh provider probe returns the installed skills. An unchanged expired fallback can also suppress composer retries.

Area

apps/server, with related packages/client-runtime and web/mobile composer refresh paths.

Impact

Minor bug or occasional failure: installed project skills can remain missing from the composer until forced discovery or catalog expiry permits recovery.

Version and environment

  • Tested revision: main at 611132c171f3a821bd2e32f22261135cef6330ac.
  • Backend reproduction used the actual ProviderRegistry, managed Codex driver, and Antigravity provider services in an isolated source copy with lockfile dependencies.
  • The Codex test used a local protocol peer and a controlled subprocess-spawn failure. It did not require a real account or model call.
  • Composer reproduction executed the exact web/mobile refresh-effect bodies with controlled RPC responses and time.
  • TestClock controlled expiry; Deferred gates controlled concurrency. No timing sleeps were used.

Reproduction 1: failed Codex discovery becomes a completed personal-only catalog

  1. Establish a healthy managed Codex machine snapshot containing a personal skill. Make a project skill available to workspace discovery for one directory.
  2. Make the first workspace probe fail while keeping the machine snapshot healthy.
  3. Call refreshWorkspaceSnapshot({ instanceId, cwd }).
  4. Restore the probe so discovery can succeed, then immediately call ordinary workspace discovery again for the same instance and directory.

Actual: the first request publishes a completed workspace catalog containing only the personal skill. The next request returns that cached catalog rather than discovering the project skill. The observed inventory is [personal]; the expected recovered inventory is [personal, project].

Controls: refreshWorkspaceSnapshot({ instanceId, cwd, fresh: true }) immediately returns both skills. An ordinary request after five-minute expiry also returns both skills when machine health is healthy. The expiry test completes machine health refresh after advancing virtual time, before measuring workspace discovery.

Expected: a failed workspace probe may leave a useful fallback visible, but must remain distinguishable from a successful workspace scan and eligible for normal retry.

The managed Codex driver catches workspace-probe failure and returns its machine snapshot. The registry then stamps that fallback with a new workspace scan time. See CodexManagedProvider and workspace publication.

Reproduction 3: overlapping discovery discards the newer forced result

  1. Seed a successful workspace catalog, then expire it.
  2. Start scan A. Let it read the older skills, then pause at the post-probe instance check before publication.
  3. Start forced scan B from the same catalog baseline. Let it read an additional newly installed skill, then pause before its probe returns.
  4. Release A and wait for it to publish.
  5. Release B and await both requests.

Actual: the final catalog contains only the older skill. B's newer result is discarded because A changed the catalog after B captured its baseline.

Expected: a newer forced scan retains publication ownership and its accepted result cannot be discarded because the older scan published during its execution.

The reproduction uses deterministic gates and asserts the final inventory. See scan claims and the baseline comparison.

Reproduction 4: an expired unchanged response suppresses composer retry

  1. Start with a complete workspace catalog older than five minutes.
  2. Trigger composer discovery, make that scan fail, and have the RPC return the registry's retained expired catalog.
  3. Restore workspace discovery.
  4. Trigger the composer refresh effect again after its 10-second retry interval.

Actual: both the web and mobile effects treat the unchanged complete catalog as a successful refresh. The observed RPC count remains one; the request timestamp suppresses another attempt for five minutes.

Control: the registry itself can recover when called again after an expired scan fails. The suppression occurs in the composer refresh decision.

Expected: retaining a complete fallback does not certify that the refresh succeeded. An expired, unchanged result should preserve the failure/retry path.

Both callbacks check completeness when handling the response: web composer, mobile composer.

Verification and limitations

The source investigation reports deterministic diagnostic tests using actual backend services, TestClock for expiry, and Deferred gates for concurrency. Failed-probe recovery, overlapping scan publication, and expired-response retry suppression were reproduced. Forced discovery and post-expiry recovery controls passed. Composer checks executed the extracted web/mobile refresh-effect bodies.

Diagnostic tests were added only to an isolated source copy. No mounted client, mobile device, remote/relay/tunnel runtime, or skill-body execution verification was performed. No repo-wide checks were run. This publication carries the existing investigation; it does not claim a new verification run against current main.

The existing focused checks passed five selected tests:

pnpm exec vp test run \
  packages/client-runtime/src/providerSkills.test.ts \
  apps/server/src/provider/ProviderRegistry.test.ts \
  -t 'workspace provider snapshots|deduplicates cwd probes'

Expected behavior

A fallback may remain visible after failed discovery, but must not certify a successful scan or suppress normal retries. A newer forced scan must retain publication ownership when an older scan finishes first. Preserve bounded/coalesced requests and directory/environment/instance isolation.

Workaround

For the failed managed Codex probe, refreshWorkspaceSnapshot({ instanceId, cwd, fresh: true }) returned the complete inventory. Ordinary discovery also recovered after five-minute expiry with healthy machine state.

Related reports and source

  • #11497: explicit provider refresh retains cached skills; that overlapping finding is not separately reported here.
  • #14801: project skills missing after worktree creation; that hook timing was not tested.
  • #11575: incomplete initial workspace command snapshot.
  • #16750: workspace-catalog TTL and composer freshness checks.
  • Original investigation, including separate clock-skew and native-command observations omitted here to keep this report focused.

Investigated with GPT-6.1 through Codex in T3 Code.

Activity

  1. juliusmarminge commented on Oct 7, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Thanks for the detailed repros. I checked all three on current main (611132c), and all three are real.

    1. A failed Codex workspace probe gets saved as a completed scan

    • CodexManagedProvider.ts:289: Effect.catch(() => snapshot.getSnapshot) turns any snapshotForCwd failure into the machine snapshot, which only has personal skills.
    • ProviderRegistry.ts:1088-1102: only status === "error" results without slashCommandsPending get thrown away. A healthy fallback is saved with checkedAt: scannedAt, so isProviderWorkspaceSnapshotCurrent (L1060) treats it as fresh for the full TTL.
    • Fix: the driver should flag a failed probe (for example, return slashCommandsPending or a "scan failed" marker), and the registry should keep the fallback visible without the fresh scan time, so the next request retries it.

    2. A newer forced scan is thrown away when an older scan publishes first

    • ProviderRegistry.ts:1052 records scannedFrom. At L1077, a fresh scan keeps going even when it didn't get the claim. At L1097, it only writes if the snapshot still Equal.equals that baseline. If the older scan A publishes while B is running, B's newer result is dropped.
    • Fix: give each (instance, cwd) a scan generation. A fresh scan bumps the generation and owns publishing, and an older scan with a lower generation can't block it. Session-event writes can still beat both.

    3. Web and mobile composers treat an expired complete fallback as a successful refresh

    • ChatComposer.tsx:2279-2287 and use-composer-command-menu.ts:325-333 only check hasCompleteProviderWorkspaceSnapshot. That function (providerSkills.ts:115-121) ignores age. In that case retryLater() is skipped, and the requestedAt guard (ChatComposer L2253-2257, mobile L299-303) blocks another request for PROVIDER_WORKSPACE_SNAPSHOT_TTL_MS.
    • Fix: test the response with hasCurrentProviderWorkspaceSnapshot(..., Date.now()), or share one helper between web and mobile, so an expired catalog takes the 10s retry path.

    Overlap

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions