Skip to content

[Bug]: Workspace skill catalogs retain failed discovery and discard newer scan results #16

Description

@Igloczek

Summary

Workspace skill discovery can record a failed probe as a successful scan, preserve an obsolete catalog after an explicit provider refresh, and discard a newer scan when discovery requests overlap. The composer can consequently offer an incomplete or outdated skill inventory even though a fresh provider probe returns the installed skills.

Refresh and freshness handling also have related failure paths: an unchanged expired catalog can suppress composer retries, native session/command updates can renew the age of skills without scanning them, and client/server clock differences can extend catalog reuse beyond its five-minute lifetime.

Area

apps/server provider discovery and catalog publication, packages/client-runtime catalog selection, and the web/mobile composer refresh paths.

Version and environment

  • Tested revision: main at 611132c171f3a821bd2e32f22261135cef6330ac.
  • Backend reproduction used the actual ProviderRegistry, managed Codex driver, and Antigravity provider services in an isolated source copy with lockfile dependencies.
  • The Codex test used a local protocol peer and a controlled subprocess-spawn failure. It did not require a real account or model call.
  • Composer reproduction executed the exact web/mobile refresh-effect bodies with controlled RPC responses and time.
  • TestClock controlled expiry; Deferred gates controlled concurrency. No timing sleeps were used.

Reproduction 1: failed Codex discovery becomes a completed personal-only catalog

  1. Establish a healthy managed Codex machine snapshot containing a personal skill. Make a project skill available to workspace discovery for one directory.
  2. Make the first workspace probe fail while keeping the machine snapshot healthy.
  3. Call refreshWorkspaceSnapshot({ instanceId, cwd }).
  4. Restore the probe so discovery can succeed, then immediately call ordinary workspace discovery again for the same instance and directory.

Actual: the first request publishes a completed workspace catalog containing only the personal skill. The next request returns that cached catalog rather than discovering the project skill. The observed inventory is [personal]; the expected recovered inventory is [personal, project].

Controls: refreshWorkspaceSnapshot({ instanceId, cwd, fresh: true }) immediately returns both skills. An ordinary request after five-minute expiry also returns both skills when machine health is healthy. The expiry test completes machine health refresh after advancing virtual time, before measuring workspace discovery.

Expected: a failed workspace probe may leave a useful fallback visible, but must remain distinguishable from a successful workspace scan and eligible for normal retry.

The managed Codex driver catches workspace-probe failure and returns its machine snapshot. The registry then stamps that fallback with a new workspace scan time. See CodexManagedProvider and workspace publication.

Reproduction 2: explicit provider refresh preserves obsolete workspace skills

  1. Successfully discover a workspace's project skill.
  2. Add a second skill to the next workspace probe result without changing the directory or provider instance.
  3. Call refreshInstance(instanceId), the service used by provider-status refresh.
  4. Request ordinary workspace discovery before the existing catalog expires.

Actual: the old one-skill workspace catalog remains effective and the new skill is missing.

Control: request ordinary discovery after the catalog expires; the new skill appears.

Expected: explicitly refreshing provider status invalidates or revalidates relevant workspace catalogs, so the next discovery can reflect installed skills without waiting for cache expiry.

Machine refresh preserves directory snapshots, and the client prefers a matching directory snapshot over the machine inventory. See refresh methods, catalog merge, and catalog selection.

Reproduction 3: overlapping discovery discards the newer forced result

  1. Seed a successful workspace catalog, then expire it.
  2. Start scan A. Let it read the older skills, then pause at the post-probe instance check before publication.
  3. Start forced scan B from the same catalog baseline. Let it read an additional newly installed skill, then pause before its probe returns.
  4. Release A and wait for it to publish.
  5. Release B and await both requests.

Actual: the final catalog contains only the older skill. B's newer result is discarded because A changed the catalog after B captured its baseline.

Expected: a newer forced scan retains publication ownership and its accepted result cannot be discarded because the older scan published during its execution.

The reproduction uses deterministic gates and asserts the final inventory. See scan claims and the baseline comparison.

Reproduction 4: an expired unchanged response suppresses composer retry

  1. Start with a complete workspace catalog older than five minutes.
  2. Trigger composer discovery, make that scan fail, and have the RPC return the registry's retained expired catalog.
  3. Restore workspace discovery.
  4. Trigger the composer refresh effect again after its 10-second retry interval.

Actual: both the web and mobile effects treat the unchanged complete catalog as a successful refresh. The observed RPC count remains one; the request timestamp suppresses another attempt for five minutes.

Control: the registry itself can recover when called again after an expired scan fails. The suppression occurs in the composer refresh decision.

Expected: retaining a complete fallback does not certify that the refresh succeeded. An expired, unchanged result should preserve the failure/retry path.

Both callbacks check completeness when handling the response: web composer, mobile composer.

Additional freshness and command observations

Reproduction Actual behavior Expected behavior
Discover Antigravity workspace skills, advance time four minutes, then publish session-start metadata without another skill scan Skills remain unchanged, but workspace checkedAt advances four minutes Metadata publication preserves the time of the accepted skill scan
Repeat with a cwd-specific Antigravity command update The same timestamp advancement occurs Command changes do not renew the age of unrescanned skills
Publish Antigravity commands without a cwd after a workspace has been discovered Machine commands update, but the workspace retains its empty command list Applicable discovered workspaces receive the current native command inventory
Evaluate freshness after five elapsed minutes with a client clock ten minutes behind the server Server time considers the catalog expired; client time considers it current, extending client reuse to approximately fifteen elapsed minutes Remote clock differences do not extend the intended catalog lifetime

These were reproduced through the actual Antigravity service and contracts freshness helper. See native metadata publication and freshness calculation.

Verification

Focused diagnostic tests produced these results:

Scenario Result
Immediate ordinary recovery after a managed Codex workspace failure Reproduced missing project skill
Managed Codex forced and post-expiry recovery controls Both passed
Explicit provider refresh followed by ordinary workspace discovery Reproduced obsolete catalog
Ordinary discovery after expiry following provider refresh Passed
Registry recovery after an expired error result Passed
Overlapping older and newer scan publication Reproduced discarded newer result
Web and mobile retry after an expired unchanged response Reproduced on both refresh-effect paths
Antigravity session/command timestamp preservation Reproduced incorrect timestamp changes in both cases
Antigravity cwd-less command propagation Reproduced stale workspace commands
Freshness with client/server clock skew Reproduced extended client cache lifetime

The following existing checks also passed five selected tests, covering TTL boundaries, catalog selection/completeness, and current-versus-expired registry behavior:

pnpm exec vp test run \
  packages/client-runtime/src/providerSkills.test.ts \
  apps/server/src/provider/ProviderRegistry.test.ts \
  -t 'workspace provider snapshots|deduplicates cwd probes'

Diagnostic tests were added only to the isolated source copy. Backend checks used real services with controlled dependencies; composer checks used extracted refresh-effect bodies. No mounted client, mobile device, remote/relay/tunnel runtime, or skill-body execution verification was performed. No repo-wide checks were run.

Expected behavior and scope

Successful workspace scans may remain cached for their configured lifetime. Explicit invalidation, failed discovery, and scan supersession must not be confused with successful reuse. Discovery timestamps should describe accepted skill scans, while native command publication should preserve current commands without making unscanned skills appear fresh.

Preserve bounded/coalesced requests, directory/environment/instance isolation, and provider-native discovery. The clock-skew and command-propagation observations are independently reproducible; the discovery and refresh failures also occur with synchronized clocks and without native command updates.

Related reports

  • #11497: provider refresh does not update cached skills after installation.
  • #14801: project skills installed after worktree creation do not appear in the picker. That hook timing was not tested here.
  • #11575: workspace command discovery captures an incomplete initial provider snapshot.
  • #16750: workspace-catalog TTL and composer freshness checks.

Investigated with GPT-6.1 through Codex in T3 Code.

Activity

  1. added
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 7, 2026
  2. changed the title [-][Bug]: Workspace skill discovery can cache failed probes and publish stale catalogs[/-] [+][Bug]: Workspace skill refresh still mishandles failures and concurrent discovery[/+] on Oct 7, 2026
  3. changed the title [-][Bug]: Workspace skill refresh still mishandles failures and concurrent discovery[/-] [+][Bug]: Workspace skill catalogs retain failed discovery and discard newer scan results[/+] on Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.needs-triageIssue needs maintainer review and initial categorization.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions