Skip to content

[Bug]: New Codex runtime entries can disable existing shadow-home providers #5817

Description

@zenDevNomad

Summary

T3's Codex shadow-home setup only links entries that already exist in the shared home when materialization runs. If a later Codex update creates a new top-level runtime entry in both the shared and shadow homes, the next materialization rejects the shadow copy and disables the provider.

Codex 0.147 introduced queue_1.sqlite, which triggers this today. The same failure can recur when Codex adds another home entry.

Sanitized T3 Code provider error showing queue_1.sqlite shadow-home conflict

Reproduction

  1. Configure the documented multi-account layout: one shared homePath and another account using that homePath plus a separate shadowHomePath.
  2. Start with no queue_1.sqlite in either home and let T3 materialize the shadow home.
  3. Run Codex 0.147 against both effective homes. Codex creates a regular queue database in each.
  4. Refresh provider status or restart T3.

The shadow-home provider becomes disabled:

Driver 'codex' failed to create instance: Cannot create Codex shadow home entry 'queue_1.sqlite' because '<shadow-home>/queue_1.sqlite' already exists and is not a symlink.

A minimal unit repro is to create regular sharedHome/queue_1.sqlite and shadowHome/queue_1.sqlite files, then call materializeCodexShadowHome(layout).

Why it happens

T3 lists the current shared-home entries, then ensureSymlink rejects any existing regular shadow entry. There is no recovery path; the conflict becomes a provider creation error.

Codex added the queue DB in openai/codex@b87981a. In the affected installation, the shared DB and two shadow DBs were automatically created within six seconds, confirming this was runtime-created state rather than a copied home.

Expected behavior

Codex upgrades that add runtime entries should not disable existing shadow-home providers. T3 should reconcile ownership safely or offer a targeted recovery path without silently deleting user data.

Acceptance criteria

  • New Codex home entries created after an earlier materialization do not leave providers disabled.
  • The solution is resilient to future filenames rather than only adding queue_1.sqlite to an allowlist.
  • Existing data is preserved; SQLite families such as the base DB, WAL, and SHM are handled together.
  • Tests cover creation between materialization passes and an existing regular shadow entry.

Environment

  • T3 main at 1a003e383ac6b10258b8100c2617d938c4f06c69
  • Codex CLI 0.147.0
  • macOS 26.5.2, Apple Silicon

Workaround

With related processes stopped, backing up and removing the shadow queue_1.sqlite, queue_1.sqlite-wal, and queue_1.sqlite-shm files allows T3 to create the shared link. This is manual, potentially data-destructive, and does not address future Codex entries.

Activity

  1. changed the title [-][Bug]: Codex 0.147 queue DB disables shadow-home provider instances[/-] [+][Bug]: New Codex runtime entries can disable existing shadow-home providers[/+] on Aug 9, 2026
  2. EwenGG commented on Aug 20, 2026

    @EwenGG

    I can confirm this issue on Windows, with an additional sidecar-only variant that may be relevant to the proposed fix in #7201.

    Environment

    • T3 Code Nightly: 0.0.34-nightly.20260820.1141
    • Codex CLI: 0.148.0
    • OS: Windows 11 Pro, version 10.0.26200, x64
    • Provider setup: two Codex provider instances using one shared home and a separate shadow home for the second account
    • Authentication: ChatGPT subscription authentication
    • The default Codex provider remains usable, but the shadow-home provider is disabled

    Sanitized configuration:

    {
      "driver": "codex",
      "enabled": true,
      "config": {
        "homePath": "C:\\Users\\<user>\\.codex",
        "shadowHomePath": "C:\\Users\\<user>\\.codex_p"
      }
    }

    Observed behavior

    This has recurred across several recent T3 Code Nightly and Codex updates.

    On startup or provider refresh, the secondary Codex provider is disabled during materializeCodexShadowHome. T3 encounters a regular SQLite runtime sidecar in the shadow home where it expects a symbolic link, then aborts provider creation.

    Exact error

    Driver 'codex' failed to create instance:
    Cannot create Codex shadow home entry 'logs_2.sqlite-shm'
    because 'C:\Users\<user>\.codex_p\logs_2.sqlite-shm'
    already exists and is not a symlink.
    

    The corresponding server trace reports:

    CodexShadowHomeEntryConflictError:
    Cannot create Codex shadow home entry 'logs_2.sqlite-shm'
    because 'C:\Users\<user>\.codex_p\logs_2.sqlite-shm'
    already exists and is not a symlink.
    
    at CodexHomeLayout.ensureSymlink
    at materializeCodexShadowHome
    

    Previous occurrences on the same machine failed on other members of the same class of runtime files:

    memories_1.sqlite-shm
    queue_1.sqlite
    logs_2.sqlite-shm
    

    Because materialization stops at the first conflict, the filename displayed by T3 does not represent the complete set of conflicting entries.

    Shadow-home state at failure time

    A complete inspection of the shadow-home root found eight regular SQLite sidecar files:

    Shadow-home entry Type Size
    logs_2.sqlite-shm regular file 32,768 bytes
    logs_2.sqlite-wal regular file 4,466,112 bytes
    memories_1.sqlite-shm regular file 32,768 bytes
    memories_1.sqlite-wal regular file 4,152 bytes
    queue_1.sqlite-shm regular file 32,768 bytes
    queue_1.sqlite-wal regular file 4,152 bytes
    state_5.sqlite-shm regular file 32,768 bytes
    state_5.sqlite-wal regular file 16,512 bytes

    The problem was therefore not limited to queue_1.sqlite.

    Importantly, the corresponding base database entries were still symbolic links to the shared Codex home:

    logs_2.sqlite       -> <shared-home>\logs_2.sqlite
    memories_1.sqlite   -> <shared-home>\memories_1.sqlite
    queue_1.sqlite      -> <shared-home>\queue_1.sqlite
    state_5.sqlite      -> <shared-home>\state_5.sqlite
    

    In this occurrence, the divergence was specifically:

    base .sqlite file: symbolic link
    -wal sidecar:       regular shadow-local file
    -shm sidecar:       regular shadow-local file
    

    This is slightly different from the original #5817 reproduction, where the shadow base database itself is described as becoming a regular file.

    Recurrence evidence

    Retained local logs show repeated failures on at least two consecutive days:

    • August 19: repeated provider-creation failures on memories_1.sqlite-shm
    • August 20: repeated provider-creation failures on logs_2.sqlite-shm
    • An earlier retained error also references queue_1.sqlite

    The problem has returned after multiple manual recoveries and subsequent updates or restarts. It is not an isolated stale file left by a single interrupted process.

    Different SQLite families become affected at different times. Some sidecars already existed from the previous day, while others were created or rewritten during the next T3/Codex run.

    Impact

    • The complete secondary Codex provider is marked disabled.
    • Provider creation fails before the account can be used.
    • The provider remains unavailable across refreshes and restarts.
    • The first conflicting filename changes between occurrences.
    • Removing only the filename shown in the UI is insufficient because additional regular sidecars may remain.
    • The default provider and shared Codex home continue to work.
    • The secondary account authentication file remains valid.
    • The failure occurs during shadow-home materialization, before normal provider use.

    Recovery confirmation

    For diagnostic purposes, I backed up and moved all eight regular *.sqlite-wal and *.sqlite-shm files out of the shadow home. I did not modify the private auth.json or models_cache.json files.

    I then disabled and re-enabled only the affected provider instance.

    After rematerialization:

    materializeCodexShadowHome: Success
    provider enabled: true
    provider status: ready
    authentication status: authenticated
    SQLite sidecar symlinks: 10
    regular conflicting sidecars: 0
    

    This immediately restored the provider, confirming that the regular SQLite sidecars directly caused provider creation to fail.

    However, the same class of conflict has returned after later updates or runs, so this procedure is only a temporary recovery.

    Relation to PR #7201

    PR #7201 appears to target the same underlying class of failure:

    #7201

    Its stated goal—keeping newly created Codex SQLite state shadow-local and treating a base database plus its WAL/SHM sidecars as one family—matches the recurring behavior observed here.

    However, my captured state includes a sidecar-only divergence: the shadow base .sqlite entry remained a symbolic link while its -wal and -shm entries were regular files.

    The PR description and diff appear to select shadow-local ownership when the shadow base database is already a regular file. I cannot confirm whether it also covers the observed case where only the WAL/SHM members have become regular files.

    This Windows reproduction may therefore be useful for validating #7201 against both forms:

    1. A regular shadow base database plus its sidecars.
    2. A symlinked shadow base database with regular shadow WAL/SHM sidecars.

    I am not proposing a separate implementation here. I am documenting the additional observed state so that the existing fix can be evaluated against the complete failure mode.

  3. tgangso commented on Oct 7, 2026

    @tgangso

    Adding updated Windows findings here as requested when #17005 was closed as a duplicate.

    Environment

    • Windows, T3 Code desktop
    • T3 Code: 0.0.46-nightly.20261007.2787
    • Codex CLI: 0.161.0
    • Multiple Codex provider instances sharing a Codex home, with separate shadow homes for authentication

    Observed failure

    Three Codex providers failed to initialize with this sanitized error:

    Driver 'codex' failed to create instance:
    Cannot create Codex shadow home entry 'memories_1.sqlite-shm'
    because '<shadow-home>/memories_1.sqlite-shm'
    already exists and is not a symlink.
    

    Inspection confirmed the sidecar-only state already described in this issue: base SQLite databases were symlinks to the shared home, while some -shm and -wal files were regular files in the shadow homes. Some WAL files were nonempty.

    In the installed build, materializeCodexShadowHome attempts to link shared-home entries, and ensureSymlink rejects these existing regular files. The initial creation of the conflicting files was not captured; no fresh-install reproduction or repository regression test was run.

    Workaround and verification

    The conflicting files were backed up and moved out of the shadow homes, allowing T3 to recreate the links. Each affected provider's launch arguments were also configured to set sqlite_home explicitly to the shared Codex directory.

    After T3 reloaded the settings, all three provider connection probes succeeded. No regular SQLite sidecar conflicts remained at that check. Long-term recurrence has not yet been tested.

    This confirms the failure on the versions above and records the observed recovery. It is not a recommendation to delete SQLite sidecars or a claim that a permanent upstream fix has been verified. Related shared-directory approach: #9229.

  4. dustenhubbard commented on Oct 8, 2026

    @dustenhubbard

    Another entry for this issue, from Codex 0.161.0: cloud-config-bundle-cache.json. This one should not be linked, either.

    Environment

    • T3 Code: 0.0.46-nightly.20261007.2774 (desktop)
    • Codex CLI: 0.161.0
    • macOS 27.0.1, Apple Silicon
    • Two Codex accounts: the default ~/.codex, and a second instance with shadowHomePath set to ~/.codex-t3/work

    What happened

    After a T3 restart, the second instance showed only "Disabled" in Settings → Providers, with its toggle still on. Its status cache had stopped updating at the restart. The settings were unchanged.

    cloud-config-bundle-cache.json existed as a regular file in both the shared home and the shadow home. In the installed build, PRIVATE_ENTRY_NAMES is auth.json and models_cache.json, so materializeCodexShadowHome tries to link this file and ensureSymlink rejects the regular copy. I did not find the error text in the server trace. I am going by the file state and the code.

    Why linking is the wrong fix for this file

    The file is per-account. It holds chatgpt_user_id, account_id, and the workspace's signed enterprise config_toml and requirements_toml, with a one-hour expiry. If it were shared, one account could pick up another account's workspace policy. It belongs with auth.json and models_cache.json.

    Linking also does not last. I deleted the shadow copy and turned the provider off and on. The provider came back, but within seconds there was a regular file in the shadow home again, and the shared copy had not changed. So the next restart would disable the provider again.

    Workaround

    A small job deletes the shadow copy every two minutes, so a restart does not find one. Codex fetches the bundle again when it needs it. This only matters until the file is treated as private.

    Suggested fix

    Add cloud-config-bundle-cache.json to PRIVATE_ENTRY_NAMES. This is a narrower case than the fix this issue asks for. The general fix should also not default to linking: a new per-account file is safer kept local than shared.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions