Skip to content

Claude thread permanently fails with "No conversation found with session ID" after fresh-session fallback resumes a newly minted id #15103

Description

@areidyOTH

What happened

A long-running Claude thread (created via the T3 MCP t3_thread_* tools and driven by an orchestrator thread) suddenly started failing every turn with:

provider error. No conversation found with session ID: 13aeed5e-6b0e-4d18-bdf4-dca34e6572b0

That session ID never existed. The thread's real Claude session (9b7e437e-…) had been running for ~4 hours and its transcript is still intact in ~/.claude/projects/<worktree>/9b7e437e-….jsonl. The thread is now permanently stuck: every new message fails the same way.

Diagnosis

Sequence from server.trace.ndjson, the per-thread provider event log, and statev2.sqlite (times UTC):

  1. 04:20:55 – Claude query opened with sessionId: 9b7e437e-…; ~6,700 events of normal work follow, all with session_id: 9b7e437e-…. The agent starts a long-running background shell.
  2. 08:08:26 and 08:08:52 – Two message.dispatch turns (runs 13, 14; t3_thread_send with mode: auto then mode: queue) fail in ClaudeAdapterV2.startTurn with ClaudeBackgroundWorkBlocksQueryReplacementError ("Claude is still running background agents or commands…"). This is the guard in ClaudeAdapterV2.ts openQuery (~L6836–6843).
  3. 08:09:19 – The background task finishes (task_notification); T3 dispatches a provider-continuation turn (run 15).
  4. 08:09:24 – ProviderTurnStartService logs Provider resume failed; attempting a fresh native session with reason: "uncertain_history_delivery" (ProviderTurnStartService.ts ~L658–704). It never calls resumeThread; it fails fast because some context handoff to this provider thread has delivery.status === "pending" for the current native id, most likely left behind by the refused runs 13/14 (I could not confirm this from the DB). It then calls ensureThread with existingProviderThread: { ...providerThread, nativeThreadRef: null }, which mints a fresh native id 13aeed5e-….
  5. 08:09:27.871 – The live query is closed. 08:09:27.882 – A new query is opened with resume: "13aeed5e-…" rather than sessionId: "13aeed5e-…". The CLI answers immediately with result/error_during_execution, errors: ["No conversation found with session ID: 13aeed5e-…"]. Runs 16 and 17 repeat this.

Root cause of step 5, in ClaudeAdapterV2.ts openQuery (~L6861–6878 at 8ed276c2; L6887–6889 on current main):

const hasPersistedProviderTurn = turnInput.providerTurnOrdinal > 1;
const shouldResume =
  resumeSessionAt !== undefined || openedWithResume || hasPersistedProviderTurn;

providerTurnOrdinal > 1 is used as proof that the native session exists. But after the fresh-session fallback in ProviderTurnStartService, the provider thread keeps its turn history while getting a brand-new native id. So the adapter passes resume for a session the CLI has never seen. The fallback can never work for any thread with more than one provider turn.

Resulting persistent state: orchestration_v2_projection_provider_threads.payload_json.nativeThreadRef = { driver: "claudeAgent", nativeId: "13aeed5e-…", strength: "strong" }, so all later turns resume the non-existent id. Two context_handoffs rows (full_thread_summary for run 15, and delta_since_target_last_seen) were written with delivery.nativeThreadId: 13aeed5e-…, but the history never reached Claude.

Possible directions (not tested):

  • Have the fallback signal "fresh session" to the adapter, so openQuery uses sessionId, not resume, for a newly minted native id regardless of providerTurnOrdinal.
  • Don't persist the new nativeThreadRef as strong until the CLI confirms the session.
  • Turns refused by ClaudeBackgroundWorkBlocksQueryReplacementError probably shouldn't leave a pending handoff delivery that later forces the fresh-session path.

Steps to reproduce

Not reproduced deterministically; reconstructed from logs:

  1. Start a Claude (claudeAgent) thread and run several turns, so providerTurnOrdinal > 1.
  2. Have the agent start a long-running background shell (e.g. a until …; do sleep; done waiter run in the background).
  3. While it runs, send another message to the thread (UI or t3_thread_send). It is refused with "Claude is still running background agents or commands…". Sending a second one in queue mode is also refused.
  4. Let the background task finish, so T3 starts a provider-continuation turn.
  5. Observe Provider resume failed; attempting a fresh native session (uncertain_history_delivery) in the trace, followed by query.open with resume: <new uuid> and No conversation found with session ID: <new uuid>. Every later turn on the thread fails the same way.

A more direct check of the core bug: force the fallback path in ProviderTurnStartService (e.g. make resumeThread fail) on a Claude thread with ≥2 provider turns, and observe that the replacement query is opened with resume for the freshly minted id.

Version

Desktop AppImage 0.0.46-nightly.20261003.2610 (commit 8ed276c). Same code is present on main as of 2026-10-03.

Environment

Linux x64 (kernel 7.0.0), Node v26.8.2, Claude Code CLI 2.1.288, desktop app (server.mode: desktop, 127.0.0.1:3773)

Evidence

# server.trace.ndjson — runs 13/14 refused
08:08:26.922 ClaudeAdapterV2.startTurn Failure
  ProviderAdapterTurnStartError: Failed to start run run:thread:mcp%3Ac5639143-…:ordinal:13 …
  [cause]: ClaudeBackgroundWorkBlocksQueryReplacementError: Claude is still running background agents or commands, and this model or setting change would end them. …
08:08:55.278 ClaudeAdapterV2.startTurn Failure   (ordinal:14, same cause)

# run 15 — fallback
08:09:24.888 orchestrationV2.providerTurnStart.start
  "Provider resume failed; attempting a fresh native session"
  { driver: "claudeAgent", runId: "run:thread:mcp%3Ac5639143-…:ordinal:15",
    reason: "uncertain_history_delivery", errorTag: "ProviderAdapterTurnStartError" }

# provider event log (events.mcp-c5639143-….log)
04:20:55.366 outgoing query.open  { sessionId: "9b7e437e-d8ba-…", model: "claude-opus-5-5[1m]", ... }
   … 6746 events with session_id 9b7e437e-… …
08:09:19.648 incoming system/task_notification (background task finished)
08:09:27.871 outgoing query.close
08:09:27.882 outgoing query.open  { resume: "13aeed5e-6b0e-4d18-bdf4-dca34e6572b0", ... }
08:09:29.084 incoming result { subtype: "error_during_execution", num_turns: 0,
  errors: ["No conversation found with session ID: 13aeed5e-6b0e-4d18-bdf4-dca34e6572b0"] }
08:09:30.862 / 08:09:33.403  same query.open with resume 13aeed5e… → same error (runs 16, 17)

# statev2.sqlite (read-only)
orchestration_v2_projection_provider_threads.nativeThreadRef
  = {"driver":"claudeAgent","nativeId":"13aeed5e-6b0e-4d18-bdf4-dca34e6572b0","strength":"strong"}
# ~/.claude/projects/<worktree>/ contains 9b7e437e-….jsonl; no 13aeed5e-* file exists anywhere

Related issues

#2336 (closed): a thread becomes permanently unusable because T3 resumes a Claude session id that has no transcript (V1 resume_cursor_json, CLI killed before first write). Same symptom and the same "permanently stuck" outcome, but a different code path. This one is the V2 ProviderTurnStartService fresh-session fallback combined with ClaudeAdapterV2 deciding resume from providerTurnOrdinal > 1.

Fix applied or workaround

Nothing was changed on the machine. Workaround: the original conversation is intact, so it can be continued outside T3 with claude --resume <original session id> in the thread's worktree, or the work can be handed to a new T3 thread. The stuck thread itself cannot recover without editing statev2.sqlite.

Filed by

Claude Code (claude-opus-5-5) via t3 triage

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions