Repository navigation
[Bug]: One transient probe timeout leaves Claude stuck on "Could not verify Claude authentication status" for hours (cause swallowed, no demand-less recheck) #13635
Description
Activity
Triage
Confirmed on current
main(e3e7cc3fc2) and on the reported nightlyv0.0.43-nightly.20260922.2110. One failed Claude capabilities probe is published as an auth warning, written to~/.t3/caches/claudeAgent.json, and left there until some later check finishes. The 22-hour-oldcheckedAtmeans no later check finished. Not a duplicate of #7111, #7230, or #7513.The warning is not cosmetic.
isProviderInstancePickerReadyonly treatsstatus === "ready"as selectable, so Claude drops out of the new-thread picker while threads that are already running keep going. That matches "logged in and working" plus a card that still says it could not verify auth.What the report gets right
probeClaudeCapabilitiesturns every failure intoundefinedand does not log it. Timeout, spawn error, and an SDK error all become the same value:Effect.ensuring( Effect.sync(() => { if (!abort.signal.aborted) abort.abort(); }), ), Effect.result, Effect.map((result) => (Result.isSuccess(result) ? result.success : undefined)), );
checkClaudeProviderStatusthen treats that as an auth warning. The version probe logs a warning; this path does not. Codex and Cursor, which failed in the same minute, keep the timeout in the message (Timed out while checking Codex app-server provider status./timed out while running `agent about`). Claude cannot.const capabilities = resolveCapabilities ? yield* resolveCapabilities(claudeSettings).pipe(Effect.orElseSucceed(() => undefined)) : undefined; // ... message: "Could not verify Claude authentication status from initialization result.",
The banner title is "Claude provider status" (warning), and Settings says "Needs attention".
auth.statusisunknown, notunauthenticated. The sentence still reads like an auth failure.The periodic loop really does refuse to run without a client lease:
const hasProviderStatusDemand = Effect.gen(function* () { // ... backgroundPolicy.shouldRunScopeWork({ type: "provider-status" }), backgroundPolicy.shouldRunScopeWork({ type: "provider-status", instanceId }), // ... ? hasProviderStatusDemand.pipe( Effect.flatMap((shouldRefresh) => shouldRefresh ? refreshSnapshot().pipe(Effect.asVoid) : Effect.void, ), )
A running thread does not create that lease. Web and mobile only send
{ type: "provider-status" }while a client is connected (BASELINE_SCOPESinbackgroundActivityReporter.tsandbackground-activity.ts). On the performance profile a connected client does not have to be focused, and a headless host with no desktop power telemetry isstale, which is not treated as constrained. So a client that stays up should get a new check within the 1-minute interval, and that check always writes a newcheckedAteven when the probe result is unchanged. The cache files staying at2026-09-24T12:59Z–13:00Zmeans those checks did not finish. That is the no-lease path, not a probe that kept failing.The failed snapshot is what a reconnecting client is shown. Disk is only the boot copy; the live process serves the in-memory snapshot. Startup does
forceRefreshonce, with no demand check, so a restart re-probes. A long-runningt3 servedoes not.There is a second, shorter stick. The probe cache TTL is 5 minutes (
CAPABILITIES_PROBE_TTLinClaudeDriver.ts), and a failed probe is a successfulundefined, so it is cached. For those 5 minutes a refresh republishes the warning without spawning Claude. That cannot explain 22 hours. It does explain a window with no new process if something was refreshing. It does not explain a frozencheckedAt.What is off
The probe options are not empty
settingSources. On that nightly and onmainthey areuser/project/local, with hooks disabled, an empty MCP map, andstrictMcpConfig. The external repro was stricter than the real probe and still returned an account, so it still shows the CLI was fine after the stall. It is not the exact call.#7111 is the same probe wiping the slash-command list. #7175 was closed without merging.
mergeProviderSnapshotonly keeps an empty command list for OpenCode, so Claude commands are still replaced. That issue expects the next probe to refill the menu. This one is the auth warning remaining because the next probe never runs. Keep them separate.#7230 / #7232 are the general "a timeout is stored as broken" bug. #7232 is still open, and it explicitly skips this probe ("already degrade without clobbering the snapshot"). They do not. This failure never becomes
ProviderProbeTimeoutError; it is published as a normal warning, so that carry-forward would not apply. #7230 also assumes a later refresh runs. The 22-hour freeze is the demand gate, which that issue does not describe.#7513 is the Codex 10s timeout, and it recovers on the next refresh. #10215 retries a non-timeout capabilities failure without
--settingsand logs the cause. It does not retry timeouts, and it does not touch the demand gate or the cachedundefined. #7691 is the opposite report (showing authenticated when the CLI is logged out).Settings already shows "Checked …" on the provider refresh control, from the newest
checkedAtacross providers. The chat banner does not, and one stale Claude card is hidden if any other provider checked more recently.Workaround
Restart
t3 serve. The new process probes at startup with an empty cache, with or without a client.On current
main, Settings → Providers → refresh sendsrefreshModels: true, which drops the capabilities cache and probes immediately (#13109). That invalidation is not inv0.0.43-nightly.20260922.2110. On that build a manual refresh inside the 5-minute TTL republishes the cachedundefinedand only movescheckedAt. DeletingclaudeAgent.jsonwhile the process is up does nothing; clients are not reading that file.Labels / next
- Type: bug, accepted.
- Labels:
bug,accepted. - Next: a capabilities miss should not replace a ready snapshot with this warning. Log the failure tag (not stderr), and say "probe timed out" when that is the cause. Do not cache
undefinedas a good lookup. Retrying on a timer with no client would fight the idle-work gate; a connected client already retries within one interval once the cached failure is gone.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 25, 2026 - added a commit that references this issue
on Sep 30, 2026
Area
apps/server(provider status / Claude capability probe)Summary
On a headless Linux server (
t3 serve, nightly 0.0.43-nightly.20260922.2110), one transient host stall made every provider check fail at once. Those failed snapshots then stayed in place for 22+ hours. The Claude card and toast kept saying "Could not verify Claude authentication status from initialization result." even though Claude was logged in and working the whole time.Evidence
~/.t3/caches/*.jsonat 2026-09-25T11:02Z:claude auth statusreportsloggedIn: true(claude.ai subscription). Claude threads kept running normally in T3.buildClaudeCapabilitiesProbeQueryOptionsequivalents: emptysettingSources,disableAllHooks, no MCP, cwd = server cwd), the bundled SDK version (@anthropic-ai/claude-agent-sdk@0.3.276) and the service's minimal env.initializationResult()returned the account in 1.3–2.0 s on 6 of 6 runs.providerHealthRefreshIntervalis 1 min (performance profile).Why it sticks (reading the bundled server)
checkClaudeProviderStatuscallsresolveCapabilities(...).pipe(orElseSucceed(() => undefined)), andprobeClaudeCapabilitiesmaps any failure toundefined. The real cause (timeout, spawn error, SDK error) is discarded and never logged, so the card can't say what went wrong.backgroundPolicy.shouldRunScopeWork({ type: "provider-status" })is true, meaning a client lease with provider-status demand exists and the host isn't constrained. If no client asks for provider status, a snapshot that failed transiently is never re-checked. It is also persisted to~/.t3/caches/<provider>.json, so every reconnecting client sees it as current.Expected
checkedAtage in the UI.Related
#7111 / #7175: same failure point, but for the slash-command list. That fix preserves commands, while the auth status still sticks. #7513: a similar probe timeout on the Codex provider.