Skip to content

[Bug]: t3-code MCP credential expires when a turn waits on the user for more than 24 hours #14076

Description

@robertnisipeanu

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Start a thread with a provider that gets the t3-code MCP server (seen with Claude).
  2. Have the agent start a turn that asks the user a question (for example with AskUserQuestion) and leave it unanswered for more than 24 hours. Keep the T3 server running the whole time.
  3. Answer the question. The same turn continues.
  4. Have the agent call any t3-code tool, for example link_pull_request.

Observed timeline: the question was asked at 2026-09-26 10:12 UTC, answered at 2026-09-28 06:19 UTC (44 hours later), and the link_pull_request call at 06:39 UTC failed. The server had been running without a restart since 2026-09-25.

Expected behavior

A provider session keeps a valid t3-code MCP credential for as long as the session is alive, however long a turn runs or waits on the user.

Actual behavior

Every t3-code tool call from that session is rejected:

MCP server "t3-code" rejected the Authorization header in its config (update it, then run /mcp to reconnect)

The session cannot recover. The provider process keeps the old bearer in its MCP config, so reconnecting with /mcp presents the same rejected token.

Cause, in apps/server/src/mcp/McpSessionRegistry.ts:

  • Credentials live in memory, and any record whose lastAliveAt is older than DEFAULT_LIVENESS_WINDOW_MS (24 hours) is pruned.
  • lastAliveAt is refreshed only by MCP traffic (resolve) and by touchActiveMcpThread, which ProviderService calls from sendTurn and compactThread. Answering a pending user-input request or an approval does not refresh it, and nothing refreshes it while a turn is in progress.
  • resolve, issue and touch all prune before they refresh, so once 24 hours have passed a later touch cannot bring the record back.

Related, but separate:

Impact

Major degradation or frequent failure

Version or commit

0.0.42; the same code is on main @ d15210c

Environment

Linux, Claude provider

Workaround

Let the turn finish, then restart the T3 server. The next message resumes the session with a fresh credential.

Activity

  1. juliusmarminge commented on Sep 28, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on main at d15210cd3d (same commit cited in the report). The diagnosis matches the code.

    The bearer is minted once in prepareMcpSession when the provider session starts (ProviderService.startSession, and session recovery only when no live process is adopted). It is not rotated on later turns. touchActiveMcpThread runs only from sendTurn and compactThread. respondToUserInput and respondToRequest do not touch it, and nothing refreshes liveness while a turn is blocked.

    pruneDead runs before the refresh in resolve, issue, and touch. After DEFAULT_LIVENESS_WINDOW_MS (24h) with no MCP traffic and no new sendTurn/compactThread, the record is gone. A later answer cannot bring that bearer back, and the provider process keeps presenting it, which is why /mcp reconnect does not help. The 401 is invalid_mcp_credential; the server warning is rejected MCP request with an unusable credential with reason unknown_or_expired_token.

    This is the same gap #4659 left open. That change refreshes liveness at the start of each turn, so idle time between turns is covered. A single turn that sits on a user question or an approval for more than 24 hours is not. The same thing happens for any in-progress turn that goes 24 hours without a t3-code call. The comment on DEFAULT_LIVENESS_WINDOW_MS says the window only bounds sessions that died without stopSession/stopAll. The implementation never checks that the session is still running.

    Touching on the answer is not enough for this report: the answer arrived at 44 hours, after the prune. Liveness has to be kept during the wait. The 24h bound should stay, because /mcp is outside environment auth and this bearer is the only guard. A still-running session (in a turn, or blocked on user input or approval) should count as alive; a session that exits uncleanly should still expire. The existing registry test only covers touches that land inside the window.

    The restart workaround holds: recovery adopts a still-running process without issuing a new credential, so the next message only gets a fresh bearer after the provider process is gone. Stopping that session and sending a new message does the same thing without restarting the whole server. Sending another message while that process is still up does not.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 28, 2026
  3. added 2 commits that reference this issue on Sep 28, 2026
    4ba7ba2
    9eaf08d
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions