Skip to content

Daemon does not reliably pick up a new credential after kcap login (live hub pins handshake identity; profile/tenant change never heals) #511

Description

@alexeyzimarev

Sibling of #509 (which makes a hook 401 tell the user to run kcap login). Once they do, a running daemon does not reliably pick the new credential up.

What is already fine

Worth stating so a fix does not touch it: every daemon HTTP path resolves a token per operation through TokenStore, which has no in-process cache — ServerConnection.ProcessEventQueueAsync (:1247), EvalRunner (:50, :99, :144), SpoolDrainLoop (:52), AgentOrchestrator (:262, :2824). Those see a new token on the next call. TokenRefreshLoop likewise drives the shared, cross-process-locked TokenStore.RefreshIfExpiringAsync, so proactive refresh is not the problem either.

A hub connection that fails to authenticate also self-heals: ConnectWithRetryAsync retries forever (backoff capped at 30s) and each StartHubAsync re-invokes AccessTokenProvider (ServerConnection.cs:192), which re-reads the store.

Problem 1 — a live hub connection pins its handshake identity

SignalR invokes AccessTokenProvider only at negotiate. kcap-server sets no CloseOnAuthenticationExpiration on the hub, so the ClaimsPrincipal captured at handshake is pinned for the life of the socket. A daemon whose token expires — or is replaced by kcap login — keeps operating as the old identity until something independently drops the transport.

This usually presents as silent staleness rather than an error, which makes it easy to miss: work attributed to the previous identity, or a tenant switch that appears not to take effect. ForceReconnectAsync (:785) already exists and is driven by DaemonHeartbeatLoop, so there is a natural hook — nothing currently triggers it on a credential change.

Problem 2 — a profile or tenant change never heals, and never explains itself

TokenStore.GetValidTokensForServerAsync resolves the active profile, then withholds the token entirely unless it is bound to the requested server:

if (!BoundToTarget(snapshot, targetBaseUrl)) {
    return new(null, AuthStatus.WrongServer, snapshot.ServerUrl, profile);
}

The daemon always asks for its own configured _config.ServerUrl. So if kcap login lands on a different tenant (discovery lets you pick one), or the user runs kcap use to switch profiles, a daemon configured for the previous server gets WrongServer forever. No amount of retrying fixes a binding mismatch, but ConnectWithRetryAsync retries anyway, every 30s, in perpetuity — the daemon needs a restart or a reconfigure and nothing says so.

Diagnosability is the other half: the hub's AccessTokenProvider just returns null on WrongServer / Expired, discarding the reason. The daemon then logs a generic connection failure instead of "this profile's token was issued by " or "run kcap login" — the same unexplained-401 complaint as #509, one layer down.

Not yet reproduced

The originally-reported symptom ("after kcap login the daemon still used the old token") was an inference, not a confirmed observation, so the exact trigger is unverified. Both problems above are read off the code rather than from a repro. A third hypothesis worth testing while reproducing: a supervised daemon (launchd/systemd) runs with its own environment, so if HOME differs from the interactive shell's it would read a different token store entirely — that one is speculation, not verified.

Suggested first step: reproduce each path deliberately (expire a token under a live hub; kcap login into a second tenant with a daemon running; run the same under a supervised daemon) before choosing between "force a re-negotiate on credential change" and "detect the binding mismatch and report it instead of retrying forever". They may well need different fixes.

Activity

  1. linear-code commented on Aug 10, 2026

    @linear-code
  2. alexeyzimarev commented on Aug 10, 2026

    @alexeyzimarev
    MemberAuthor

    Narrowing this: #516 (AI-1843) investigated the same symptom from the MCP side and disproved one of the hypotheses in the issue body, so nobody should re-investigate it.

    From that commit message:

    the issue's suspected "caches auth at process start" bug does not exist; the recovery path was working as designed on every surface … every kcap process re-reads the same token store on 401

    That matches what this issue already recorded from the code (every daemon HTTP path resolves a token per operation through TokenStore, which has no in-process cache), and it adds a real-incident data point: what looked like stale-credential behaviour was a tenant server rotating its ephemeral signing key across pod replacements, 401'ing tokens that kcap status concurrently and correctly reported as valid.

    So the "supervised daemon reads a different token store because HOME differs" hypothesis at the bottom of this issue is now the weakest of the three and should be tested last, if at all.

    What #516 does not address, and what this issue should narrow to:

    1. A live SignalR hub pins its handshake identity. AccessTokenProvider runs only at negotiate and kcap-server sets no CloseOnAuthenticationExpiration, so the socket keeps its original ClaimsPrincipal until something drops the transport. Per-request re-reads don't help here, because there is no per-request auth on an established hub connection. ForceReconnectAsync exists and is already driven by DaemonHeartbeatLoop, so there is a hook — nothing triggers it on a credential change.
    2. A profile or tenant change never heals and never explains itself. GetValidTokensForServerAsync withholds the token as WrongServer unless it is bound to the daemon's configured ServerUrl, so logging into a different tenant leaves ConnectWithRetryAsync retrying every 30s forever. The hub's AccessTokenProvider discards the reason and returns bare null, so the daemon logs a generic connection failure rather than "this profile's token was issued by <other server>".

    Point 2 is now cheap to fix well: #516's AuthRejectionNotice.Classify/Render already produces exactly the right sentence for the WrongServer case, naming both servers and offering kcap login / kcap use.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions