Skip to content

[Bug] PR panel shows "Could not load pull requests" on the first open after a turn #14113

Description

@simonsteiner

What happened

Clicking a pull request link in a thread opens the PR panel in the right sidebar, but the first open after launching the app often shows Could not load pull requests — All fibers interrupted without error with a Retry button. Retry always loads the PR. Opening PRs later in the same session works.

Diagnosis

Two bugs combine. The server hands a request an interrupted result, and the client stores that interrupt as a permanent error.

1. The client cancels and resends detail right after the panel opens (by design, harmless alone).
The PR panel's detail query refreshes on the pullRequests.subscribeRefreshes signal (makeRefreshOnSignal in packages/client-runtime/src/state/runtime.ts). That stream is SubscriptionRef.changes(pullRequestRefreshes) filtered to revision > 0 (apps/server/src/pullRequest/PullRequestService.ts, subscribeRefreshes). So once the server has completed any agent turn since it started, the stream emits its current value as soon as it is subscribed. When the panel is the first PR reader to mount, it subscribes as it opens. The first value arrives about 40 ms later and restarts the query: the client sends Interrupt for the first detail request and a second identical detail request. This happened on every one of 20+ attempts.

2. The server lets the resent request join the read that is being torn down.
detail reads are coalesced through an effect/Cache (detailCache in PullRequestService.ts). When the only caller is interrupted, Cache's entry await interrupts the lookup fiber once the awaiter count drops to 0. The entry is removed from the map only after that fiber exits (the observer in Cache.get), and stopping the gh processes it is running takes a few ms. A request that reaches Cache.get during that window finds the still-present entry, joins the dying fiber, and exits with a bare Interrupt. The RPC server sends that back as the request's result.

3. The client turns that interrupt into a sticky error.
The atom stores Failure(Interrupt) and never retries. useEnvironmentQuery (apps/web/src/state/query.ts) formats it as the error string ("All fibers interrupted without error"), writableQueryFamily (packages/client-runtime/src/state/pullRequests.ts) passes it through, and PullRequestDetailPanel renders PullRequestsUnavailableState with Retry. Commands already ignore interrupt-only failures (isAtomCommandInterrupted); queries don't.

Whether step 2 hits depends on timing: the resent request must reach Cache.get inside the few ms while gh is stopping. On my machine it does often: the desktop app on Windows with the WSL backend, a real database, and gh against github.com. In an isolated dev server with an empty database it never hit in 19 attempts, although step 1 happened every time. Widening the window (below) makes it deterministic.

Possible fixes, for whoever picks this up:

  • Server: don't hand an entry whose lookup is being interrupted to a new caller. For example, remove it from the map when the last awaiter leaves, or retry Cache.get when the joined exit is interrupt-only and the caller itself wasn't interrupted. The same Cache pattern backs activity, preview, list and diff reads.
  • Client: treat an interrupt-only Failure as still pending, or retry it, instead of an error.
  • Optionally, don't refresh on the refresh stream's initial value, only on changes after subscribing.

Steps to reproduce

In the app (intermittent, depends on timing):

  1. Start T3 Code and let an agent turn complete in any thread (the refresh revision is 0 until then).
  2. Without having opened any PR panel in this session, click a GitHub PR link in a thread whose project matches the repository. The PR must not have been read by the server in the last ~15 s.
  3. The PR panel sometimes shows "Could not load pull requests — All fibers interrupted without error". Retry loads it.

Deterministic, in an isolated dev server (vp run dev with a fresh --home-dir):

  1. Widen the teardown window by adding Effect.onInterrupt(() => Effect.sleep("250 millis")) to the detailUncached(...) pipe in the detailCache lookup in PullRequestService.ts. This only simulates a slower gh shutdown; it does not change which code path runs.
  2. Add a project inside a GitHub repo and let an agent reply with a PR link from that repo.
  3. In a fresh browser session, click the link. The panel fails on 4 of 4 attempts. The attached video shows this setup.

Deterministic, as unit tests (no timing involved):

// apps/server: effect/Cache hands a dying lookup to the next caller
import { assert, it } from "@effect/vitest";
import * as Cache from "effect/Cache";
import * as Deferred from "effect/Deferred";
import * as Effect from "effect/Effect";
import * as Exit from "effect/Exit";
import * as Fiber from "effect/Fiber";

it.effect("a read arriving while an abandoned lookup shuts down gets its interrupt", () =>
  Effect.gen(function* () {
    const started = yield* Deferred.make<void>();
    const stopped = yield* Deferred.make<void>();
    let lookups = 0;
    const cache = yield* Cache.make({
      capacity: 10,
      lookup: (_: string) =>
        ++lookups === 1
          ? // Like `gh`, the first lookup takes a moment to stop once interrupted.
            Effect.yieldNow.pipe(
              Effect.andThen(Deferred.succeed(started, undefined)),
              Effect.andThen(Effect.never),
              Effect.onInterrupt(() => Deferred.await(stopped)),
            )
          : Effect.succeed("fresh"),
    });
    const first = yield* Cache.get(cache, "detail").pipe(Effect.forkChild);
    yield* Deferred.await(started);
    yield* Fiber.interrupt(first).pipe(Effect.forkChild({ startImmediately: true }));
    const second = yield* Cache.get(cache, "detail").pipe(
      Effect.forkChild({ startImmediately: true }),
    );
    yield* Deferred.succeed(stopped, undefined);
    // Fails: Failure(Cause([Interrupt])), lookups === 1
    assert.deepStrictEqual(yield* Fiber.await(second), Exit.succeed("fresh"));
  }),
);
// packages/client-runtime: an interrupt reply becomes a permanent failure with no retry.
// Uses a harness like makeTestRuntime in src/state/pullRequests.test.ts.
it.effect("detail recovers when the server answers a read with an interrupt", () =>
  Effect.scoped(
    Effect.gen(function* () {
      let calls = 0;
      const client = {
        [WS_METHODS.pullRequestsSubscribeRefreshes]: () => Stream.never,
        [WS_METHODS.pullRequestsDetail]: () =>
          ++calls === 1 ? Effect.interrupt : Effect.succeed({ ...reference, title: "ok" }),
      } as unknown as WsRpcProtocolClient;
      const { atoms, registry } = yield* setup(client);
      const detail = atoms.detail({ environmentId: TARGET.environmentId, input: reference });
      // ...mount and wait for the atom to stop waiting...
      // Fails: final is Failure(Cause([Interrupt])) after calls === 1; nothing retries.
      expect(AsyncResult.isSuccess(registry.get(detail))).toBe(true);
    }),
  ),
);

Version

0.0.43-nightly.20260928.2375 (desktop, WSL backend); also reproduced on main at d15210c

Environment

Windows 11 (10.0.26200) desktop app with the WSL2 backend (kernel 6.18.33.2), gh 2.101.0, effect 4.0.0-rc.115

Evidence

# Real session, server trace (ws.rpc.pullRequests.*), first PR opened after launch:
16:25:52.261 ws.rpc.pullRequests.subscribeRefreshes  (stream opened)
16:25:52.263 ws.rpc.pullRequests.detail  81 ms  Interrupted  (client abort; its gh processes killed)
16:25:52.340 ws.rpc.pullRequests.detail   3 ms  Interrupted  "interrupted by fiber #14087"
             # #14087 is the detailCache lookup fiber created by the first request
             # (its gh children are #14088..#14094); the second request only ran canonicalRef
16:25:52.426 ws.rpc.pullRequests.activity 962 ms Success
16:25:54.503 ws.rpc.pullRequests.detail 681 ms  Success  (user pressed Retry)

# Isolated dev repro, client WebSocket frames (ms since page start):
 9653 click PR link
 9804 → Request 16 pullRequests.subscribeRefreshes
 9804 → Request 17 pullRequests.detail {"repository":"pingdotgg/t3code","number":13190,...}
 9840 → Interrupt 17
 9840 → Request 19 pullRequests.detail {"repository":"pingdotgg/t3code","number":13190,...}
10091 ← Exit 17 Failure [Interrupt fiberId 4921]
10163 ← Exit 19 Failure [Interrupt fiberId 5054]   # the dying detailCache lookup
10169 panel: "Could not load pull requests / All fibers interrupted without error"
13684 click Retry → loads

Video (isolated dev server, widened window as described above): first click on the link → error panel → Retry → loads.

pr-panel-interrupted-first-open.mp4

Related issues

None covers this interrupt race.

Fix applied or workaround

None applied. Workaround: press Retry.

Filed by

Claude Opus 5.5 (Claude Code, running in T3 Code), investigated with t3 triage's playbook

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions