Skip to content

[Bug]: PullRequestSyncReactor's gh calls fail with an opaque error — VcsProcessExitError discards stderr, so the real cause can't be diagnosed #12224

Description

@Artic0din

Edit: the original diagnosis below ("not a git repository" because cwd defaults to $HOME) is retracted — see the correction comment below. cwd is confirmed (by reading the actual source) to be the real project.workspaceRoot, not a wrong default, and gh pr view --repo <owner>/<repo> (what T3 actually runs) does not require cwd to be a git repo at all. The real cause of the 25 failed gh calls in the evidence below is unknown, because VcsProcessExitError discards the actual stderr text (see the comment for detail — same defect class as #4380, different type/file). Leaving the raw evidence (timing, cwd, frequency) as filed since it's still accurate; only the diagnosis is wrong.

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

This is a background-reactor bug, not directly user-triggered, so there isn't a UI click-path repro. It shows up passively:

  1. Run T3 Code server/desktop with at least one project open that has a GitHub remote and an active/linked thread (PR-related state).
  2. Leave the app running normally (this is not idle-only — it also happened while a thread was actively being worked in the same window).
  3. Tail apps/server trace logs (~/.t3/userdata/logs/server.trace.ndjson*) and filter for "name":"PullRequestSyncReactor.syncGroup" (or PullRequestReadCache.get, GitHubCli.executeRaw) with "_tag":"Failure".
  4. Every occurrence has gh (/Users/<user>) in the error chain. That cwd is the real project.workspaceRoot (confirmed by reading PullRequestService.summaryUncached → requireProject → getChangeRequestSummary), not a bug in cwd resolution — see the edit note above.

VcsProcessExitError only records stderrLength/stderrTruncated, not the stderr text, so the actual reason gh exited 1 is not recoverable from this trace. It was classified as the generic "command-failed" bucket by classifyNonZeroExit (not auth, not rate-limited, not not-found), which only narrows it down, doesn't identify it.

Expected behavior

VcsProcessExitError (and its gh/glab/az callers through GitHubCli/VcsProcess) should carry enough of the real stderr to diagnose a command-failed exit, instead of only stderrLength. See #4380 for the same defect in GitCommandError (plain git commands) — this is asking for the same fix applied to VcsProcessExitError.

Actual behavior

25 distinct gh invocations failed with exit 1 over a 3.5h window (below), roughly 60s apart, matching the reactor cadence already described in #11220. The failure is silent from the user's perspective — nothing surfaces in the UI, only the trace logs — so PR data for the affected thread(s) may simply never populate, with no way (yet) to tell why.

Observed over a 3.5h window: 25 distinct failed gh invocations (PullRequestSyncReactor.syncGroup → PullRequestService.persistedRead → PullRequestReadCache.get → GitHubCli.executeRaw), each one logged at up to 15+ nested effect-span layers (382 raw "_tag":"Failure" log lines total for these 25 logical failures), spaced roughly 60s apart, all with gh (/Users/ryan) in the cause chain.

Impact

Major degradation or frequent failure

Version or commit

T3 Code Nightly 0.0.43-nightly.20260917.1837

Environment

macOS 27.0 (build 26A428), Bun 1.4.2, Node v26.8.2, gh 2.83.1, GitHub provider = github.com with personal OAuth token (not a GitHub App installation token — rules out #11247).

Logs or stack traces

PullRequestOperationError: Pull request operation summary failed: GitHub CLI command failed.
    at PullRequestReadCache.get (Nightly)
    at PullRequestService.persistedRead (Nightly)
    at PullRequestSyncReactor.syncGroup (Nightly)
    at PullRequestSyncReactor.sweep (Nightly) {
  [cause]: PullRequestProviderError: github failed in getChangeRequestSummary: GitHub CLI command failed. {
    [cause]: GitHubCliCommandError: GitHub CLI failed in execute: GitHub CLI command failed.
        at fromVcsError (bin.mjs:86949:9) {
      [cause]: VcsProcessExitError: VCS process failed in GitHubCli.execute: gh (/Users/ryan) exited with 1 - Process exited with a non-zero status.
          at VcsProcessExitError.fromProcessExit (bin.mjs:31123:10)
    }
  }
}

Note: gh pr view --repo <owner>/<repo> --json ... (what T3 actually runs here) does not require cwd to be a git repo — verified directly (cd /Users/ryan && gh pr view 11225 --repo pingdotgg/t3code --json number,title succeeds, exit 0). An earlier manual repro using gh pr list (no --repo) is not representative of this call shape and has been removed.

Workaround

None found — the real cause is unknown pending a stderr excerpt on VcsProcessExitError.

Activity

  1. juliusmarminge commented on Sep 17, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on current main (0150c6a53). Real server GitHub CLI bug on packaged desktop. Not a duplicate of #11220 (cadence) or #11247 (App installation token). The gh ($HOME) string is real; the command that produces it is not the pr view the reactor thinks it is running.

    What happens

    Packaged desktop starts the backend with cwd = the user’s home directory:

        backendCwd: input.isPackaged ? homeDirectory : appRoot,

    PullRequestSyncReactor.syncGroup then calls pullRequests.summary with recoverTransientFailure: false every time an open, unsettled link is due (the 1-minute sweep from Schedule.spaced("1 minute")):

        const syncGroup = Effect.fn("PullRequestSyncReactor.syncGroup")(function* (
          key: string,
          entries: ReadonlyArray<LinkEntry>,
        ) {
          // ...
          const summary = yield* pullRequests.summary(ref, { recoverTransientFailure: false });

    summaryUncached does pass the project root into the provider:

      const summaryUncached: PullRequestService["Service"]["summary"] = (input) =>
        requireProject(input).pipe(
          Effect.flatMap((project) => {
            const providerInput = {
              cwd: project.project.workspaceRoot,
              repository: project.repository,
              host: project.host,
              number: input.number,
            };
            const read =
              project.api.getChangeRequestSummary === undefined
                ? project.api.getChangeRequest(providerInput)
                : project.api.getChangeRequestSummary(providerInput);

    and getPullRequestSummary runs gh pr view <n> --repo <host>/<repo>, which does not need a local git checkout:

        getPullRequestSummary: (input) =>
          github
            .execute({
              cwd: input.cwd,
              args: [
                "pr",
                "view",
                String(input.number),
                ...repositoryArgs(input),
                "--json",
                PULL_REQUEST_DETAIL_JSON_FIELDS,
              ],
            })

    That execute never reaches pr view until a quota probe succeeds. Since #11888 the probe is hardcoded to the server cwd:

      const quota = yield* Cache.makeWith(
        (key: string) => {
          const host = key.split("\0")[0]!;
          return executeRaw({
            cwd: globalThis.process.cwd(),
            args: [
              "api",
              "rate_limit",
              "--hostname",
              host,
              "--jq",
              ".resources.graphql | {data:{rateLimit:{cost:1,limit:.limit,remaining:.remaining,resetAt:(.reset|todateiso8601)}}}",
            ],
          }).pipe(
            // ...
          );
        },
        {
          capacity: 32,
          timeToLive: (exit) => (Exit.isSuccess(exit) ? Duration.seconds(30) : Duration.zero),
        },
      );
      // pr list | pr view | repo view then:
      yield* Cache.get(quota, `${host}\0${credential?.credentialFingerprint ?? ""}`);
      return yield* executeRaw(input);

    VcsProcessExitError prints gh (<cwd>). That is why every failure shows gh (/Users/ryan) — it is process.cwd(), which packaged desktop set to $HOME. A failed probe is not cached (timeToLive is 0 on failure), so every sweep retries it and pr view never runs. recoverTransientFailure: false also drops any last-good summary.

    The pasted stack does not include gh stderr. VcsProcessExitError.detail is only Process exited with a non-zero status. The cd ~ && gh pr list --json number repro is a different command (pr list with no --repo). T3’s actual read is pr view --repo. The command that matches cwd=$HOME on every failure is the quota gh api rate_limit.

    ThreadPullRequestReactor is the contrast: it picks worktree or project.workspaceRoot before calling git/gh. The sync reactor never sets cwd itself; GitHubCli.execute overwrites it for the probe.

    Nightly 0.0.43-nightly.20260917.1837 includes #11888 (merged 2026-09-15). Auth is fine; this is not #11247.

    Not a duplicate

    No existing issue for the quota probe using process.cwd().

    Likely code

    Fix direction

    1. Pass input.cwd (already the project workspace root) into the quota executeRaw, not process.cwd().
    2. Do not fail the real pr view when the probe fails. Treat a dead probe as “unknown budget” and still run the read. A host-level rate_limit should not take the PR sync path down.
    3. Test: packaged-style process.cwd() === $HOME (not a git repo) + execute({ cwd: "/repo", args: ["pr", "view", …] }) must still call pr view with /repo. Cover probe-failure-must-not-block-read.
    4. If a follow-up log still has stderr not a git repository on pr view --repo itself, then also check workspaceRoot; that is a second bug, not what this stack shows.

    Workaround

    None in the UI. Unpacked / npx t3 serve from a repo may hide it because process.cwd() is then a git root. No user setting points the probe at the project.

    Accepting as a server GitHub CLI bug. The reactor cadence in #11220 makes the failed probe loud; it is not the cause.

  2. Artic0din commented on Sep 17, 2026

    @Artic0din
    Author

    Correction

    I traced the exact code path (PullRequestSyncReactor.syncGroup → PullRequestService.summary → requireProject → getChangeRequestSummary → GitHubPullRequestCli.getPullRequestSummary) in the actual source rather than guessing from the packaged trace alone.

    getPullRequestSummary always passes --repo <host>/<repository> explicitly:

    github.execute({
      cwd: input.cwd,
      args: ["pr", "view", String(input.number), ...repositoryArgs(input), "--json", ...],
    })

    gh pr view --repo owner/repo does not require cwd to be inside a git repository. I confirmed this directly:

    $ cd /Users/ryan && gh pr view 11225 --repo pingdotgg/t3code --json number,title
    {"number":11225,"title":"fix(cursor): retry prompts that only returned a transport failure"}
    

    Exit 0, from the bare home directory. So "not a git repository" is not what's failing these calls — retracting that part of the original report. cwd really is project.workspaceRoot for the pr view call itself (confirmed via requireProject), it just isn't the cause of that call failing.

    I couldn't find the actual failing call site from the trace alone, since VcsProcessExitError only keeps stderrLength, not the text. See the triage comment below — the real mechanism is the gh api rate_limit quota probe inside GitHubCli.execute, which does hardcode process.cwd() rather than input.cwd. That's the actual cwd bug; it's just one call earlier in the chain than I looked.

  3. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 17, 2026
  4. changed the title [-][Bug]: PullRequestSyncReactor runs gh with server cwd ($HOME) instead of project root, failing every sync sweep[/-] [+][Bug]: PullRequestSyncReactor's gh calls fail with an opaque error — VcsProcessExitError discards stderr, so the real cause can't be diagnosed[/+] on Sep 17, 2026
  5. josephv123 commented on Oct 6, 2026

    @josephv123

    The blocking part of this appears fixed on main by #14673.

    The probe that failed with gh (<home>) was gh api rate_limit with process.cwd(). #14673 replaced it with a GraphQL rateLimit reading (budgetReading in apps/server/src/sourceControl/GitHubCli.ts), and guardedRead now ignores a failed reading: Cache.get(budgetReading, …).pipe(Effect.ignore), commented "A failed reading leaves the budget unknown; it never blocks the read itself." That's fix direction 2 from the triage, so pr view --repo runs even when the reading fails. The reading still runs from process.cwd(), but gh api graphql --hostname doesn't need a git checkout.

    What remains is the title's point that VcsProcessExitError doesn't include stderr. That's a contract-level choice, so it may be worth its own issue if wanted. Otherwise this one can probably be closed.

  6. juliusmarminge commented on Oct 6, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Resolved by the GitHub API stack that landed today (#16319, #16320, #16321). Pull request sync no longer shells out to gh; it calls GitHub's API in-process, and GitHubCli.execute is gone, so the opaque VcsProcessExitError from those calls can't happen anymore. API failures now map to specific reasons (unauthenticated, rate-limited with a retry time, not found). Closing; open a new issue with the new error if PR sync still fails.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions