Skip to content

[Bug]: Antigravity/Gemini randomly hangs in thinking, then session/cancel fails and the thread cannot recover #14619

Description

@heykerim

Before submitting

  • I searched existing issues and did not find an exact duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server + apps/desktop (Antigravity / ACP session lifecycle)

Summary

Antigravity/Gemini sessions in T3 Code intermittently get stuck in a running/thinking state even though the agent has stopped making progress. The elapsed timer keeps increasing and the thread still looks active, but no further output or tool activity arrives.

When I try to stop the stuck run, T3 Code often surfaces:

ACP transport operation call-rpc failed for method session/cancel.

and/or:

The request was canceled by the client.

After that, the thread may remain in a bad/stuck state instead of cleanly settling or recovering.

This has been happening repeatedly for roughly two weeks and, in my usage, only with the Antigravity/Gemini provider. I have not seen the same behavior with the other providers I use in T3 Code.

Steps to reproduce

This is intermittent rather than perfectly deterministic, but it happens frequently enough to block normal use:

  1. Open a T3 Code thread using the Antigravity provider.
  2. Select a Gemini model (currently observed with Gemini 3.8 Flash (High)).
  3. Run a normal coding task, especially one that takes several minutes and uses tools.
  4. At some point the agent stops producing output/activity, but T3 Code continues showing the turn as running/thinking and the elapsed timer keeps increasing.
  5. Wait several more minutes; the session does not recover.
  6. Click Stop / cancel the run.
  7. T3 Code reports:
    ACP transport operation call-rpc failed for method session/cancel.
    
    In the affected run shown in my screenshot, the UI also showed:
    The request was canceled by the client.
    
  8. The thread does not reliably return to a clean idle state.

A recent occurrence ran for 11m 39s before I stopped it and got the session/cancel transport error.

Expected behavior

  • If the Antigravity/Gemini stream stops or the provider session becomes unhealthy, T3 Code should either reconnect/recover or clearly fail the turn.
  • Clicking Stop should always settle the current turn cleanly as cancelled.
  • session/cancel transport failures should not leave the thread appearing active or unusable.
  • The UI should never remain indefinitely in a thinking/running state when the provider has stopped making progress.

Actual behavior

  • The thread randomly stops progressing while still appearing to run.
  • The timer continues indefinitely.
  • The agent produces no further output/tool activity.
  • Cancelling can fail with:
    ACP transport operation call-rpc failed for method session/cancel.
    
  • The thread can remain stuck instead of cleanly settling/recovering.

Impact

Blocks work completely for Gemini inside T3 Code.

I currently have to use Google Antigravity directly when I want to use Gemini because T3 Code is not reliable enough for these sessions.

Version or commit

T3 Code (Nightly) 0.0.43-nightly.20260929.2450 (27bdf1aa14e6)

This is the exact build from the latest reproduced occurrence shown in the attached screenshot.

The problem has been recurring for approximately two weeks across multiple Gemini sessions.

Environment

  • macOS
  • T3 Code desktop app
  • Provider: Antigravity
  • Model observed in latest occurrence: Gemini 3.8 Flash (High)

Logs or stack traces

Most visible error from the affected run:

ACP transport operation call-rpc failed for method session/cancel.

The UI also reported:

The request was canceled by the client.

I can provide the corresponding server.trace.ndjson / provider event logs if maintainers tell me which thread/session identifiers are most useful.

Related issues / prior work

This looks closely related to PR #11626:

#11626

That PR explicitly described and attempted to fix the same user-facing error:

ACP transport operation call-rpc failed for method session/cancel.

It also described Antigravity prompt cancellation failures and a short cancel timeout, but the PR was closed without being merged.

There is also #11670, which covers Antigravity ACP stream drops / failed recovery:

#11670

My report may share part of that lifecycle/recovery path, but the visible failure here is specifically a turn that remains stuck as running and then fails on session/cancel. I am not assuming the root cause is identical without the trace logs.

Workaround

Use Antigravity directly instead of T3 Code for Gemini sessions.

Activity

  1. juliusmarminge commented on Oct 1, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the report, @heykerim. This is a real bug, and it's still present on main (5cc99e1c23). It isn't a duplicate of #11670 (open issue). In that issue the Antigravity stream dies with a raw EOF. Here the thread sits in thinking with no error until you press Stop, and then Stop fails. You don't need to look up your version, because this code path hasn't changed since your build.

    What's happening

    An Antigravity turn has no progress timeout. The session stays running as long as session/prompt is open, so if Gemini stops sending updates and never finishes that call, the timer keeps counting until you click Stop.

    Stop for Antigravity sends session/cancel and then waits up to 15 seconds for the prompt to finish (cancelBehavior: "wait-for-prompt"):

    cancelBehavior: "wait-for-prompt",

    const cancel = Effect.gen(function* () {
    const started = yield* getStartedState;
    const activePrompt = yield* Ref.get(activePromptRef);
    if (options.cancelBehavior !== "wait-for-prompt") {
    if (Option.isSome(activePrompt)) {
    yield* Fiber.interrupt(activePrompt.value.fiber).pipe(Effect.ignore);
    }
    // Write cancel before a replacement prompt can reach the agent.
    yield* acp.agent.cancel({ sessionId: started.sessionId }).pipe(Effect.ignore);
    return;
    }
    yield* acp.agent.cancel({ sessionId: started.sessionId });
    if (Option.isNone(activePrompt)) {
    return;
    }
    const completed = yield* Effect.gen(function* () {
    const result = yield* Fiber.await(activePrompt.value.fiber);
    yield* Deferred.await(activePrompt.value.completed);
    if (Option.isNone(yield* Ref.get(terminationErrorRef))) {
    yield* drainEvents;
    }
    return result;
    }).pipe(Effect.timeoutOption(options.cancelTimeout ?? defaultCancelTimeout));
    if (Option.isNone(completed)) {
    const error = new EffectAcpErrors.AcpTransportError({
    operation: "call-rpc",
    method: "session/cancel",
    detail: "The ACP agent did not finish cancellation. Its process was stopped.",
    cause: undefined,
    });
    yield* retireRuntime(error);
    return yield* error;
    }
    if (Exit.isFailure(completed.value)) {
    return yield* Effect.failCause(completed.value.cause);
    }

    session/cancel is a notification, not a request, so the error you saw doesn't come from the cancel itself. If the prompt doesn't finish within 15 seconds, T3 kills the agent process and raises a made-up AcpTransportError for session/cancel. Its message leaves out the useful detail ("The ACP agent did not finish cancellation. Its process was stopped."), so all you see is ACP transport operation call-rpc failed for method session/cancel.:

    override get message() {
    const method = this.method ? ` for method ${this.method}` : "";
    return this.operation
    ? `ACP transport operation ${this.operation} failed${method}.`
    : "ACP transport operation failed.";

    The other message, "The request was canceled by the client.", is the open session/prompt coming back after the cancel. Instead of treating it as a normal cancellation, interruptTurn passes it on as a session/cancel failure, so the turn is never marked as cancelled:

    Effect.mapError((cause) =>
    isAcpError(cause) ? mapAntigravityError(input.threadId, "session/prompt", cause) : cause,
    ),
    Effect.tapError((cause) =>
    Effect.suspend(() =>
    intent
    ? context.promptLock.withPermit(
    finishTurn(intent, { state: "failed", errorMessage: cause.message }),
    )
    : Effect.void,
    ),
    ),
    Effect.onInterrupt(() =>
    context.promptLock.withPermit(
    Effect.gen(function* () {
    const turn = intent;
    if (!turn || turn.settled || context.stopped || context.generation !== turn.generation)
    return;
    const promptFiber = context.promptFiber;
    yield* cancelRequests(context);
    yield* Effect.ignore(context.runtime.cancel);
    if (promptFiber) yield* Fiber.interrupt(promptFiber);
    yield* finishTurn(turn, { state: "cancelled", stopReason: "cancelled" });
    }),
    ),
    ),
    );
    });
    const interruptTurn: Adapter["interruptTurn"] = (threadId) =>
    Effect.gen(function* () {
    const context = yield* requireSession(threadId);
    // A command that outlived its turn keeps running in the agent, and
    // session/cancel only stops a prompt. The agent kills its background
    // commands when its session closes, so Stop with nothing else running
    // ends the session, as Claude's does. The next turn resumes it.
    let idleWithCommands = false;
    yield* context.promptLock
    .withPermit(
    Effect.gen(function* () {
    // Decided under the prompt lock so a turn cannot start in between.
    if (!context.promptFiber && [...context.commands.values()].some((c) => c.promoted)) {
    context.stopped = true;
    idleWithCommands = true;
    return;
    }
    yield* cancelRequests(context);
    yield* context.runtime.cancel;
    }),
    )
    .pipe(
    Effect.mapError((cause) => mapAntigravityError(threadId, "session/cancel", cause)),
    // Once marked stopped the session must close, even if this call is

    Because the process gets killed, the runtime stays in a failed state, and the next prompt runs into the same unrecoverable session as #11670.

    What's upstream and what's ours

    The agent going silent and not finishing the cancel quickly is upstream, in agy and Gemini. T3 owns the rest: there's no stall timeout while a prompt is open, cancel errors show up as a transport failure, and Stop doesn't mark the turn as cancelled. #11626 (closed PR, not merged) targeted these same error strings. It also raised the cancel wait from 15 seconds to 60, which is a product decision.

    Proposed fix

    1. When the prompt fails because a cancel was requested ("The request was canceled by the client." or a context-canceled error), complete the turn as cancelled instead of reporting a session/cancel transport error.
    2. When the 15-second wait runs out, still stop the process, but mark the turn as cancelled, and if an error is shown at all, show the real detail.
    3. Keep the 15-second default unless maintainers agree to change it.

    A trace is optional. If you still have the run, the useful part is from the last session/update to the moment you clicked Stop: the thread id, when you clicked Stop, and whether a process exit follows the cancel.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 1, 2026
  3. GorlikItsMe commented on Oct 2, 2026

    @GorlikItsMe

    Note

    Antigravity responding on behalf of @GorlikItsMe.

    Adding the deterministic repro notes and trace details as requested from #14854:

    Deterministic Reproduction: Active Background Task + Queued User Message

    While this issue also occurs intermittently during long "thinking" turns, we found a 100% reproducible trigger involving background tasks:

    1. Background task started: During a turn, the agent runs a persistent service (e.g. bun run dev / npm run dev), which the Antigravity harness tracks as a running background command task.
    2. Text generation finishes: The assistant streams its complete answer text to the UI.
    3. Turn remains open on backend: Even though no more text or tool calls are pending, the session/prompt call never completes because the background process is alive and occasionally emitting activity updates (tool.updated). In our logs, AntigravityAdapter.sendTurn remained active for 462,564 ms (~7.7 minutes).
    4. User queues a follow-up message: Seeing the response on screen while the composer shows the turn still running, the user submits a new prompt. The message enters the queue.
    5. Preempting the turn triggers session/cancel: T3 Code picks up the queued message, dispatches thread.turn-start-requested, and calls interruptTurn -> session/cancel.
    6. 15s cancel timeout expires: Because Antigravity uses cancelBehavior: "wait-for-prompt", T3 Code awaits activePrompt.value.fiber with defaultCancelTimeout (15s). With the background task still attached, Antigravity doesn't exit the prompt fiber within 15s.
    7. Fatal session crash: At 15,019 ms, timeoutOption raises AcpTransportError, invokes retireRuntime, and stops the session (status: "stopped" / session.exited).

    Relevant Log Traces

    From server.trace.ndjson:

    {"level":40,"time":1759428618287,"pid":486518,"name":"AntigravityAdapter.sendTurn","durationMs":462564.6,"threadId":"c285cbba-ab4d-45f0-8123-ae9debb8d315","error":{"name":"ProviderAdapterRequestError","provider":"antigravity","method":"session/cancel","detail":"ACP transport operation call-rpc failed for method session/cancel."}}

    Timeline from SQLite orchestration_events & provider event stream:

    18:05:22.312Z  Assistant finished streaming final response text
    18:06:22.608Z  Background task activity update received (bun run dev status: inProgress)
    18:10:03.255Z  User submits new message -> queued -> thread.turn-start-requested -> interruptTurn called
    18:10:03.255Z  T3 Code invokes session/cancel (cancelBehavior: "wait-for-prompt", timeout: 15s)
    18:10:18.287Z  (+15,032 ms) cancel timeout hit -> AcpTransportError: "The ACP agent did not finish cancellation. Its process was stopped."
    18:10:18.307Z  session/prompt fails with errorTag: "Interrupt" ("The request was canceled by the client.")
    18:10:18.349Z  session.exited emitted (exitKind: "error", reason: "Antigravity process stopped.")
    

    This confirms the proposed fix in triage: handling cancel timeouts / prompt interrupts without retiring the entire adapter runtime will completely prevent this crash.

  4. guillaume-murat commented on Oct 3, 2026

    @guillaume-murat

    Same issue here, hope it will be fixed soon ! ❤️

  5. added 2 commits that reference this issue on Oct 5, 2026
    7963774
    49793d0
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.upstreamvia-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions