Skip to content

[Bug]: Grok stays working after the final answer, then Stop files that answer as a superseded partial #15489

Description

@letrandat

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Searched open and closed issues for Superseded attempt and Partial output retained. No issue uses those labels. Related but not the same:

Area

apps/web

Steps to reproduce

  1. Open a Grok thread in T3 Code Nightly. Model was grok-4.7.
  2. Send one long coding task (commit a small change, then install it for the local agent harnesses).
  3. Leave the turn running. The header stays on "Worked for …" for about 21 minutes.
  4. The timeline already contains the full final answer: the task is done, and the last message says so.
  5. Stop the run, because the thread still looks in progress.
  6. Expand the row labeled Superseded attempt / Partial output retained.

The same pattern has not shown up on the other providers in this app (Claude, Codex, Cursor).

Expected behavior

Once Grok has written the final answer, the run indicator clears and that answer stays in the main timeline. Stop is not required. If an attempt is truly superseded, "Partial output retained" means text that was cut off, not a complete final reply.

Actual behavior

The thread kept looking in progress after the final answer was already on screen ("Worked for 21m"). Stopping it collapsed that finished answer under Superseded attempt / Partial output retained. Expanding the row shows the complete reply, including the tool summaries and the closing paragraph. Nothing in that fold was cut off.

The banner is the copy in apps/web/src/components/chat/MessagesTimeline.tsx (Partial output retained), applied to every superseded-attempt fold from deriveSupersededAttemptFolds in MessagesTimeline.logic.ts. A Grok attempt that has already finished can still be marked superseded, so a complete answer is presented as a partial.

Impact

Major degradation or frequent failure

A finished Grok turn is indistinguishable from a live one. Stopping it, which is the obvious next step, hides the finished answer behind a banner that says the output is partial.

Version or commit

T3 Code (Nightly) 0.0.46-nightly.20261004.2644

Environment

  • macOS
  • T3 Code (Nightly) 0.0.46-nightly.20261004.2644
  • Provider: Grok
  • Model: grok-4.7
  • Grok CLI: 1.0.46 (2765805b9442) stable

Logs or stack traces

No session log is attached. The timeline evidence is the header "Worked for 21m", a resume line timestamped 9:36 PM local, and the expanded fold containing the full final assistant message under the superseded banner.

Screenshots, recordings, or supporting files

Not attached. The capture includes repository and account names.

Workaround

Expand Superseded attempt. The full answer is in that fold. The work itself had already finished before Stop.

Activity

  1. juliusmarminge commented on Oct 4, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the careful write-up, @letrandat. This looks like a real Grok issue, and it isn't a duplicate of the ones you linked. #3580 and #7210 were earlier stuck-on-Working bugs that are already closed, #8283 and #9154 are about mid-turn sends, and open PR #15413 covers threads staying in Working after a background command, which is a different path.

    What I found

    It looks like two separate things are happening here.

    • The turn staying open. For Grok, the provider turn only settles when turn_completed / x.ai/session/prompt_complete arrives. Finalizing is also held back while a background tool or subagent is still running, while monitor hydration is pending (capped at 60s), or while an injected report is outstanding. So a complete reply can already be on screen while the run still counts as active and Stop is still offered. While a turn is open the header should read Working for; Worked for is the settled fold.
    • The "Superseded attempt" row. deriveSupersededAttemptFolds in MessagesTimeline.logic.ts only builds that fold for an attempt whose status is superseded, and the only place that status is set is a steering restart (dispatchSteerIntoRun, reason: "steering_restart"). A plain Stop marks the attempt interrupted and should leave the final message in the main timeline. Grok steers by interrupting and restarting, which replays session/resume, so the resume line you saw suggests a replacement attempt started before the first one settled. The "Partial output retained" text is shown for every superseded fold, which is why a complete reply ended up labeled as partial.

    If you can, a redacted copy of any of these from the affected turn would help a lot (please leave out tokens, pairing credentials, and home-directory paths):

    • server.trace.ndjson around that turn, or the Grok provider event log
    • Whether a second attempt with steering_restart shows up before the stop, and whether turn_completed ever arrived
    • A screenshot of the resume row and the header exactly as rendered (Working for vs Worked for)

    Likely fix area

    • Grok turn finalization: why turn_completed doesn't arrive, or why finalize stays deferred, after the final answer.
    • Whether a steering restart should mark an attempt superseded when it already produced a complete final message.
    • Possibly making the "Partial output retained" copy depend on whether the attempt's output was actually cut off.

    Until then, expanding Superseded attempt is the workaround, since the full reply is in that fold.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 4, 2026
  3. letrandat commented on Oct 4, 2026

    @letrandat
    Author

    Thanks for the split. The local projection and the Grok session log answer the three questions. Paths, account names, and the task text are left out.

    Header

    The capture of the settled fold reads Worked for 21m, not Working for. The resume row under it is timestamped 9:36 PM local, which is 2026-10-04T04:36:07Z.

    That duration matches the superseded attempt, not the provider finish:

    start end duration
    Grok turn_ended 04:36:07.242Z 04:43:09.332Z 7m 2s, outcome: completed
    T3 run attempt (ordinal 1, reason: initial) 04:36:07.193Z 04:56:58.194Z 20m 51s, status superseded

    04:56:58Z is the timestamp of the follow-up message ("what are you working on, it has been 20 minutes"). There is no capture of the header from the 13 minutes between those two timestamps, so this does not show whether it said Working for while it still looked active. After the fold settled, it said Worked for 21m.

    Second attempt

    Yes. A steering_restart was recorded, and it starts at the follow-up, before any separate Stop:

    • Attempt 1: reason: initial, status superseded, completedAt 04:56:58.194Z
    • Attempt 2: reason: steering_restart, status completed, 04:56:58.225Z → 04:57:07.357Z

    The provider-turn row for attempt 1 is interrupted (04:36:07.231Z → 04:56:58.205Z). The run-attempt row is superseded. So the follow-up was applied as a steer. It was not stored as a plain Stop that leaves the final message on the main timeline.

    turn_completed / prompt_complete is not in the retained server trace for this window (the trace rotated). The Grok session log is unambiguous: turn_ended with outcome: completed at 04:43:09.332Z, then no events at all until turn_started at 04:56:58.258Z. The provider had finished. T3 kept the attempt open until the next message restarted it.

    The 9:36 PM resume line matches a session reconnect at the same second as turn_started (mcp_config_resolved at 04:36:07.120Z, then turn 46). The previous provider turn had been cancelled at 04:35:26Z with Grok cancellation_context.trigger: session_close.

    What the fold contains

    The expanded "Superseded attempt / Partial output retained" row is the full final reply from the attempt that Grok marked completed at 04:43:09Z. The text is not cut off. The partial label is on a complete answer because that attempt was later marked superseded.

    Happy to pull a narrower redacted slice of the Grok turn_started / turn_ended records if that is useful. I am not attaching the screenshot; it includes repository and account names.

  4. juliusmarminge commented on Oct 8, 2026

    @juliusmarminge
    Member

    Thanks for the projection and session-log details, @letrandat. They settle the labeling question: the "Superseded attempt / Partial output retained" fold holds a reply that was complete. The label is applied to every superseded attempt, whether or not its output was cut off. We'll treat that as its own fix.

    The part we can't explain yet is why the attempt stayed open from 04:43:09Z (Grok's turn_ended, outcome: completed) until 04:56:58Z. The trace for that window rotated away, so we can't see whether prompt_complete reached T3 or whether finalizing was held back by background work or monitor hydration.

    If it happens again, could you capture the following before sending a follow-up or pressing Stop?

    1. ~/.t3/userdata/logs/server.trace.ndjson (copy it right away so it doesn't rotate), redacted.
    2. The Grok provider event log for the thread from ~/.t3/userdata/logs/provider/, covering the final answer through the next message.
    3. Whether the thread showed any background task, subagent or monitor as still running at that point (the footer or pending-work indicator).

    That will tell us which hold kept the turn open.

  5. Mkassabov commented on Oct 8, 2026

    @Mkassabov

    Fresh evidence for the unexplained window, from a t3 triage session on another machine. The traces had not rotated yet, so this run shows which hold kept the turn open: background subagents whose inference stalled for an hour. The final answer did reach T3.

    Environment: T3 Code Nightly 0.0.46-nightly.20261008.2813 (30cc788), desktop app on Linux x64 (CachyOS, kernel 7.1.8), Node 26.8.2, Grok CLI 1.0.46 (2765805b9442), model cpa-alchemy-claude-opus-5-5 through a custom gateway. Seen from both the desktop app and the Android app (over Tailscale).

    Timeline (run 17 of one thread, UTC, 2026-10-08)

    time source event
    08:49:24 T3 run 17 attempt 1 starts (session/prompt)
    08:50:08 Grok root calls spawn_subagent twice (toolu_01LJjY…, toolu_012pZJ…)
    08:50:44 / 08:51:47 Grok each subagent starts an inference call that never returns a chunk
    09:08:29 T3 user steers → attempt 1 superseded, attempt 2 starts
    09:08:50 both root reply finishes ~20 s later; Grok logs shell.handle_prompt.done ok:true
    09:15:56, 09:21:55, 09:32:05 T3 more steers → attempts 3, 4, 5. Each one replies in 15–30 s and then holds
    09:32:20.133 Grok shell.handle_prompt.done {"ok":true} for attempt 5
    09:32:20.151 T3 RpcClient.session/prompt exits Interrupted (the benign end-of-turn event from #15888), and the final assistant_message item completes at the same ms
    09:32:20 → 09:50:04 — nothing happens. Both T3 subagent rows are running. Root Grok session is idle
    09:51:51 Grok both subagents: inference_failed kind=idle_timeout "inference idle timeout after 3600s with no chunks"
    09:51:54–55 both subagent failed → T3 subagent rows failed → run 17 attempt 6 finalizes completed at 09:51:58

    So for this run, deferFinalizeForBackgroundWork held the turn open correctly, because the subagents really were alive. The UX problem is in how that hold behaves:

    1. Nothing tells the user the thread is waiting on subagents. From their side the agent has answered and the thread is stuck on "working".
    2. The obvious next step is to steer, and that makes it worse. Every steer supersedes the attempt (6 attempts, 5 superseded), and each completed reply gets folded as "Superseded attempt / Partial output retained". To the user, history disappears. Steering never releases the hold, because the subagents are still running.
    3. Queued follow-ups wait the whole time, so the queue looks like it ignores an idle agent.
    4. Grok's 3600 s inference idle timeout is the only thing that ever ends the hold. A stalled model stream on a subagent wedges the parent thread for an hour.

    Per-attempt evidence (orchestration_v2_projection_provider_turns, with the last turn-item update in each turn):

    pturn  status       started     completed   last item
    17     interrupted  08:49:24    09:08:29    (subagents spawned 08:50:08)
    18     interrupted  09:08:29    09:15:56    09:08:50
    19     interrupted  09:15:56    09:21:55    09:16:25
    20     interrupted  09:21:55    09:32:05    09:25:30
    21     interrupted  09:32:05    09:50:04    09:32:20
    22     completed    09:50:04    09:51:58    09:50:30   (released only by subagent failure at 09:51:55)
    

    A scheduled task (status-update-5m) was also posting into the thread every 5 minutes. Its queued posts were turned into steers at 09:32:05 and 09:50:04, adding to attempts 5 and 6. Two identical queued posts (runs 23 and 24) were then cancelled at 09:50:05–06 and never delivered.

    Possible directions: show a "waiting on N subagents" state on held turns; let a steer on a settled-and-held turn start a new turn instead of superseding a complete attempt; enforce a T3-side stall bound or a per-subagent cancel affordance.

    Posted from a t3 triage session by Claude Code (Opus 5.5).

  6. creedants commented on Oct 10, 2026

    @creedants

    More data on the open window, from the delegated-task side rather than the UI.

    Environment: T3 Code 0.0.46-nightly.20261008.2819 (5e2225671f70), Linux, provider grok, model grok-4.7, full-access child runs started with delegate_task in mode: "async".

    Same symptom, different consequence. A Grok child posts its complete final message and the run stays open. For a delegated task this means the parent never gets its completion wake. task_status keeps returning status: running, workState: working, hasPendingChildRuns: false until the parent calls task_cancel, and the run then ends as interrupted.

    How often. I counted three days of delegated runs in one project:

    Provider Runs Final message posted, run still open more than 2 minutes later
    Grok (grok-4.7) 160 13 (8.1%)
    Claude, Codex, OpenCode, Cursor combined 657 0

    All 13 ended as interrupted after a cancel. None closed late on its own. The median gap from the final message to the end of the run was about 19 minutes and the longest was about 24 hours. Ten grok-4.7-build-fast runs in the same window did not show it. This is a count from normal use, not a controlled benchmark, and most Grok runs close promptly (one closed 0.157 s after its final message).

    Three cases where I confirmed the parent's task_cancel in the activity log:

    Final message (UTC) Run ended by cancel (UTC) Gap
    2026-10-08 10:00:04 2026-10-08 10:59:56 59 m 52 s
    2026-10-09 02:57:51 2026-10-09 03:17:14 19 m 23 s
    2026-10-09 04:44:57 2026-10-09 04:56:03 11 m 6 s

    hasPendingChildRuns was false before the cancel in these, so on the T3 side the child was not waiting on children of its own. That fits the background-subagent explanation above only if the hold is inside the Grok session and not visible to T3's task state.

    I did not capture the Grok ACP notifications for an affected run. I can collect traces from the next one if you say which files help. If the delegated-task case should be its own issue, I can open one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.needs more infoInitial triage showed no bug. Awaiting more infovia-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions