Skip to content

Oversized Codex diff event exhausts desktop backend heap #10924

Description

@alexfertel

What happened

The T3 Code desktop backend stops without warning. The UI then marks active work as failed and shows:

Provider session did not survive a server restart. Send a new message to continue.

The fault happened three times over two days.

Diagnosis

The desktop backend runs out of V8 heap while it handles a large Codex turn/diff/updated notification.

CodexAdapter places the same diff in two fields of the canonical event:

  • raw.payload.diff
  • payload.unifiedDiff

Source:

  • apps/server/src/provider/Layers/CodexAdapter.ts:985
  • apps/server/src/provider/Layers/CodexAdapter.ts:1646

ProviderService sends this canonical event to the provider event logger before it publishes the event:

  • apps/server/src/provider/Layers/ProviderService.ts:916

EventNdjsonLogger serializes the complete event and holds the resulting line in memory:

  • apps/server/src/provider/Layers/EventNdjsonLogger.ts:497
  • apps/server/src/provider/Layers/EventNdjsonLogger.ts:620

In the observed crash, one redacted canonical log record exceeded 250 MiB. The duplicate diff and temporary serialization copies pushed the backend past its heap limit.

After the backend restarts, reconcileProviderSessions marks sessions whose provider process is gone as failed. This creates the user-facing restart message:

  • apps/server/src/serverRuntimeStartup.ts:338

Expected behavior: T3 Code should bound, omit, or summarize large raw diff data before diagnostic serialization. One provider event should not stop the backend or unrelated work.

Steps to reproduce

Proposed source-level reproduction:

  1. Start the T3 Code desktop backend with the Codex provider.
  2. Send a valid turn/diff/updated notification with a large generated string in params.diff.
  3. Let CodexAdapter convert it to a canonical event.
  4. Let the canonical provider event logger serialize the event.
  5. Observe that the serialized record contains the diff twice.
  6. With a large enough diff or existing heap use, V8 stops the backend with an out-of-memory error.

A standalone synthetic test has not yet confirmed the minimum payload size.

Version

0.0.41-nightly.20260909.1439, commit 6c583620ff7a

This was the newest nightly, and its commit matched main when checked.

Environment

  • T3 Code desktop app with local backend
  • Darwin 25.6.0, arm64
  • Node.js 25.6.0
  • Codex CLI 0.153.4

Evidence

<--- Last few GCs --->

Scavenge 1856.5 (...) -> 1856.7 (...) MB
Scavenge 1856.7 (...) -> 1848.5 (...) MB

FATAL ERROR: Zone Allocation failed - process out of memory

Sanitized provider log measurement:

stream: CANON
method: turn/diff/updated
record size: greater than 250 MiB
payload content: redacted

No secrets, source paths, thread identifiers, thread titles, commands, query data, or payload content are included.

Related issues

Fix applied or workaround

Restarting the desktop app restores the backend for a time. Avoiding very large single-event diffs reduces the risk. No local data was changed.

Filed by

Codex, GPT-5, via t3 triage

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions