Skip to content

[Bug]: Orchestration read model and per-thread client/VCS state grow unbounded over uptime #4178

Description

@RusiruSadathana

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Run the server (t3 serve) and keep it up for an extended session (many hours), using several concurrent threads and creating/archiving/deleting threads over time.
  2. Watch the server process RSS and CPU time (for example with ps/htop), and watch the web UI responsiveness.
  3. Compare a freshly started server against one that has been up for 12+ hours with the same workload.

Observed on a real server: the process reached roughly 1 GB RSS with hundreds of minutes of accumulated CPU time, and the web UI became progressively laggy the longer the server stayed up.

Expected behavior

Steady-state memory and per-event CPU should stay roughly flat over uptime. Command processing cost should not scale with the total number of threads ever created, and deleted threads should not keep costing work forever.

Actual behavior

The in-memory orchestration command read model stores threads and projects as arrays. Every domain event runs an O(N) linear .find() over the threads array and produces a full-array .map() copy through the projector, so per-event work and allocation scale with the total thread count. Deleted and archived threads are never removed from the model (only timestamped), so N grows monotonically for the life of the process and every subsequent event gets slower, increasing GC pressure.

Two client-side contributors compound this in long-lived browser tabs:

  • Several per-thread stores (preview state, right panel, diff panel, ui state) accumulate one entry per thread ever visited and are never pruned on thread deletion; the ready-made cleanup functions exist but are not wired into the delete flow. previewStateAtom is keepAlive, so its atoms are retained for the tab lifetime.
  • The VcsStatusBroadcaster per-cwd status cache is never pruned; it grows one entry per distinct cwd for the life of the server process.

Impact

Major degradation or frequent failure

Version or commit

main @ 2640e6d

Environment

Linux server (t3 serve), Node 24; reproduced with multiple concurrent threads over long uptime.

Workaround

Restarting the server clears the in-memory model and temporarily restores responsiveness, but the growth returns over time.

Activity

  1. juliusmarminge commented on Jul 20, 2026

    @juliusmarminge
    Member

    PR #4176 has been refreshed without opening a new PR: branch codex/refresh-pr-4176 at commit 56b6615afd.

    The refresh replays cleanly over current main, preserves #4177 catch-up behavior, and adds the missing VCS cache-eviction regression. Verification passed 181 focused backend tests, 76 client-store tests, 6 catch-up/recovery cases, targeted lint/typechecks, and React Doctor. A deterministic 20k-thread benchmark measured 297.2 ms for the prior array path versus 2.65 ms for the HashMap path (~111x). No new PR has been opened.

  2. danieliser commented on Jul 22, 2026

    @danieliser

    A related 4.43 GiB real-world profile supports keeping archived history separate from resident-memory management.

    The immediate desktop OOM was fixed without deleting archived tasks: retain shell metadata globally, load task detail only when viewed, page large activity collections, replace stale event replay with a bounded snapshot, and dispose unsubscribed detail state after its idle TTL.

    The proven crash path was stale subscribeThread cursors replaying hundreds of thousands of global events and filtering by task only after SQLite/V8 decoding. Electron’s first-ten detail prewarmer multiplied those catch-ups. With #3510 plus removal of that prewarmer, the same database ran in the low hundreds of MiB instead of aborting near 2.8 GiB.

    #4176 appears to address the distinct long-uptime concerns from this issue: O(1) command state and cleanup for deleted-task/client/VCS stores, while deliberately retaining archived tasks. That separation matches what worked in this reproduction.

    Archived content can remain durable and cold on disk until requested. Disk-retention policy may still be valuable, but archive should not imply history loss or permanent in-memory hydration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.🚧 In Progress

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions