Skip to content

Keep the Editor responsive while a Forge write stalls - #503

Merged
sandersonstabo merged 8 commits into
masterfrom
concurrent-requests
Oct 10, 2026
Merged

sandersonstabo merged 8 commits into
masterfrom
concurrent-requests

Conversation

@sandersonstabo

@sandersonstabo sandersonstabo commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Under disk contention one Forge SQLite write can stall for ~44 s, and every navigation request queued behind it ("the Editor is still busy", "Images are still uploading").

  • Forge answers up to 16 requests per connection concurrently. Subscribe/Unsubscribe keep accept order under the delivery driver, and write-gate stalls are logged.
  • Transport: SessionRequester runs requests beside the session owner's, each on its own stream (up to 8 at once). Requesters share one correlation registry, and dropped requests are abandoned cleanly.
  • Editor: draft saves, reads, uploads, and sends run on a mutation lane, so they no longer block navigation. Each draft scope runs one job at a time in queue order, so saves stay ordered and a send follows its save. Subscriptions and deliveries stay on the command loop. Losing the connection causes one reconnect, durable writes retry once under their stable identity, and shutdown drains the lane.

Draft saves were already coalesced app-side (DraftSync: one save in flight per scope, latest body wins).

Tests: new mutation_lane_tests::navigation_is_answered_while_a_draft_save_is_unanswered (a scripted Forge holds a save until a navigation request is answered). The frontend lib's 20 failures from master are unchanged.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Keeps the Editor responsive when a Forge SQLite write stalls on a busy disk: navigation is answered while draft saves, uploads, and sends queue behind the write, instead of piling up behind "Editor is still busy".

  • Forge answers up to 16 requests per connection concurrently; Subscribe/Unsubscribe stay in accept order under the delivery driver, and a write-gate wait or hold longer than two seconds is logged with the operation's name.
  • Transport gains SessionRequester, which runs requests beside the session owner's on their own streams (up to 8 at once) against a shared correlation registry, with clean abandonment of dropped requests.
  • The Editor's draft saves, reads, uploads, and sends run on a mutation lane, so they no longer block navigation; each draft scope runs one job at a time in queue order, keeping saves ordered and a send after its save. Subscriptions and deliveries stay on the command loop; a lost connection reconnects once, durable writes retry once, and shutdown drains the lane.
  • Adds a scripted-Forge test that holds a save until a navigation request is answered; the frontend lib's 20 failures from master are unchanged.

Written for commit dca5f4f. Summary will update on new commits.

View guided diff

sandersonstabo and others added 8 commits October 11, 2026 00:58
Every write now names its operation when it queues for the write gate.
A write that waits for the gate, or keeps it, longer than two seconds is
reported to the process log and the flight recorder with the operation
that holds it, so a write stalled in the kernel shows up by name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`receive_server_request` and `send_server_reply` are the two halves of
`dispatch_server_request_with_receipt`, for a server that bounds reading,
answering, and replying separately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A connection answered one request at a time under the request limit, so
one write stuck behind a slow disk held up every request behind it and
then failed the whole connection. Requests now run side by side, up to
the sixteen streams a peer may keep open. Subscription requests still
hold the delivery driver for their whole stage, in accept order, and a
lifecycle request runs alone. Every other request's answer is unbounded;
reading it and writing its reply stay bounded, and its failure resets
only its own stream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clippy's pedantic single_match_else rejected the two-arm match on the
gate's holder.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An independent request wrote its reply and then queued for the driver,
so a wake scan could run between the two and push against a listing
the peer had not yet been shown as current. It now takes the driver
before writing its reply and keeps it through its follow-up.

A subscription request took the driver before the answering lock while
every other lane takes them the other way round; it now waits for its
accept-order turn, then the answering lock, then the driver, and gives
the turn up only once it holds the driver. The completed count is taken
when a reply settles, so a follow-up that fails no longer hides it.
Large stage futures are boxed.

The stalled-write test opens the pool's connections before the stall,
as a running Forge has them: opening a fresh pooled connection while
another process holds the write lock can itself wait on that lock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A SessionRequester runs requests on the session's connection beside the
owner's, each on its own stream, up to CONCURRENT_REQUEST_CAPACITY at
once. Requesters share the owner's correlation registry; a dropped or
failed request abandons its entry so its slot frees and its identity
stays retired.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Draft saves, reads, attachment uploads and reads, and sends now run on a
mutation lane: each job in its own task on a SessionRequester of the
service's one session, so a stalled Forge write no longer holds up
Unsubscribe, Subscribe, ReadActiveRun, snapshots, or any other command.
Jobs of one draft scope run one at a time in queue order, so saves apply
in order and a send follows the save that stored its body. Subscriptions
and deliveries stay on the command loop. A job that loses its connection
asks the loop for a session; the loop reconnects only if the lost
connection is still current, and a durable write retries once under its
stable identity as before. Shutdown drains the lane first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Formats the branch's files and keeps client_session.rs under its
file-size allowance.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sandersonstabo
sandersonstabo merged commit bcec3db into master Oct 10, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant