Skip to content

🧭 feat: Mid-Run Steering and Queued Messages for Agent Runs - #14220

Merged
danny-avila merged 33 commits into
devfrom
claude/steering-queuing-planning-f9d487
Jul 14, 2026
Merged

danny-avila merged 33 commits into
devfrom
claude/steering-queuing-planning-f9d487

Conversation

@danny-avila

@danny-avila danny-avila commented Jul 11, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

I added mid-run steering and post-run message queuing to agent runs: while a response is generating, Enter now sends the message into the live run — injected into the graph at the next tool-batch boundary — or queues it to auto-send after the run finishes, per a new user preference, with an "interrupt & send" escape hatch for urgent redirects.

  • Added POST /api/agents/chat/steer: validates text (moderation, 16k cap), authorizes against the job (userId/tenant), and enqueues into a new job-store steer queue — cross-instance safe (Redis Lua with an atomic running-status guard and depth cap of 10; single-threaded equivalent in the in-memory store), exposed as GenerationJobManager.steering (SteeringLifecycle). Rejection codes tell the client how to degrade: NO_ACTIVE_RUN → normal send, RUN_PAUSED/STEER_UNSUPPORTED/STEER_QUEUE_FULL → client-side queue.

  • Registered a run-scoped PostToolBatch drain hook (packages/api/src/agents/steering/) that empties the queue FIFO at each tool-batch boundary, injects each steer into graph state as its own user message via the SDK's hook injectedMessages, and skips subagent scopes and replaced jobs. createRun now passes a hook registry independent of the tool-approval policy; steering needs neither HITL nor a checkpointer.

  • Recorded each injected steer as an inline ContentTypes.STEER part on the response message at the live content index, with a mutable index-offset wrapper (createSteerIndexOffsetHandlers, the mid-run analog of the HITL resume offsets — composes with them on resumed runs) so subsequent SDK step indices land past the insertion. Emitted on_steer_applied over the existing SSE; the Redis chunk-log reconstruction splices steer parts at their recorded index for cross-replica reconnects.

  • Replayed persisted steer parts as standalone user messages in formatAgentMessages on later turns, preserving tool-call/result adjacency; token counting reads the flat steer: string shape for free. Exempted steer parts from the hide_sequential_outputs filter and included them in conversation export (You (steered) label).

  • Reported steers that never reached an injection boundary on the final/abort events and resumeState.pendingSteers, so the client converts them to queued follow-ups instead of dropping them. Every terminal path also parks leftovers under a dedicated bounded-TTL store key (independent of the job record, which the default completeJob path deletes immediately): the normal final, generation errors, aborts, resume failures, and approval expiry (snapshot before the requires_action→aborted CAS, park only when it wins). The payload carries its owner, so the status route recovers steers even on its jobless branch (the common reload-after-terminal case) with claim-on-read semantics — a second reload can't re-mint dismissed items, and a non-owner claim returns nothing and re-parks.

  • Closed the Redis snapshot→subscribe resume gap: after attaching, subscribeWithResume re-peeks the steer queue and re-surfaces any on_steer_applied events published in the window (synthesizeAppliedSteerEvents rebuilds them from the durable content view), updating resumeState.pendingSteers to the live queue.

  • Built the composer UX: Enter during a run steers by default (configurable in Settings → Chat, duringRunDefaultAction), degrades to queue while paused on a tool approval or before a conversation id exists. The send button reuses the send/stop slot — with composer text it replaces Stop, and hovering it reveals the action list with shortcuts (Steer ⏎ / Queue ⌘⏎ / Interrupt & send ⌥⏎; ⌘⏎ runs the non-default action, ⌥⏎ interrupts). A submitted steer enters the chat history immediately as a standard user message (SteerPart: user icon + author header + user text presentation) at its projected injection point — the end of the streamed content — and on_steer_applied swaps the optimistic entry for the persisted part at its authoritative index, so live, reload, share, and search views all render the same thing and the visible order matches what the next turn replays. Rows above the textarea remain only for recoverable states: failed steers (retry / edit / convert-to-queue / dismiss) and queued messages (send-now, edit, remove, default-mode toggle).

  • Implemented the queue drain (useQueueDrain): a one-shot runEndByIndex signal from the SSE final/error handlers dispatches one queued message per clean completion (FIFO — each new turn's own final drains the next), migrates a queue keyed under a new conversation to the real id, and leaves chips in place on user aborts or errors unless the one-shot interrupt-and-send flag is armed. HITL answer mode always wins the composer; ask-question and tool-approval flows are untouched.

  • Fixed the interrupt & send settlement race: the aborted run's final SSE starts the next submission while the abort POST is still in flight, and the response handler's unconditional submission clear was tearing down the new run's stream before it attached (the follow-up ran server-side but the live placeholder finalized empty). useAbortCleanup captures the submission before the abort round-trip and both settlement paths clear only when it is still the current one.

  • Hard-gated everything on SDK support: isSteeringSupported() probes both halves of the SDK contract (HOOK_INJECTED_MESSAGES_CAPABLE for injection and the steer content type for replay), the controller returns 501 and createRun skips the drain wiring on older versions, so a queued steer can never be drained by an SDK that would silently drop it.

  • Added a Playwright e2e suite (e2e/specs/mock/steering.spec.ts) driving all three flows against the mock harness with a real MCP tool boundary (new E2E_STEER_TOOL_REPLY fake-model marker): steer mid-run (202 → immediate in-thread part → applied at the tool boundary → survives after run end), ⌘⏎ queue → auto-send after clean completion, and ⌥⏎ interrupt & send with the follow-up streaming live. Writing it surfaced two real bugs — the SDK's top-level agentId stamping (fixed upstream, below) and the abort settlement race (fixed above).

Depends on @librechat/agents ≥ 3.2.63, pinned in this PR (api/ + packages/api/ + lockfile). 3.2.63 scopes the hook agentId subagent-scope marker to child graphs (LibreChat-AI/agents#307): on 3.2.62 the event-driven ToolNode stamped it at the top level, so the drain hook treated every boundary as a subagent scope and mid-run injection never fired — steers only degraded to the run-end path. 3.2.62 carried the injection/replay seam itself (LibreChat-AI/agents#299 + #304).

Change Type

  • New feature (non-breaking change which adds functionality)

Testing

  • packages/api (npx jest): new suites — stream/__tests__/steering.spec.ts (lifecycle: FIFO drain, status guard, depth cap, terminal cleanup, abort/resume reporting, park/claim incl. survival of the default job-deleting cleanup, non-owner claim protection, approval-expiry parking, resume-gap event synthesis), agents/steering/__tests__/runtime.spec.ts (subagent/job-replacement guards, injectedMessages shape, applySteer error isolation), agents/steering/__tests__/offset.spec.ts (mutable offset, composition with resume offsets); RedisJobStore.stream_integration.spec.ts steering block passes against live Redis (USE_REDIS=true, 57 passed) including the dedicated parked-steers key (survives deleteJob, claim-once, reset by createJob). Full workspace suite green apart from pre-existing environment-gated integration specs (Redis cache, MCP OAuth, YouTube).
  • api (npx jest): new steer.spec.js controller guard ladder (400/403/404/409/413/429/501, sanitation, 202 shape); extended formatAgentMessages.spec.js (steer between tool steps, text flush, trailing steer); full HITL regression set (askUserQuestion.e2e, hitlCheckpoint.e2e, resume, jobReplacement, request.resumeMetadata) — 106 tests pass; full suite has zero failing tests.
  • client (npx jest): new suites for applySteerPart/findSteerMessageIndex, the useQueueDrain matrix (completed/aborted/error/interrupt-armed/double-fire/new-convo migration), the useSteering action matrix (preference, approval-pause degrade, POST error fallbacks), the in-thread PendingSteers slot (filtering, user-message presentation, attachments), and useAbortCleanup (clear vs keep-replacement matrix); SSE suites extended for on_steer_applied replay paths — 3,163+ tests pass (the single failure, ConversationsSection.spec, also fails on dev). Typecheck clean.
  • Playwright (npx playwright test --config=e2e/playwright.config.mock.ts steering.spec.ts): 3/3 green twice consecutively (queue and interrupt flows against their final contracts), chat.spec.ts canary green. The steer flow's applied-at-boundary assertions were flipped in with the 3.2.63 pin; a local re-run against the pinned SDK is the one outstanding verification.
  • Verified packages/api typechecks against the pre-steering 3.2.61 dist and the pinned 3.2.63, and that isSteeringSupported() returns false/true respectively.
  • Recommended manual pass: agent with a slow tool → steer mid-batch (message appears in-thread immediately, injects right after the in-flight tool batch, model acknowledges next turn), regenerate follow-up (steer replays as user message), queue two messages (sequential auto-send), interrupt & send (follow-up streams live), Stop with pending steers (converted to queued rows), reload mid-run (in-thread steers re-seed from resumeState).

Test Configuration:

  • Node v24, MongoDB via mongodb-memory-server, Redis 7 local (USE_REDIS=true for the integration block), published @librechat/agents 3.2.63 installed; Playwright mock harness (fake model + stdio MCP fixture) for the e2e flows.

Checklist

  • My code adheres to this project's style guidelines
  • I have performed a self-review of my own code
  • I have commented in any complex areas of my code
  • My changes do not introduce new warnings
  • I have written tests demonstrating that my changes are effective or that my feature works
  • Local unit tests pass with my changes

Loading
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants