Repository navigation
🐛 Keep the control connection responsive while a broker ask waits - #247
Conversation
…nect (#243) An unanswered broker_ask blocked the control connection's reader for up to the broker's 10-minute ask timeout, so every later request on that connection (status, spawn, peers, follow_up) timed out with "read response: timeout" while the child kept running. - dispatch broker_ask on a background goroutine bound to a per-connection context; connection loss and server stop cancel the pending wait - BrokerAsk takes a context, sends wait_reply with a request context, and withdraws a cancelled ask via /cancel_message using a fresh context - broker_ask failures return error data with the ask's message_id so the caller can cancel it instead of retrying blind; optional timeout_ms param bounds the wait - /wait_reply extends its socket write deadline (new configurable writeTimeout field); the previous fixed 30s WriteTimeout could drop a reply that arrived after the deadline but before the long-poll timeout - Go and TS control clients preserve structured RPC error data (RPCError / RpcError); the ask tool surfaces the message_id for avenor_cancel Fixes #243 AI-Generated-By: glm-5.3-flash, sparky/qwen3.8:27b
There was a problem hiding this comment.
This PR is marked... OUT! ✊
Blocking
timeout_ms parameter silently capped at 30s on every ask path
Files: packages/core/src/tools/ask.ts:45, packages/core/src/supervisor.ts:160/299, packages/core/src/tools/get-supervisor-client.ts:24, packages/core/src/client.ts:218/315-318
The new timeout_ms parameter is silently capped at 30s on every ask path. The singleton client is dialed with Supervisor's default callTimeoutMs = 30_000 (supervisor.ts:160), the external path uses plain dial(id) with no opts (get-supervisor-client.ts:24 → client.ts:218 defaults to 30s), and Client.call() unconditionally arms a 30s timer (client.ts:315-318). An agent passing timeout_ms: 120000 to avenor_ask fails at 30s with a generic "read response: timeout" (a plain Error, not an RpcError) while the server keeps processing the ask and later discards the out-of-order reply. The advertised timeout semantics never hold and the failure message misrepresents the configured bound.
Fix: Thread timeoutMs through Client.call() as a per-call override (e.g. setTimeout(..., Math.max(this.callTimeout, timeoutMs)) for ask-shaped calls), or at minimum have brokerAsk reject/validate timeoutMs > callTimeout so the mismatch is explicit instead of silent.
…rrors - broker test: set writeTimeout before Start so the server is actually created with the 100ms deadline and SetWriteDeadline is load-bearing - control server: skip the response frame for JSON-RPC notifications (req.ID == nil) on the async broker_ask path - AskError carries pending: false when the ask edge is provably cleaned up (send failure, broker-side cleanup, already-deleted edge) and true only when cleanup failed; the ask tool's cancel guidance now matches - TS client: per-call timeout override so timeout_ms above the default 30s no longer fails client-side first Refs #247 AI-Generated-By: glm-5.3-flash, sparky/qwen3.8:27b
…imeouts - withdrawAsk helper classifies ask-edge state on every failure path (send and wait): pending=false when the broker cleaned the edge or reports it gone; true only when cleanup itself fails - broker: SetWriteDeadline failure is non-fatal instead of a 500 that orphaned the registered edge; fix the writeTimeout doc comment - control server: reject negative timeout_ms, cap it at the broker's 10-minute ask ceiling, and bound concurrent in-flight asks per connection (8) instead of accumulating unbounded parked goroutines - broker test: send before starting the replier poll loop and continue on transient poll errors; tests for the classification and dead-broker send path - control tests: pin the asyncResponse guard (no double write), the notification contract (id-less ask gets no frame), the server-side timeout_ms deadline, Stop cancellation, and the in-flight cap - TS: ignore non-positive timeoutMs, clamp the per-call timer to setTimeout's range, and make the timer monkeypatch deterministic Refs #247 AI-Generated-By: glm-5.3-flash, sparky/qwen3.8:27b
…afe test fixtures - blockingAskHandler: sync.Once for the entered signal and a mutex around askCtx, safe under the in-flight cap test's 8 concurrent asks - clamp the server-side timeout_ms to just below the broker's DefaultAskTimeout so clamped values expire server-side with structured error data instead of racing the broker's 504 - classify ask-edge cleanup on the 404 status, not just the body text; a refused send connection (request never delivered) reports pending false - TS client: clamp the final computed timer value, not just the input - tests: strict first-frame reader pins the asyncResponse guard; the notification test completes the handler before asserting silence; the timeout test uses a 500ms window AI-Generated-By: glm-5.3-flash
|
Addressed 4 inline comments across 3 commits (ffd1236, 466abd5, 61d62b4), all as fixes:
Blocking verdict — Internal adversarial review after the fixes found and fixed further issues: send-path edge cleanup with a Review loop: 2 iterations. Deferred as accepted: per-call timeout for the Go @umpire-bot review again |
|
Round 2 addressed 3 inline comments in commit 799ed28, all as fixes:
Review loop: 3 iterations. All 7 review threads resolved; no remaining findings. Go suites pass; |
|
@umpire-bot review again |
- declare avenor_ask's timeout_ms as Type.Integer so fractional values cannot reach the int64 server parser - rewrite the duplicated per-call-timeout test so it covers the branch where the base callTimeout exceeds timeoutMs + margin - cover the send-failure cleanup path end-to-end: a 404 target on a live broker runs withdrawAsk and classifies pending=false - assert data["pending"] in the ask-error wire tests (both false and true variants) - add an avenor_ask wiring test for timeout_ms -> timeoutMs AI-Generated-By: glm-5.3-flash, sparky/qwen3.8:27b
|
Round 3 addressed 5 inline comments in commit 49a4900, all as fixes:
Review loop: 3 iterations. All 12 threads resolved; no remaining findings. Go: control/stable pass; |
|
@umpire-bot review again |
…tests - the write-timeout test now records and asserts the wait_reply elapsed time exceeds the deadline, so a reply that lands before the handler parks fails loudly instead of passing vacuously - collapse the three timeout-normalization tests into one it.each case - the delay-capture assertions stay: the +10s margin is a documented contract, and the behavioral case is covered by the slow-response test AI-Generated-By: glm-5.3-flash
|
Round 4 addressed 3 inline comments in commit e196f4d:
Review loop: 4 iterations. 15/15 threads resolved; no remaining findings. Go: broker suite passes; |
|
@umpire-bot review again |
|
No new commits since the last review — skipping this re-review. Latest verdict: #247 (review) |
An unanswered
broker_askblocked the read loop on the control connection for up to the broker's 10-minute ask timeout. While the child kept running, every later request on that connection returnedread response: timeout. Dispatch recovered on a second supervisor.Changes
internal/control:broker_asknow runs on a background goroutine bound to a per-connection context. The connection keeps serving requests such asstatus,spawn,broker_peers, andfollow_upwhile an ask waits. Disconnect and server stop cancel the pending wait.internal/stable:BrokerAsktakes a context and sends/wait_replywith a request context. A cancelled wait withdraws the pending ask via/cancel_messagewith a fresh context. A late child reply therefore cannot pair with a dead waiter.broker_askfailures return error data with the ask'smessage_id. Callers can cancel the pending ask withbroker_cancelinstead of retrying blind. A new optionaltimeout_msparameter bounds the wait.internal/runtime/broker:/wait_replynow extends its socket write deadline through a new configurablewriteTimeoutfield. The previous fixed 30sWriteTimeoutrisked dropping a reply that arrived after the deadline but before the long-poll timeout.client(Go) andpackages/core(TS): RPC errors preserve structureddata(RPCError/RpcError). Theavenor_asktool returns themessage_idin the error text so the agent can withdraw the ask withavenor_cancel. It accepts optionaltimeout_ms.Non-ask handlers remain serial. Two hazards rule out per-request goroutines elsewhere. The
promptandshutdownhandlers rely on arrival-order semantics.ensureOwnercan also reassign ownership after a disconnect. Those handlers stay out of scope.Follow-up work (out of scope here)
run_idin durable event logs.avenor_follow_upadvertises resume that the Pi backend cannot support (--no-session).Validation
go test ./internal/control/ ./internal/stable/ ./internal/runtime/broker/ ./client/— okgo test -raceoninternal/control,internal/runtime/broker,internal/stable— ok, no racesbun testinpackages/core: 320 pass.packages/pi: 139 pass.statuswhile an ask is blocked. Disconnect cancels the in-flight ask. Ask errors carrymessage_id. A reply arriving after the old write deadline still reaches the waiter.cmd/avenorhas 26 pre-existing failures on a clean tree (macOS unix-socket path length in$TMPDIR); unchanged by this branchFixes #243