Skip to content

Commit ec172ee

Browse files
Abort stall after post-tool and post-compact silence (#1095)
* fix(tui): abort stall after post-tool and post-compact silence Awaiting-response with a null stream is no longer a permanent stall exemption. Execution-watchdog-exempt tools stay bounded by stall timeout instead of pinning forever. Verification: bun run check (7664 pass). * fix(tui): abort stall while a sibling collect is still in flight A sibling tool.done clears currentToolName while shell_collect stays in activeToolCalls, so keying only the last name left that poll unbounded. * docs(tui): qualify stall abort cannot-pin-forever claim Post-tool and post-compact silence still abort. Concurrent same-name shell_collect can still lose the stall bound via callIdByName collision, so docs must not promise an absolute cannot-pin-forever.
1 parent 9bdb9d2 commit ec172ee

8 files changed

Lines changed: 251 additions & 95 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -335,11 +335,11 @@ Data-only agent plugins (`src/plugins/data-only-agent.ts`) synthesize `agentPlug
335335

336336
### System Prompt (`src/agent/prompts.ts`)
337337

338-
The primary session identity is **Skywalker** (`buildChatRole` → `createSkywalkerSystemPrompt`). Product name remains Corbits Code; when asked its name, the primary answers Skywalker. Role: orchestrate — classify, DIY tiny/single-file/one-route product edits, dispatch closed directors via `spawn_agent`/`wait_agents` for substantial work, track the fleet, synthesize. Product mutation tools (`write_file` / `edit_file` / `delete_file`) are mounted on the primary session (CORE and `SKYWALKER_TOOLS`) so Skywalker can DIY bounded edits; spawn remains the default for substantial, multi-file, parallel, or specialist work. Shell file-writes stay denied by auto-shell policy. MCP tools are not re-filtered by a product-write deny list (that list is gone). There is no static per-worker write-path lock; concurrent lanes sharing a cwd are instead flagged (not blocked) as a `conflict` intervention. A frontier model already knows how to code; the static prompt carries harness-specific facts and the closed-fleet orchestration policy. The base is three individually-exported sections:
338+
The primary session identity is **Skywalker** (`buildChatRole` → `createSkywalkerSystemPrompt`). Product name remains Corbits Code; when asked its name, the primary answers Skywalker. Role: orchestrate — classify, DIY tiny/single-file/one-route product edits, dispatch closed directors via `spawn_agent` (mailbox mail on TUI; `wait_agents` on exec primary) for substantial work, track the fleet, synthesize. Product mutation tools (`write_file` / `edit_file` / `delete_file`) are mounted on the primary session (CORE and `SKYWALKER_TOOLS`) so Skywalker can DIY bounded edits; spawn remains the default for substantial, multi-file, parallel, or specialist work. Shell file-writes stay denied by auto-shell policy. MCP tools are not re-filtered by a product-write deny list (that list is gone). There is no static per-worker write-path lock; concurrent lanes sharing a cwd are instead flagged (not blocked) as a `conflict` intervention. A frontier model already knows how to code; the static prompt carries harness-specific facts and the closed-fleet orchestration policy. The base is three individually-exported sections:
339339

340340
- `buildChatRole` — Skywalker primary identity (orchestrate; DIY tiny/bounded product edits; spawn for substantial work).
341341
- `buildHarnessFacts` — the non-derivable rules: shell file-writes are blocked, path tools are the DIY surface on primary (spawn builder/docs directors for substantial work), dependency installs and off-limits paths need approval, images are native multimodal input, core tools plus the advertised catalog (including `skill_search`) are resident (MCP and other unadvertised tools load via `tool_search`; use `search_agents` before dispatching specialists), workflows run only from slash-command steps, and session memory lives at `.corbits/MEMORY.md`.
342-
- `buildGuidelines` — be concise, prefer `spawn_agent`/`wait_agents` for substantial product work, DIY tiny/bounded edits on the parent, answer questions and diagnose visual/product feedback before editing, work autonomously for explicit coding tasks, use `lsp` for symbol work, and require implementation agents to run the repository-defined typecheck, relevant tests, and every defined full verification command. Agents report exact commands, outcomes, and exit statuses; a repository with no typecheck command produces an explicit Blocker backed by project-configuration evidence rather than an invented command or silent skip.
342+
- `buildGuidelines` — be concise, prefer `spawn_agent` (mailbox mail on TUI; exec-primary `wait_agents`) for substantial product work, DIY tiny/bounded edits on the parent, answer questions and diagnose visual/product feedback before editing, work autonomously for explicit coding tasks, use `lsp` for symbol work, and require implementation agents to run the repository-defined typecheck, relevant tests, and every defined full verification command. Agents report exact commands, outcomes, and exit statuses; a repository with no typecheck command produces an explicit Blocker backed by project-configuration evidence rather than an invented command or silent skip.
343343
- `buildPromptDisciplineBlock` — a shared, prohibition-form section appended exactly once to every built prompt (chat and sub-agent, every provider family). Primary vs worker wording differs for product writes: workers are told to use `read_file`/`edit_file`/`write_file`; Skywalker is told to DIY tiny/bounded edits with those path tools and spawn directors for substantial work. Shared rules: never `cat`/`sed`/heredoc/`echo` for file work, no setting or exporting environment variables (recurring needs belong in project settings), `web_fetch`/`web_search` instead of `curl`/`wget`/hand-rolled queries, one operation per `run_shell` call, turn semantics (a tool-less reply is the final answer, no repeat searches, stop and change approach after three failed attempts, batch independent reads in parallel), and TTY output rules (short bold headers, one-line bullets, backticks for paths/commands, no wide tables).
344344

345345
**Provider-conditional residuals.** Per-family additions layer on top of the shared block via the same `ModelFamilyPolicy` mechanism the directors use (`src/subagent/provider-family.ts`, `src/agent/model-family-policy.ts`) — additive lines, never prompt forks. **Grok** leaves get `buildGrokLeafAntiThrashNote` (gated by `shouldApplyGrokAntiThrash` / `applyGrokFinishBias`, withheld from orchestrators): a compact finish-bias reinforcement plus a one-line reminder to route file/web work through the dedicated tools rather than `run_shell`, motivated by observed tool-routing thrash on the same harness. **Kimi** intentionally has no residual yet — `detectModelFamily` already resolves the family so callers can branch on it, but the prompt seam is left unfilled pending eval characterization of Kimi's behavior, mirroring the provisional (permissive-default) policy in `model-family-policy.ts`.
@@ -421,7 +421,7 @@ tool call
421421

422422
**Approval log** (`src/permission/approval-log.ts`, CL-5666): every consequential decision the gate makes — auto-mode allow/deny or an interactive prompt's allow-once/allow-with-scope/deny/timeout/abort — is appended as one JSONL record to `approvals.jsonl` in the session dir, carrying the classifier/auto-shell rule name that fired (the existing `auto-shell-policy.ts`/`classify.ts` rule names, plus a small closed set of additional fixed literals the log itself defines for decisions those modules don't otherwise name — `auto-allowed-tool`, `non-interactive`, `mega-chain` — never model- or user-authored text), whether the decision was `auto` or `interactive`, a shell chain's segment count, and queued/displayed/settled timestamps. `displayedAt` is set by `PermissionRequest.markDisplayed`, called from `gate-wire.ts`'s `open()` the moment a request actually reaches the overlay host — distinct from when it was raised, so the gap it exposes is the CL-5664 signal (a queued gate arming its timeout before the operator could see it). No command text, file content, path, credential, or other free text is ever recorded — only tool name, rule, mode, segment count, and timing; a sub-agent's free-text dispatch label is deliberately left out, even though it would enable a per-agent breakdown, because nothing constrains what a model puts in it. A hard size cap on the serialized line is defense in depth against a future field reintroducing free text. Writes are fire-and-forget and swallow their own errors; the log defaults to a no-op so nothing depends on it being wired. `scripts/approval-forensics.ts` aggregates across local sessions the same way `intervention-forensics.ts` does for stop/nudge events: per-tool counts by outcome and mode, duration/display-delay percentiles, mega-chain counts, and a duplicate-rate proxy (sessions that hit the same rule more than once).
423423

424-
**Tool wall-clock budget vs. permission prompts.** Each tool `run()` is wrapped by an outer execution watchdog (`src/tui/tool-execution-watchdog.ts`). The watchdog arms when Settings set `tools.timeoutMs` / `tools.maxTimeoutMs`, when foreground `run_shell` has an effective timeout (120s default or per-call, plus 1000ms slack so this layer cannot beat shell-guard), or for `mcp__*` calls. Fleet wait tools are exempt, regardless of Settings: the generic per-tool budget never aborts a sub-agent run while the parent waits for it. That exemption is unconditional, not because the worker is otherwise bounded — there is no turn budget; `deadlineMs` is opt-in, and there is no no-progress or thrash stop. A stuck worker that never trips stall or deadline runs until the parent cancels it or, in eval mode, `--agent-timeout-ms` bounds it. `spawn_agent`, `wait_agents`, `ask_director`, `shell_collect`, and `run_shell` with `background: true` stay exempt. By default (`tools.waitForApproval`, Settings → Tools, **On**), an armed budget freezes while the operator is deciding on a permission prompt, so a late approve still runs the tool and the agent waits for the decision instead of timing out under the modal. When **Off**, the budget keeps ticking during the prompt; if it expires first the tool is skipped and the permission modal is dismissed via the budget AbortSignal (auto-deny with a timeout message). The TUI permission queue (`src/tui/gate-wire.ts`, backed by `src/permission/queue.ts`) attaches that signal so ghost prompts cannot outlive an already-aborted tool.
424+
**Tool wall-clock budget vs. permission prompts.** Each tool `run()` is wrapped by an outer execution watchdog (`src/tui/tool-execution-watchdog.ts`). The watchdog arms when Settings set `tools.timeoutMs` / `tools.maxTimeoutMs`, when foreground `run_shell` has an effective timeout (120s default or per-call, plus 1000ms slack so this layer cannot beat shell-guard), or for `mcp__*` calls. Fleet wait tools are exempt, regardless of Settings: the generic per-tool budget never aborts a sub-agent run while the parent waits for it. That exemption is unconditional, not because the worker is otherwise bounded — there is no turn budget; `deadlineMs` is opt-in, and there is no no-progress or thrash stop. A stuck worker that never trips stall or deadline still runs until the parent turn's stall budget fires, the operator cancels, or, in eval mode, `--agent-timeout-ms` bounds it. `spawn_agent`, `wait_agents`, `ask_director`, `shell_collect`, and `run_shell` with `background: true` stay exempt from the per-tool watchdog. On TUI, an in-flight `shell_collect` or `ask_director` is still bounded by the stall watchdog — including after a tool batch resolves and after compact continuation re-entry — so a typical poll does not pin the turn. Concurrent same-name `shell_collect` calls still collide in `callIdByName` (one slot per name); a sibling collect finishing can drop the mapping and the remaining poll can lose that stall bound. That residual lives in the name-keyed tracker (`src/tui/turn-state.ts`), not in the post-tool / post-compact abort itself. `wait_agents` is exec-primary only and is not mounted on TUI. By default (`tools.waitForApproval`, Settings → Tools, **On**), an armed budget freezes while the operator is deciding on a permission prompt, so a late approve still runs the tool and the agent waits for the decision instead of timing out under the modal. When **Off**, the budget keeps ticking during the prompt; if it expires first the tool is skipped and the permission modal is dismissed via the budget AbortSignal (auto-deny with a timeout message). The TUI permission queue (`src/tui/gate-wire.ts`, backed by `src/permission/queue.ts`) attaches that signal so ghost prompts cannot outlive an already-aborted tool.
425425

426426
`mcp__*` tool calls are the exception to "arms only when Settings set it": they arm unconditionally with a 5-minute default (`DEFAULT_MCP_TOOL_TIMEOUT_MS`), overridable via `mcp.timeoutMs` and still capped by `tools.maxTimeoutMs` (CL-6895). Nothing else bounds an MCP call — the stall watchdog treats an in-flight tool as activity by design, so a wedged MCP server previously hung a tool call, and the turn, forever. On expiry the call returns a normal tool-error result ("MCP tool `<name>` timed out after `<n>`s — the server may be wedged; retry or continue without it"); the turn is never aborted. The MCP client itself (`src/mcp/client.ts`, wrapping `@modelcontextprotocol/sdk`) multiplexes concurrent requests over one connection by JSON-RPC message id with no serial queue or mutex in our code or in the vendored SDK's `Protocol.request()` — so concurrent calls to the same server are not expected to deadlock each other. Live forensics for CL-6895 showed multi-minute MCP calls that eventually completed successfully, consistent with a slow server response rather than a client-side deadlock.
427427

‎docs/IMPLEMENTATION.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -204,7 +204,7 @@ Listings are list-free, dumps are dump-locked: a bounded `ls`/`tree` prints name
204204

205205
`ChatInputProps` carries `isProcessing?: boolean` and `onInterrupt?: (message: string) => void`. When `isProcessing` is true, drain timing is **parent-idle** vs **session-idle**:
206206

207-
- **Enter** soft-steers while the parent is busy — enqueues kind `"steer"` and delivers at the next **parent** `tool.boundary` (the parent tool finishing, not a child). Does not interrupt. **Parent-idle** is when the primary Skywalker turn is not inside an in-flight parent tool; a long parent **foreground** `run_shell` or awaiting `wait_agents` is parent-busy and holds steers. A `run_shell` started with `background: true` returns at once and releases the boundary; its completion is delivered as a system message (`buildShellBackgroundMessage`, mailbox `system`, no operator-originated flag — it re-enters the reactor without counting as operator input) on a later turn.
207+
- **Enter** soft-steers while the parent is busy — enqueues kind `"steer"` and delivers at the next **parent** `tool.boundary` (the parent tool finishing, not a child). Does not interrupt. **Parent-idle** is when the primary Skywalker turn is not inside an in-flight parent tool; a long parent **foreground** `run_shell` is parent-busy and holds steers. A `run_shell` started with `background: true` returns at once and releases the boundary; its completion is delivered as a system message (`buildShellBackgroundMessage`, mailbox `system`, no operator-originated flag — it re-enters the reactor without counting as operator input) on a later turn.
208208

209209
#### Background shell mode
210210

@@ -288,7 +288,7 @@ Provider and model configuration lives in JSON settings files. The global file h
288288
}
289289
```
290290

291-
- `timeoutMs` / `maxTimeoutMs` — outer execution watchdog around each tool `run()`. Unset leaves the generic watchdog unarmed; set these to arm it. `maxTimeoutMs` clamps non-shell tools when set and does not cap a longer requested `run_shell`. Foreground `run_shell` always arms at the effective shell timeout (120s default, `settings.shell.timeoutMs` override, or per-call) plus 1000ms slack. Fleet wait tools are exempt: a dispatched sub-agent is bounded by stall, opt-in `deadlineMs`, and operator cancel, not the generic per-tool budget. Background shell is exempt too: a `run_shell` with `background: true` arms nothing (only a per-call timeout bounds the process; there is no 120s default) and `shell_collect` never arms (a bounded poll over a process that outlives the turn).
291+
- `timeoutMs` / `maxTimeoutMs` — outer execution watchdog around each tool `run()`. Unset leaves the generic watchdog unarmed; set these to arm it. `maxTimeoutMs` clamps non-shell tools when set and does not cap a longer requested `run_shell`. Foreground `run_shell` always arms at the effective shell timeout (120s default, `settings.shell.timeoutMs` override, or per-call) plus 1000ms slack. Fleet wait tools are exempt: a dispatched sub-agent is bounded by stall, opt-in `deadlineMs`, and operator cancel, not the generic per-tool budget. Background shell is exempt too: a `run_shell` with `background: true` arms nothing (only a per-call timeout bounds the process; there is no 120s default) and `shell_collect` never arms the per-tool watchdog (a bounded poll over a process that outlives the turn). An in-flight TUI `shell_collect` or `ask_director` is still bounded by the TUI stall watchdog, including after a tool batch resolves and after compact continuation re-entry, so a typical poll does not pin the turn. Concurrent same-name `shell_collect` calls still collide in `callIdByName` (one slot per name); a sibling collect finishing can drop the mapping and the remaining poll can lose that stall bound — a residual of the name-keyed tracker, not a hole in the post-tool / post-compact abort. `wait_agents` is exec-primary only and is not mounted on TUI.
292292
- `waitForApproval` (default **true** when unset) — freeze that budget while a permission prompt is open so a late approve still runs the tool. **Settings → Tools** toggles this live for the next tool call and persists it here. When **false**, the budget keeps ticking during the prompt; on expiry the tool is skipped and the modal is auto-dismissed. The freeze is bounded: after **30 minutes** with the prompt still unanswered the budget resumes ticking on its own, so a prompt that never becomes visible (overlay open, UI gone) cannot hang a tool run indefinitely.
293293

294294
Optional `mcp` block bounds MCP tool calls (`mcp__*` names) specifically — unlike `tools.*`, this arms **unconditionally** even with no settings at all, defaulting to **5 minutes**, since a wedged MCP server otherwise hangs a call forever with nothing to bound it (CL-6895):

0 commit comments

Comments
 (0)