You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(server): file a papercut when the stall watchdog flags a turn #28
Part of #22. Wave 2, after #26 for useful content.
Problem
A papercut exists only if I notice a problem and tap. The stall that hurt most (#1) is already detected by the server, and the evidence is freshest at that moment. Waiting for me costs minutes of trace rotation, and a hung client cannot report at all.
What the code shows
ThreadTurnWatchdog.ts flags a running turn silent for 10 minutes (STALLED_TURN_THRESHOLD_MS), appends provider.turn.stalled, and logs orchestration.turn.stalled (around lines 84-132). It covers claudeAgent only (WATCHED_PROVIDERS) and skips turns waiting on approvals, user input, or background work.
When markStalled succeeds, create a server-originated papercut: no screenshot, no client evidence, a server value added to PapercutClientSurface, a one-line summary ("stalled turn, idle N s"), and the full server snapshot from feat(papercuts): capture the server's full picture at report time #26.
At most one per thread and turn (deterministic report id from thread, turn, and stall time). Resume then stall again on a later turn creates another.
Failure to create never affects the sweep (catchCause, as the existing append).
Test with the existing sweep harness (ThreadTurnWatchdog.test.ts): a stalled thread yields exactly one record across repeated sweeps; a new turn yields another; a watched-out provider yields none.
The record's overlay can be marked and triaged like any other.
Surfaces
Server only. Contract: one enum value. Providers: inherits the watchdog's Claude-only scope; widening it is its own decision. Other server-side triggers (provider start failure, a send that never settles) need an exhibit before they are added.
Part of #22. Wave 2, after #26 for useful content.
Problem
A papercut exists only if I notice a problem and tap. The stall that hurt most (#1) is already detected by the server, and the evidence is freshest at that moment. Waiting for me costs minutes of trace rotation, and a hung client cannot report at all.
What the code shows
ThreadTurnWatchdog.tsflags a running turn silent for 10 minutes (STALLED_TURN_THRESHOLD_MS), appendsprovider.turn.stalled, and logsorchestration.turn.stalled(around lines 84-132). It coversclaudeAgentonly (WATCHED_PROVIDERS) and skips turns waiting on approvals, user input, or background work.docs/fork/papercuts.mdstill says the server has no stalled-turn detection; stale, corrected in the decision commit for feat(papercuts): a local feedback loop from one-tap report to verified fix #22.Plan
markStalledsucceeds, create a server-originated papercut: no screenshot, no client evidence, aservervalue added toPapercutClientSurface, a one-line summary ("stalled turn, idle N s"), and the full server snapshot from feat(papercuts): capture the server's full picture at report time #26.catchCause, as the existing append).Acceptance
ThreadTurnWatchdog.test.ts): a stalled thread yields exactly one record across repeated sweeps; a new turn yields another; a watched-out provider yields none.Surfaces
Server only. Contract: one enum value. Providers: inherits the watchdog's Claude-only scope; widening it is its own decision. Other server-side triggers (provider start failure, a send that never settles) need an exhibit before they are added.
Related
#22, #26, #1, #30.