Repository navigation
Restore the WebSocket, stop the 5s repaint, share one launch roster (v0.5.0) - #7
Merged
Merged
Conversation
added 2 commits
August 27, 2026 10:32
…v0.5.0) Three defects, grouped because they share two files and one deploy cycle. WebSocket (vikunja#532). v0.4.0 tightened the upgrade guard to reject a *missing* Origin as well as a wrong one, reasoning that only non-browser clients omit it. The one non-browser client here is CloudCLI's own plugin WS proxy, which uses the `ws` library — and that sends no Origin unless explicitly passed. Every handshake was 403'd from 2026-08-02 to 2026-08-27: 2239 error lines, and a tab reading "disconnected" throughout. The guard now gates on the peer (loopback only; anything else refused outright) and applies the Origin allowlist only when an Origin is present. A present-but-wrong Origin is still refused, so this is a narrowing, not a relaxation — the loopback bind was always the real boundary. Extracted to ws-guard.ts as a pure function with tests, since server.ts listens at import time. AGENTS.md carried the old rule as a project invariant. Left there, the next reviewer restores the defect on principle, so that sentence is rewritten rather than merely contradicted by the code. Repaint (vikunja#532). Each failed reconnect emitted _disconnected, which called render() — and render() does root.innerHTML = ''. The whole panel was torn down and rebuilt every 5s while disconnected, losing scroll position and closing any open filter dropdown, even though wsConnected was already false. Connection state no longer reaches render(): the header badge is mutated in place, and only on a genuine transition. Reconnect backoff added at 5s -> 10s -> 30s, capped, reset on open; the first delay is unchanged so a transient blip recovers as fast as before. Launcher (vikunja#523). AGENT_PROJECTS was a second, drifted copy of task-dispatcher.py's roster with no steward entry, so Start refused steward outright. The literal is deleted rather than extended: adding an entry would have made launchSession() spawn `claude` as the plugin's own user, bypassing the launcher whose whole purpose is that agent's isolation — a session appearing as steward in every log while holding none of steward's credentials. Both consumers now read ~/scripts/agent-launch.yml. Validation re-establishes what the literal provided for free: every field checked against a closed set, whole document rejected on any violation, and a missing or malformed file is a named error rather than an empty policy — an empty policy makes run_as_user absent for every agent, which is exactly the impersonation above. An agent with run_as_user is launched via `sudo -n -u <user> <launcher> --workflow-mode <mode> -- <prompt>`; a missing or non-executable launcher is refused by name, never falling back. Also: - Launch logs move to ~/.claude/comms/artifacts/task-launches/<agent>-<task8>.log, matching the dispatcher. Two destinations for one concept meant nothing could list "the launches". - Start's `review` maps to the queue's `semi-auto` rather than passing through, which the launcher would refuse by name. For a run-as agent `review` is prompt-enforced only — run-steward.sh sets --dangerously-skip-permissions itself and accepts no permission mode — and the toast now says so instead of implying a tool gate. - lookupAgent() guards the policy lookup: target_agent comes from a queue YAML, and a plain policy['constructor'] is truthy, so a bare !entry check would accept a non-entry. Policy objects are built with a null prototype. - A refused upgrade logs why. v0.4.0's refusals were silent on this side; the only signal was a 403 on the far side of the proxy, naming neither leg nor the cause. Verified against a live instance of the new binary on a scratch port: no Origin -> 101, allowed Origin -> 101, wrong Origin -> 403. Negative test confirmed a broken launcher fails by name with nothing spawned. Tests 12 -> 41; 13 mutations applied across the three modules, all caught. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> agent-id: developer
Audit of task-queue-plugin-repair-2026-08 found the two hand-written launch-policy validators computed their containment root differently: Python called .resolve() on it, this side uses a plain path join. Neither resolves the CANDIDATE project_dir, so resolving only the root compares a canonical path against an uncanonical one. The fix is on the Python side (the .resolve() is removed there). What changes here is the comment: the plain join is now stated as a rule the other implementation must match, rather than being an unexplained coincidence that the next edit could break. Reproduced before fixing rather than taken on trust — with ~/.claude/projects made a symlink, the same document was REJECTED by Python and ACCEPTED here. After the fix both return the identical unresolved path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> agent-id: developer
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three defects, grouped because they share two files and one deploy cycle.
Covers vikunja#532 and vikunja#523.
1 — The WebSocket has been dead since v0.4.0 (#532)
v0.4.0 tightened the upgrade guard from
if (origin && !allowed.includes(origin))toif (!allowed.includes(origin)), so a handshake with noOriginis rejected. The statedreasoning was that only non-browser clients omit
Origin.On this deployment the only non-browser client is CloudCLI's own plugin WS proxy, which uses
the
wslibrary — and that sends noOriginunless one is passed. Every upstream handshakewas 403'd from
2026-08-02T17:45(the v0.4.0 deploy) to today: 2239WS proxy error … 403lines, and a tab readingdisconnectedthroughout.The guard now gates on the peer — loopback only, anything else refused outright — and
applies the
Originallowlist only when anOriginis actually present. A present-but-wrongOriginis still refused.This is a narrowing, not a relaxation. The loopback bind on an ephemeral port was always the
real boundary;
Originwas never doing the work this deployment needed. Two independentcontrols remain either way: CloudCLI's
verifyClientauthenticates the browser leg beforehandlePluginWsProxyis ever called, and the server binds127.0.0.1.AGENTS.mdcarried the defect as a project invariant — "The WebSocket upgrade rejects amissing Origin, not just a wrong one." Contradicting that in code alone would invite the next
reviewer to restore the bug on principle, so the sentence is rewritten with the loopback rule
and the history.
New
src/ws-guard.tsholds the decision as a pure function;server.tscallslisten()atimport time, so this is the same extraction rationale as
control-api.ts.Verified against a live instance of the new binary on a scratch port:
Origin(what the proxy sends)101 Switching ProtocolsOrigin: http://127.0.0.1:3001101 Switching ProtocolsOrigin: http://localhost:3001101 Switching ProtocolsOrigin: http://evil.example403 Forbidden2 — The five-second full-panel repaint (#532)
ws-client.tsemitted_disconnectedfromoncloseand reconnected every 5000ms forever.index.tshandled it by settingstate.wsConnected = falseand callingrender()unconditionally — and
render()doesroot.innerHTML = ''.So the entire panel was torn down and rebuilt every 5 seconds while disconnected, losing
scroll position and closing any open filter dropdown, even though
wsConnectedwas alreadyfalse.Connection state no longer reaches
render()at all.updateConnectionBadge()mutates theheader's dot colour and label in place, and only on a genuine transition. Reconnect backoff
added: 5s → 10s → 30s, capped, reset on a successful open — the first delay is unchanged, so a
transient blip still recovers as fast as before; the widening is for outages. A fixed 5s retry
is what turned a three-week outage into 2239 identical log lines.
3 — Start could not launch a run-as agent (#523)
AGENT_PROJECTSwas a hardcoded five-agent map with nostewardentry, solaunchSession()refused it. Adding an entry would have been worse than the bug:
launchSession()spawnedclaudedirectly as the plugin's own user, so a "fixed" map would have started a session insteward's project dir, as the wrong user, with none of steward's credentials — a session
appearing as steward in every log while holding nothing of steward's. The launcher's identity
guard never fires when the launcher is bypassed.
The literal is deleted, not extended. Both this plugin and
task-dispatcher.pynow readone file,
~/scripts/agent-launch.yml(override:AGENT_LAUNCH_POLICY). Two rosters of onefact is exactly what drifted into this ticket.
Validation re-establishes what the literal gave for free — nothing user-supplied reached
spawn. Every field is checked against a closed set (agent name shape,project_dirunder~/.claude/projectswith..normalised first,run_as_usermatchingagent-*,launcherunder
/usr/local/sbin/forge/), and the whole document is rejected on any violation. Aloader that skipped bad entries would pass nearly every test while silently dropping an agent
from the queue's reach.
A missing or malformed file is a named error, never an empty policy: an empty policy makes
run_as_userabsent for every agent, which is precisely the impersonation above. A missingor non-executable launcher is refused by name — there is no fallback path, deliberately.
Negative test, live: pointing steward's
launcherat a nonexistent path returnsLauncher missing or not executable for run-as agent 'steward': … — deploy it with forge-scripts-deploy.sh, and nothing is spawned.The
sudoargv the plugin now builds was exercised end to end against the real launcher withan empty prompt, so the launcher's own guard fires before it reaches
claude: sudo acceptsthe argv and the identity guard passes.
No new sudoers grant. The plugin runs as ted, and ted already holds
(agent-steward) CWD=* NOPASSWD: /usr/local/sbin/forge/run-steward.sh. This change makes theplugin stop bypassing that launcher, which is a net tightening.
Two things reviewers should know
Mode vocabulary drift. Start sends
review | auto;run-steward.shandtask-queue-mcp'sVALID_WORKFLOW_MODEStakesemi-auto | auto | manual-then-auto.reviewis mapped explicitly to
semi-autorather than passed through — passed through, the launcherrefuses it by name (confirmed live:
FATAL: invalid --workflow-mode: review). vikunja#533covers the wider unification.
For a run-as agent,
reviewis prompt-enforced only:run-steward.shsets--dangerously-skip-permissionsitself and accepts no permission mode, so--permission-mode planis not reachable. Rather than let the toast imply a tool gate that is not there, thebackend returns a
noteand the UI shows it. Filed as a ticket rather than widening thelauncher's argv surface in this build.
Prototype-chain lookup (found in pre-audit baseline).
target_agentcomes from a queueYAML, and
policy['constructor']on a plain object is truthy — a bareif (!entry)guardwould carry a non-entry into the launch path. It failed safe at the next check, but the
ordering was load-bearing by accident. Policies are now built with a null prototype and looked
up via
lookupAgent().Also
~/.claude/comms/artifacts/task-launches/<agent>-<task8>.log, matchingthe dispatcher. It wrote
~/.pm2/logs/agent-launch-*while this wrote<taskId>.log— twodestinations for one concept.
~/.claude/commsis the side both can read;~/.pm2/logsisnot in
PREVIEW_ALLOWED_PREFIXESand must not be added, since it covers every PM2 servicelog on the host. vikunja#534 depends on this.
was a 403 on the far side of the proxy, naming neither leg nor the cause.
env:CLOUDCLI_ORIGINadded to manifest permissions.package.jsonandmanifest.json.Companion changes
This half alone restores service. Two companion PRs land the other legs so neither side can
silently re-break it:
TadMSTR/claudecodeuifix/plugin-ws-origin-2026-08— the proxy sends anOrigin;CLOUDCLI_ORIGINadded toPLUGIN_ENV_ALLOWLIST; recorded inPATCHES.mdas a carriedpatch with a both-directions probe.
host-forge/scriptsfix/agent-launch-policy-2026-08—agent-launch.ymlitself, thedispatcher's loader, and
CLOUDCLI_ORIGINincloudcli.sh.Testing
npm run build && npm test— 41 passing, up from 12. New coverage: the upgrade guard (allthree cases, including the loopback-no-Origin one v0.4.0 broke), the reconnect schedule, and
the launch policy (every closed-set rejection, whole-document rejection, both argv shapes,
prototype-chain lookup).
15 mutations applied across
ws-guard.ts,ws-client.tsandlaunch-policy.ts— includingrestoring the exact v0.4.0 defect and making a run-as agent fall through to a
claudeargv —all 15 caught.
Both loaders were run against the shipped policy file and produce byte-identical rosters.