Skip to content

Restore the WebSocket, stop the 5s repaint, share one launch roster (v0.5.0) - #7

Merged
TadMSTR merged 2 commits into
mainfrom
fix/ws-repaint-launcher-2026-08
Aug 27, 2026
Merged

TadMSTR merged 2 commits into
mainfrom
fix/ws-repaint-launcher-2026-08

Conversation

@TadMSTR

@TadMSTR TadMSTR commented Aug 27, 2026

Copy link
Copy Markdown
Owner

Three defects, grouped because they share two files and one deploy cycle.
Covers vikunja#532 and vikunja#523.

1 — The WebSocket has been dead since v0.4.0 (#532)

v0.4.0 tightened the upgrade guard from if (origin && !allowed.includes(origin)) to
if (!allowed.includes(origin)), so a handshake with no Origin is rejected. The stated
reasoning was that only non-browser clients omit Origin.

On this deployment the only non-browser client is CloudCLI's own plugin WS proxy, which uses
the ws library — and that sends no Origin unless one is passed. Every upstream handshake
was 403'd from 2026-08-02T17:45 (the v0.4.0 deploy) to today: 2239 WS proxy error … 403 lines, and a tab reading disconnected throughout.

The guard now gates on the peer — loopback only, anything else refused outright — and
applies the Origin allowlist only when an Origin is actually present. A present-but-wrong
Origin is still refused.

This is a narrowing, not a relaxation. The loopback bind on an ephemeral port was always the
real boundary; Origin was never doing the work this deployment needed. Two independent
controls remain either way: CloudCLI's verifyClient authenticates the browser leg before
handlePluginWsProxy is ever called, and the server binds 127.0.0.1.

AGENTS.md carried the defect as a project invariant — "The WebSocket upgrade rejects a
missing Origin, not just a wrong one."
Contradicting that in code alone would invite the next
reviewer to restore the bug on principle, so the sentence is rewritten with the loopback rule
and the history.

New src/ws-guard.ts holds the decision as a pure function; server.ts calls listen() at
import time, so this is the same extraction rationale as control-api.ts.

Verified against a live instance of the new binary on a scratch port:

Handshake Result
no Origin (what the proxy sends) 101 Switching Protocols
Origin: http://127.0.0.1:3001 101 Switching Protocols
Origin: http://localhost:3001 101 Switching Protocols
Origin: http://evil.example 403 Forbidden

2 — The five-second full-panel repaint (#532)

ws-client.ts emitted _disconnected from onclose and reconnected every 5000ms forever.
index.ts handled it by setting state.wsConnected = false and calling render()
unconditionally — and render() does root.innerHTML = ''.

So the entire panel was torn down and rebuilt every 5 seconds while disconnected, losing
scroll position and closing any open filter dropdown, even though wsConnected was already
false
.

Connection state no longer reaches render() at all. updateConnectionBadge() mutates the
header's dot colour and label in place, and only on a genuine transition. Reconnect backoff
added: 5s → 10s → 30s, capped, reset on a successful open — the first delay is unchanged, so a
transient blip still recovers as fast as before; the widening is for outages. A fixed 5s retry
is what turned a three-week outage into 2239 identical log lines.

Noted, not fixed: render() is called on every filter change and every detail load and
always does root.innerHTML = ''. Restructuring that was out of scope here, but it is why
any future state change is a full repaint by default.

3 — Start could not launch a run-as agent (#523)

AGENT_PROJECTS was a hardcoded five-agent map with no steward entry, so launchSession()
refused it. Adding an entry would have been worse than the bug: launchSession() spawned
claude directly as the plugin's own user, so a "fixed" map would have started a session in
steward's project dir, as the wrong user, with none of steward's credentials — a session
appearing as steward in every log while holding nothing of steward's. The launcher's identity
guard never fires when the launcher is bypassed.

The literal is deleted, not extended. Both this plugin and task-dispatcher.py now read
one file, ~/scripts/agent-launch.yml (override: AGENT_LAUNCH_POLICY). Two rosters of one
fact is exactly what drifted into this ticket.

Validation re-establishes what the literal gave for free — nothing user-supplied reached
spawn. Every field is checked against a closed set (agent name shape, project_dir under
~/.claude/projects with .. normalised first, run_as_user matching agent-*, launcher
under /usr/local/sbin/forge/), and the whole document is rejected on any violation. A
loader that skipped bad entries would pass nearly every test while silently dropping an agent
from the queue's reach.

A missing or malformed file is a named error, never an empty policy: an empty policy makes
run_as_user absent for every agent, which is precisely the impersonation above. A missing
or non-executable launcher is refused by name — there is no fallback path, deliberately.

Negative test, live: pointing steward's launcher at a nonexistent path returns
Launcher missing or not executable for run-as agent 'steward': … — deploy it with forge-scripts-deploy.sh, and nothing is spawned.

The sudo argv the plugin now builds was exercised end to end against the real launcher with
an empty prompt, so the launcher's own guard fires before it reaches claude: sudo accepts
the argv and the identity guard passes.

No new sudoers grant. The plugin runs as ted, and ted already holds
(agent-steward) CWD=* NOPASSWD: /usr/local/sbin/forge/run-steward.sh. This change makes the
plugin stop bypassing that launcher, which is a net tightening.

Two things reviewers should know

Mode vocabulary drift. Start sends review | auto; run-steward.sh and
task-queue-mcp's VALID_WORKFLOW_MODES take semi-auto | auto | manual-then-auto. review
is mapped explicitly to semi-auto rather than passed through — passed through, the launcher
refuses it by name (confirmed live: FATAL: invalid --workflow-mode: review). vikunja#533
covers the wider unification.

For a run-as agent, review is prompt-enforced only: run-steward.sh sets
--dangerously-skip-permissions itself and accepts no permission mode, so --permission-mode plan is not reachable. Rather than let the toast imply a tool gate that is not there, the
backend returns a note and the UI shows it. Filed as a ticket rather than widening the
launcher's argv surface in this build.

Prototype-chain lookup (found in pre-audit baseline). target_agent comes from a queue
YAML, and policy['constructor'] on a plain object is truthy — a bare if (!entry) guard
would carry a non-entry into the launch path. It failed safe at the next check, but the
ordering was load-bearing by accident. Policies are now built with a null prototype and looked
up via lookupAgent().

Also

  • Launch logs move to ~/.claude/comms/artifacts/task-launches/<agent>-<task8>.log, matching
    the dispatcher. It wrote ~/.pm2/logs/agent-launch-* while this wrote <taskId>.log — two
    destinations for one concept. ~/.claude/comms is the side both can read; ~/.pm2/logs is
    not in PREVIEW_ALLOWED_PREFIXES and must not be added, since it covers every PM2 service
    log on the host. vikunja#534 depends on this.
  • A refused upgrade now logs why. v0.4.0's refusals were silent on this side; the only signal
    was a 403 on the far side of the proxy, naming neither leg nor the cause.
  • env:CLOUDCLI_ORIGIN added to manifest permissions.
  • Version 0.5.0 in both package.json and manifest.json.

Companion changes

This half alone restores service. Two companion PRs land the other legs so neither side can
silently re-break it:

  • TadMSTR/claudecodeui fix/plugin-ws-origin-2026-08 — the proxy sends an Origin;
    CLOUDCLI_ORIGIN added to PLUGIN_ENV_ALLOWLIST; recorded in PATCHES.md as a carried
    patch with a both-directions probe.
  • host-forge/scripts fix/agent-launch-policy-2026-08 — agent-launch.yml itself, the
    dispatcher's loader, and CLOUDCLI_ORIGIN in cloudcli.sh.

Testing

npm run build && npm test — 41 passing, up from 12. New coverage: the upgrade guard (all
three cases, including the loopback-no-Origin one v0.4.0 broke), the reconnect schedule, and
the launch policy (every closed-set rejection, whole-document rejection, both argv shapes,
prototype-chain lookup).

15 mutations applied across ws-guard.ts, ws-client.ts and launch-policy.ts — including
restoring the exact v0.4.0 defect and making a run-as agent fall through to a claude argv —
all 15 caught.

Both loaders were run against the shipped policy file and produce byte-identical rosters.

developer-agent added 2 commits August 27, 2026 10:32
…v0.5.0)

Three defects, grouped because they share two files and one deploy cycle.

WebSocket (vikunja#532). v0.4.0 tightened the upgrade guard to reject a *missing*
Origin as well as a wrong one, reasoning that only non-browser clients omit it. The
one non-browser client here is CloudCLI's own plugin WS proxy, which uses the `ws`
library — and that sends no Origin unless explicitly passed. Every handshake was
403'd from 2026-08-02 to 2026-08-27: 2239 error lines, and a tab reading
"disconnected" throughout.

The guard now gates on the peer (loopback only; anything else refused outright) and
applies the Origin allowlist only when an Origin is present. A present-but-wrong
Origin is still refused, so this is a narrowing, not a relaxation — the loopback bind
was always the real boundary. Extracted to ws-guard.ts as a pure function with tests,
since server.ts listens at import time.

AGENTS.md carried the old rule as a project invariant. Left there, the next reviewer
restores the defect on principle, so that sentence is rewritten rather than merely
contradicted by the code.

Repaint (vikunja#532). Each failed reconnect emitted _disconnected, which called
render() — and render() does root.innerHTML = ''. The whole panel was torn down and
rebuilt every 5s while disconnected, losing scroll position and closing any open
filter dropdown, even though wsConnected was already false. Connection state no
longer reaches render(): the header badge is mutated in place, and only on a genuine
transition. Reconnect backoff added at 5s -> 10s -> 30s, capped, reset on open; the
first delay is unchanged so a transient blip recovers as fast as before.

Launcher (vikunja#523). AGENT_PROJECTS was a second, drifted copy of
task-dispatcher.py's roster with no steward entry, so Start refused steward outright.
The literal is deleted rather than extended: adding an entry would have made
launchSession() spawn `claude` as the plugin's own user, bypassing the launcher whose
whole purpose is that agent's isolation — a session appearing as steward in every log
while holding none of steward's credentials.

Both consumers now read ~/scripts/agent-launch.yml. Validation re-establishes what
the literal provided for free: every field checked against a closed set, whole
document rejected on any violation, and a missing or malformed file is a named error
rather than an empty policy — an empty policy makes run_as_user absent for every
agent, which is exactly the impersonation above. An agent with run_as_user is
launched via `sudo -n -u <user> <launcher> --workflow-mode <mode> -- <prompt>`; a
missing or non-executable launcher is refused by name, never falling back.

Also:
- Launch logs move to ~/.claude/comms/artifacts/task-launches/<agent>-<task8>.log,
  matching the dispatcher. Two destinations for one concept meant nothing could list
  "the launches".
- Start's `review` maps to the queue's `semi-auto` rather than passing through, which
  the launcher would refuse by name. For a run-as agent `review` is prompt-enforced
  only — run-steward.sh sets --dangerously-skip-permissions itself and accepts no
  permission mode — and the toast now says so instead of implying a tool gate.
- lookupAgent() guards the policy lookup: target_agent comes from a queue YAML, and
  a plain policy['constructor'] is truthy, so a bare !entry check would accept a
  non-entry. Policy objects are built with a null prototype.
- A refused upgrade logs why. v0.4.0's refusals were silent on this side; the only
  signal was a 403 on the far side of the proxy, naming neither leg nor the cause.

Verified against a live instance of the new binary on a scratch port: no Origin -> 101,
allowed Origin -> 101, wrong Origin -> 403. Negative test confirmed a broken launcher
fails by name with nothing spawned. Tests 12 -> 41; 13 mutations applied across the
three modules, all caught.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

agent-id: developer
Audit of task-queue-plugin-repair-2026-08 found the two hand-written launch-policy
validators computed their containment root differently: Python called .resolve() on it,
this side uses a plain path join. Neither resolves the CANDIDATE project_dir, so
resolving only the root compares a canonical path against an uncanonical one.

The fix is on the Python side (the .resolve() is removed there). What changes here is
the comment: the plain join is now stated as a rule the other implementation must match,
rather than being an unexplained coincidence that the next edit could break.

Reproduced before fixing rather than taken on trust — with ~/.claude/projects made a
symlink, the same document was REJECTED by Python and ACCEPTED here. After the fix both
return the identical unresolved path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

agent-id: developer
@TadMSTR
TadMSTR merged commit 0f77c58 into main Aug 27, 2026
2 checks passed
@TadMSTR
TadMSTR deleted the fix/ws-repaint-launcher-2026-08 branch August 27, 2026 14:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant