Skip to content

perf(guardrails): run the guard chain in-process and answer payload fields without jq - #4185

Merged
kyle-sexton merged 7 commits into
mainfrom
guardrails-run-guards-in-process-no-subs
Sep 16, 2026
Merged

kyle-sexton merged 7 commits into
mainfrom
guardrails-run-guards-in-process-no-subs

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

No related issue: handoff-inbox item 20260913-034031 under program item 20260915-153000 (spawn budget per tool call); no GitHub issue was filed for either.

Summary

Every Bash and PowerShell tool call waited on the guardrails PreToolUse chain, and on Windows Git Bash that chain cost 23 process creations and about 880 ms before the tool's own command started. Almost all of it was the dispatcher's own shape: one command-substitution subshell per guard, a $(declare -f), a < <(printf | jq) process substitution, and a fork per telemetry subject. This change runs the chain inside one shell and answers payload fields without spawning jq, without changing any guard's decision.

Fix

  • plugins/guardrails/hooks/run-guards.sh sources each guard into its own shell. A guard's exit is a dispatcher function that records the status and runs the next guard from inside the call; stdout documents go through hook::emit_document; a guard that dies of a hard error hands its status to the abort boundary's new chain slot (_GAB_CONTINUE in abort-boundary.sh), which settles that guard's posture and runs the guards still owed in one subshell. Per-invocation analysis state (the alias memo) is reset before each guard.
  • lib/hook-utils.sh gains hook::_fast_fields, a builtin JSON field parser that answers a well-formed payload's plain-string fields and falls back to jq only for a shape it cannot prove (NUL escape, duplicate key, non-string value); hook::jq_fields_uncached, hook::emit_document, hook::extract_bash_subject_to, hook::reset_analysis_state. Synced to the 17 carrying plugins' hooks/hook-utils.sh with a patch bump each; guardrails is 0.34.0.
  • block-*.sh and flag-commit-pr-skill-bypass.sh resolve their telemetry subject in-shell via the _to helpers; block-windows-drive-tmp.sh masks quoted redirects in-shell (mask_quoted_redirect_ops_to).

Design notes: bash 5.3 does not propagate break across function calls, so exit-chaining replaced that; a second fatal error inside an EXIT trap ends bash with no result, so the still-owed guards run in one subshell rather than in the trap. Stdout collection is by contract, not fd capture: a guard that printfs a document directly bypasses the merge (documented in the run-guards.sh header).

Verification

Measured on Windows 11 + Git Bash 5.3 with an exact job-object census (TotalProcesses) of the harness's own bash -c "<hooks.json command>" invocation, n=5 isolated p50:

Lane Creations before Creations after p50 before p50 after
Bash (true) 23 3 880 ms 285 to 297 ms across two runs
PowerShell (exit 0) 100 80 3.3 s 2.4 to 2.7 s across two runs

The 3 remaining Bash-lane creations are the harness's bash -c, the env of the shebang, and bash itself; the chain spawns nothing. The 80 PowerShell-lane creations are all inside lib/powershell/ps-command.sh and are the next item (20260914-183000).

Decisions byte-identical (rc, stdout, stderr) against 0.33.11 via a two-root differential:

  • perf baseline guard_path_differential.json 17-command corpus x {Bash, PowerShell}: 34 cases, 0 mismatches
  • 653 commands harvested from the guard suites, Bash lane: 0 mismatches (9 before the alias-memo reset, which this differential exposed and the suite did not)
  • 200 of those commands, PowerShell lane: 0 mismatches
  • Write lane 17, Edit lane 13, drive-root-tmp lane 16 cases: 0 mismatches

Suites: lib/hook-utils.test.sh 504 pass (tests 22 to 24 added: fast fields vs jq, emit_document, subject _to); plugins/guardrails/hooks/run-guards.test.sh adds in-process chain, hard-error, alias-memo regression and no-fork xtrace cases; its 11 failures on this host (stale-path lane and if-row cases) are identical on main and pre-exist this change. scripts/sync-hook-utils.sh --check-bump origin/main and check-changelog-parity.sh --check-bump origin/main pass. A grep of the chained guards for process-global state escapes (builtin exit, exec, trap, shopt, set -e) finds only each guard's set -uo pipefail, which the dispatcher already runs under.

The wide affected-suite set from scripts/affected-tests.sh --base origin/main (every plugin carrying hook-utils.sh) is left to CI.

Security review (fresh-context reviewer over the diff, both trees runnable): no confirmed bypass. Three residuals it could not close were closed afterwards: _HOOK_UTR_TARGET_PHYSICAL is set and reset around each under_temp_root call (hook-utils.sh 1393 to 1400), so it never leaks between guards; the jq merge in run_guards::finish runs under the same open 0 2 boundary the pre-change dispatcher armed at its line 85; and a 33-payload fuzz of hook::_fast_fields against the library's jq arm (escapes, \u and surrogate forms, NUL, duplicate keys at both depths, nested same-name keys, key-in-value, non-string values, CRLF, a 200 KB value, invalid escapes, trailing junk) shows no divergence under either LC_ALL=C or en_US.UTF-8; every shape the parser cannot prove falls back to jq.

Residual: hooks.json keeps the shebang exec form. bash "<script>" would drop the env hop (3 to 2 creations), but the docs say shell-form commands run under sh -c on macOS/Linux, where /bin/sh is bash 3.2, so the form change is cross-platform unsafe and not applied.

Related

  • Program item 20260915-153000 (spawn budget per tool call). Later items build on this dispatcher and lib: 20260914-183000 + 20260915-144500 (fork-free ps-command.sh, quoted-git-token narrowing), 20260913-034035 (session-event-log), 20260913-034032/034033/034034/034036/034037 (if-gate and fast-exit hooks).
  • perf: drop leftover forks in hook-utils, always-on hooks, and CI scanners #3878 (precedent for the lib-side fan-out to 17 plugins).

🤖 Generated with Claude Code

…ields without jq

run-guards.sh sources its guards into its own shell instead of one
command-substitution subshell per guard: `exit` is a dispatcher function
that records the guard's status and runs the next guard from inside the
call, stdout documents are collected through hook::emit_document, and a
guard that dies of a hard error hands its status to the abort boundary's
new chain slot (_GAB_CONTINUE), which runs the guards still owed in one
subshell. Per-invocation analysis state (the alias memo) is reset before
each guard, so a later guard's alias walk is not answered by an earlier
guard's memo.

lib/hook-utils.sh answers a well-formed payload's plain-string fields
with its builtin JSON parser (hook::_fast_fields) and runs jq only for a
shape it cannot prove; hook::jq_fields_uncached names the same body for
the dispatcher's cache, hook::emit_document is the one stdout path, and
hook::extract_bash_subject_to is the in-shell telemetry subject. The 17
carrying plugins take the synced copy with a patch bump.

Measured on Windows 11 + Git Bash with an exact job-object census of the
harness's own invocation: the Bash-lane chain went from 23 process
creations to 3 and 880 ms to 285 ms isolated p50; the PowerShell lane
from 100 to 80 creations and 3.3 s to 2.7 s. Decisions are byte-identical
(rc, stdout, stderr) against 0.33.11 over the perf baseline's 17-command
corpus in both tool modes, over 653 commands harvested from the guard
suites on the Bash lane and 200 on the PowerShell lane, and over the
Write, Edit and drive-root-tmp lanes.

Handoff item: 20260913-034031 (program 20260915-153000).

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
@kyle-sexton
kyle-sexton marked this pull request as ready for review September 15, 2026 18:34
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-15T18:41:04.237824Z 952befe Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@claude

claude Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 7m 12s —— View job


Code review complete

This is a large diff (65 files) but almost all of it is mechanical: hook-utils.sh synced verbatim into 17 plugins with a patch-version bump and matching CHANGELOG.md entry, plus one-line conversions of SUBJECT=$(hook::extract_bash_subject ...) call sites to the new _to form. I scoped the review to the substantive logic: the in-process guard chain in plugins/guardrails/hooks/run-guards.sh and abort-boundary.sh, and the new builtin JSON field parser hook::_fast_fields in lib/hook-utils.sh. Security-shaped findings are scoped to the claude-security-review lane per REVIEW.md, so I omitted them here even where the diff touches trust-relevant code (e.g. the fail-open/fail-closed posture logic).

Traced the trickiest parts by hand — the exit shadowing/chaining through nested functions, real subshells, and pipelines; the hard-error → _GAB_CONTINUE → subshell-resume protocol; and the fast-field parser's uniqueness/depth proofs against the skeleton — and didn't find a case where a guard's decision changes. That matches the PR's own extensive differential testing (byte-identical rc/stdout/stderr against 0.33.11 over 34 + 653 + 200 + 46 harvested/corpus cases) and the new unit tests (22–24) that pin the fast parser against jq on ~24 payload shapes including duplicates, NULs, and unicode escapes.

Two findings posted inline:

  1. plugins/guardrails/hooks/abort-boundary.sh:69 — the pre-existing "chain slot" guidance a few lines above (56–62, untouched by this PR) says a function plugged into _GAB_CONTINUE must be "builtins only, never exits, never touches the trap." The dispatcher's own _GAB_CONTINUE=run_guards::guard_died (run-guards.sh:504) violates that by design — it spawns jq, sources further guard scripts, and ends in builtin exit, never returning. I didn't find a functional bug from this (traced it through; no double-settle, no lost trap), but it leaves the file's own documented invariant contradicted by its first real consumer, which will mislead whoever adds the second one.
  2. lib/hook-utils.sh:1102 — hook::_fast_fields skips indexing any JSON string body over 60 characters when matching key names, but places no length limit on the filter names it accepts. A filter naming a key longer than 60 chars that is genuinely present in the payload gets reported "proven absent" (empty string, rc 0) instead of falling back to jq (rc 2), silently violating the function's own prove-or-fallback contract. No current caller's filters are long enough to trigger this (all are well under 60 chars), so it's latent, not live — but this function is a public entry point synced into 17 plugins and feeds guard block/allow decisions directly, so it's worth closing off for future callers.

Both are linked with a "Fix this →" Claude Code URL in the inline comments.
(branch)

@claude

claude Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 5m 42s —— View job


Security review (/review:security-review)

  • Gather context (PR diff, custom instructions)
  • Run /review:security-review skip-gate check
  • Read changed guardrails files (dispatcher, abort boundary, hook-utils, block-*.sh)
  • Hunt for logic/trust-boundary/injection issues introduced by the in-process refactor and the new JSON field parser
  • Adversarially verify candidate findings
  • Post final findings

No security issues found.

Scope: this PR's security-relevant surface is the guardrails dispatcher rewrite (guard chain now runs in-process instead of one subshell per guard) and the new builtin JSON field parser that replaces most jq calls. Both are trust-boundary-adjacent: a bug in either could make a guard silently see the wrong command/field and fail to block something it should, across all 17 plugins that carry hook-utils.sh. I read this code independently rather than relying on the PR's own stated security review.

What I checked closely:

  • exit override / chain-continuation in run-guards.sh#L310 and the hard-error path via _GAB_CONTINUE in abort-boundary.sh#L128/#L130. Traced the BASHPID guard, the real-subshell passthrough, the "chosen status" check in guard::_abort_settle, and the double-hard-error path via run_guards::run_rest_in_subshell (#L401) — a guard's deliberate block decision (exit 2) is preserved correctly through the chain, and a hard error in one guard doesn't skip the guards still owed. The run-guards.test.sh additions (nested-function exit, real-subshell exit, fall-through, no-arg exit, single/double hard error under open and closed posture, and the alias-memo cross-guard regression) exercise exactly these edges and match what I traced by hand.
  • hook::_fast_fields in lib/hook-utils.sh#L1078, the hand-rolled JSON field parser. This is the highest-risk piece: a misparse here means a guard reasons about the wrong .tool_input.command. Verified the uniqueness/proof logic (duplicate key, key-spelled-as-a-value-elsewhere, same key at two nesting depths, non-flat nested object, non-string values, NUL escapes, non-object root) all fall back to return 2 (real jq) rather than guessing, which is the safe direction. The added hook-utils.test.sh cases (tests 22–24) specifically pin these adversarial shapes against the jq path and matched my manual trace.
  • Per-invocation state reset (hook::reset_analysis_state, lib/hook-utils.sh#L3798) is called before every guard in the chain, closing the alias-memo cross-guard leak the PR's own differential found; grepped hook-utils.sh for other armed/memo-style globals and found none left unreset.
  • Mechanical conversions (hook::extract_bash_subject_to, mask_quoted_redirect_ops_to, the block-*.sh telemetry call-site updates) are output-mechanism changes only (printf → printf -v), with no algorithm changes; telemetry SUBJECT/subject values feed only hook::emit_telemetry, never a guard decision.
  • Confirmed lib/hook-utils.sh is byte-identical (same blob SHA 0b66f9e0...) across lib/ and all 17 carrying plugins, so there's no drift introducing a divergent copy.

I could not execute the shell test suites or the differential harness myself in this sandbox (non-git Bash invocations require approval I don't have here), so this is a manual/static read rather than a re-run of the PR's own verification — I'd weight the PR's own differential/fuzz results (byte-identical decisions over 850+ harvested commands, 33-payload parser fuzz) alongside this.

Out of scope per this lane's charter: GitHub Actions/workflow hardening (zizmor's lane) and general code quality/style (/review:code-review's lane) — neither is what I was asked to check here.

@github-actions

Copy link
Copy Markdown
Contributor

Last security-reviewed head: 952befe1106481931f0658253fc9fd898fc24fbd. On the next push, the relevance gate compares only the commits since this SHA; delete this comment to force a full re-review.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 952befe110

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread lib/hook-utils.sh
Comment thread plugins/guardrails/hooks/abort-boundary.sh
Comment thread lib/hook-utils.sh Outdated
Comment thread plugins/guardrails/hooks/abort-boundary.sh
@github-actions

Copy link
Copy Markdown
Contributor

Claude has reviewed this PR 1 time. The lane skips further automatic reviews after 5; deleting this comment resets the count.

hook::extract_bash_subject_to assigns SUBJECT through a nameref, which
shellcheck 0.11 cannot follow; CI's lint lane failed SC2154 on the two
guards that read it in emit_tel. Declaring the variable empty first is
the idiom block-hook-bypass already uses. block-no-verify 256/0 and
flag-commit-pr-skill-bypass 35/0 unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
CI's machine-specific-paths hygiene check refuses a Windows user path
in a fixture. The value is opaque to the parser under test; hook-utils
504/0 unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
kyle-sexton and others added 4 commits September 15, 2026 21:26
…ty lint

The shell-portability lint reads the backslash-w in the previous spelling as a GNU-only regex class. The value is opaque to the parser under test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ts key bound

hook::_fast_fields indexes the payload's key strings with an associative
array, which Bash added in 4.0, and it ran on every hook::jq_fields call.
On the 3.2 shell macOS ships, and which these hooks document support for,
`local -A` fails per call. hook::_fast_fields_supported is the predicate,
split out the way hook::read_supports_nchars is so a test can force the
below-floor branch on a modern host, and hook::jq_fields_uncached asks it
before entering the fast path. Below the floor jq answers, unchanged.

The index loop also skipped any string body longer than a fixed 60 bytes
before decoding it, while nothing capped the key names a caller may ask
for. Two wrong answers came out of that: a requested key longer than 60
characters was proven ABSENT while present, and a key of 11 or more
characters spelled with \u escapes (hook_event_name is 15, 90 escaped)
was missed the same way. The bound is now six times the longest requested
key name, the width of `\uXXXX` per identifier character, which is the
bound the header comment always described.

The suite gains four cases: the below-floor branch forced with the fast
path replaced by a tripwire, a present 70-character key, an absent one,
and hook_event_name spelled entirely in \u escapes. Each compares the
fast path against jq rather than against an expectation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The chain-slot paragraph asks a function plugged into _GAB_CONTINUE to be
builtins only, never exit, and never touch the trap, and the slot's own
comment says it must not return. run-guards.sh's consumer does none of
that: run_guards::guard_died forks a subshell for the guards still owed,
spawns jq to merge their documents, and ends at `builtin exit`.

Say so. run-guards.sh is the one documented exception, and it is one
because it is the dispatcher finishing the run the process owes rather
than a hook doing exit-time work. The discipline is unchanged for
everyone else, and the slot comment now matches it: a chained function
that returns hands control back and the handler settles the status.

Comment only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…lliding plugins

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@kyle-sexton
kyle-sexton merged commit 0b0f38f into main Sep 16, 2026
12 checks passed
@kyle-sexton
kyle-sexton deleted the guardrails-run-guards-in-process-no-subs branch September 16, 2026 12:29
kyle-sexton added a commit that referenced this pull request Sep 16, 2026
No related issue: handoff-inbox item 20260913-034034 under program item
20260915-153000 (spawn budget per tool call); no GitHub issue was filed.

## Summary

disk-hygiene's `Stop` row started a whole Python interpreter on every
interactive stop to read the transcript and discover that this session
never launched `destructive_guard.py`, which is true of most sessions
because the engine-gate rows are `if`-gated on the engine's file name.
`Stop` rows accept neither `matcher` nor `if`, so the gate has to live
in the bash launcher, before the interpreter is resolved.

## Fix

`hooks/run-python-hook.sh` takes three optional leading flags, consumed
there and never forwarded to Python:

- `--marker-root <dir>`: the root the marker tree lives under, spelled
identically on the writer and the reader rows
(`"${CLAUDE_PLUGIN_DATA}"`, the same literal the rows already pass as
`--authorized-data-root`). An explicit root rather than the environment
variable, because a writer and a reader that disagree about the root
skip silently, which is the missed detection this monitor exists to
prevent.
- `--launch-marker <subdir>` (the engine-gate rows): write
`<root>/<subdir>/<session>.launched` before exec'ing the target, so a
guard that launches and dies still leaves the marker that keeps the
monitor running.
- `--skip-unless-marker <subdir>` (the `Stop` row): exit 0 without
exec'ing anything when that file is absent.

Candidates mirror the monitor's own `_marker_path_candidates` (the data
root, then a `${TMPDIR:-/tmp}` fallback beside the `.warned` marker).
The session id is recovered by a bash regex anchored to the payload's
opening key, with an unanchored fallback so a reordered payload degrades
to running Python rather than to keying on nothing; a payload it cannot
key on records nothing and skips nothing. The payload is buffered with
the `read` builtin and replayed to Python on a here-string.

disk-hygiene is bumped 0.23.11 to 0.23.12 with the numbers and the
marker semantics in the CHANGELOG.

## Verification

Job-object census (n=5, Git Bash `-c "<row string>"`,
`CLAUDE_PLUGIN_DATA` at a temp dir, harness floor `bash -c ':'` = 1):

| Row | Creations before | Creations after | Wall p50 before | Wall p50
after |
|---|---|---|---|---|
| Stop, no guard launched this session | 5 | 3 | 271 ms | 120 ms |
| Stop, marker present | 5 | 5 | 271 ms | 276 ms |
| engine-gate, later launches | 5 | 5 | 268 ms | 277 ms |
| engine-gate, first launch in a data root (`mkdir`) | 5 | 7 (n=1) | |
348 ms |

The `mkdir` is paid once per data root, never per session or per launch.

Byte identity: the fixture target writes `sys.stdin.buffer.read()`;
stdin after equals the payload plus the one trailing newline `<<<`
appends, which both consumers (`json.load(sys.stdin)` in the guard,
`sys.stdin.read()` + `json.loads` in the monitor) ignore. End to end on
a transcript carrying a real `hook_non_blocking_error`: stdout 534 bytes
on both roots, `cmp` identical; stderr identical;
`guard-decisions/decisions.jsonl` identical with timestamp and session
normalized. Skipped case: rc 0, 0 bytes out, 0 files written.

Suites: `run-python-hook.test.sh` 50 pass / 0 fail (32 pre-existing + 18
new; the pre-existing no-python case now also stubs `py`, because this
host's Windows py launcher resolved a real 3.13 and aborted the suite
before this change); `test_guard_launch_monitor.py` 28/28 both sides.
shellcheck and `check-shell-portability.sh --paths` clean;
`check-changelog-parity.sh --check-bump origin/main`,
check-hook-exec-form, check-hook-userconfig-argv,
check-hook-wiring-liveness, validate-plugin-contracts.mjs and
validate-plugins.sh green. Pre-existing and reproduced on the unmodified
tree: `test_hook_telemetry` 2 sink-timeout failures, `test_hygiene` 10
git-fixture errors.

Residuals, for review:

- The marker is one empty file per session that launched a guard, never
removed, and nothing sweeps the data root (`lib/guard_decision_log.py`
rotates only its own log). Stated in the wrapper and the CHANGELOG; a
retention sweep is a separate decision.
- A skipped turn emits no telemetry envelope where it used to emit an
`ok` one.
- The hooks reference shows `session_id` as the opening key but
documents no ordering guarantee; anchored-first plus fallback covers
both, not proven against a live wire payload.

## Related

- Program item 20260915-153000; siblings #4185 (item 1), #4188 (item 2),
#4189 (item 3), and the context-guard, hook-failure-audit and autonomy
hooks as their own PRs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 16, 2026
…4189)

No related issue: handoff-inbox item 20260913-034035 under program item
20260915-153000 (spawn budget per tool call); no GitHub issue was filed.

## Summary

`hooks/session-event-log.sh` is registered on 30 events, including every
PreToolUse, PostToolUse, PostToolBatch, UserPromptSubmit and Stop. With
the `session_event_log_enabled` option off (the default) each row still
spawned a three-process bash chain (the harness's `bash -c`, the `env`
of the shebang, bash) only to exit on the script's first check: 9 to 12
creations per Bash tool call, 0.3 to 0.5 s each on the critical path
under load.

## Fix

The 30 generated rows now carry an inline POSIX shell gate instead of a
bare script path:

```
[ "$CLAUDE_PLUGIN_OPTION_SESSION_EVENT_LOG_ENABLED" = true ] || exit 0; exec "${CLAUDE_PLUGIN_ROOT}"/hooks/session-event-log.sh
```

The harness's own shell evaluates the option and exits; only an enabled
logger execs the script. The row template lives in
`scripts/gen-hook-event-registry.sh` (the rows are generated from the
committed hook-event registry), so the template and its test changed and
the rows were regenerated; `--check` is clean over 33 events and the
registry is byte-identical. The script's own line-41 gate stays for
direct invocation. claude-ops is bumped to 0.56.14; a second commit
corrects the script's and README's now-false "pays the kill-switch read"
clause.

Why not an `if` field: the current hooks reference says `if` holds one
permission rule and is evaluated only on tool events, so it cannot read
a plugin option and cannot gate Stop or UserPromptSubmit rows at all.

## Verification

Exact job-object census of the generated command string as Git Bash runs
it, PreToolUse payload, n=5:

| Row command | Switch | Creations | p50 |
|---|---|---|---|
| old | off | 3 | 107 ms |
| new | off | 1 | 41 ms |
| old | on | 3 | 182 ms |
| new | on | 3 | 117 ms |

Per Bash tool call attributable to this hook while off: 9 to 12 → 3 to 4
(one per firing row). With the switch on the new command wrote the same
`sessions/<id>.jsonl` rows as the old one (10 lines, `status: ok`).

Suites: gen-hook-event-registry and session-event-log report the same
named pre-existing failures as before the edit on this host (a jq CRLF
case and a symlink-root case); check-killswitch-hoist,
check-hook-exec-form, check-hooks-description,
check-hook-userconfig-argv, check-hook-wiring-liveness,
validate-plugin-contracts, shellcheck, and changelog-parity `--check`,
`--check-bump origin/main` and `--check-preserved` all pass. The wider
affected-suite set is left to CI.

Residual: a registered shell-form row always costs the harness's one
shell process, so the item's target of 0 is reachable only by not
registering the per-tool rows while the option is off. hooks.json cannot
vary by option. Dropping the `PreToolUse *` and `PostToolUse *` rows in
favor of PostToolBatch, as the item also suggests, is a product call
left open.

## Related

- Program item 20260915-153000; siblings #4185 (guardrails in-process
chain) and #4188 (PowerShell guard path).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 16, 2026
… its fail-close to executable git (#4188)

No related issue: handoff-inbox items 20260914-183000 and its correction
20260915-144500, under program item 20260915-153000 (spawn budget per
tool call); no GitHub issue was filed.

Stacked on #4185 (base branch
`guardrails-run-guards-in-process-no-subs`): this branch needs that PR's
in-process dispatcher and lib. GitHub retargets it to main when #4185
merges; #4185's commits are merged in here.

## Summary

A PowerShell tool call paid 80 process creations and about 2.6 s in the
guardrails `PreToolUse` chain against 3 and 0.27 s for a Bash call, all
of it inside `lib/powershell/ps-command.sh`: `$(ps::…)` captures (some
inside per-character loops) and `printf | sed` pipelines, which on
Windows Git Bash each cost one or more process creations. Separately,
the fail-closed "cannot be parsed with confidence and could reach git"
sink fired on read-only pipelines that only name git as data, such as
`Get-Process | Where-Object { $_.Name -eq 'git' }`, because the git
probe is quote-intact by design.

## Fix

Two changes, each with its own differential contract.

- `perf(guardrails): make ps-command.sh fork-free` (no behavior change).
Every `$(ps::…)` capture becomes an out-parameter `_to` helper assigning
with `printf -v`; every `printf | sed` pipeline and `< <(printf …)`
reader becomes a pure-bash substitution or split (`ps::_gsub_to`,
`ps::_split_lines_to`). Two host behaviors the captured forms carried
are reproduced deliberately, because the command arrives with Windows
line endings and goes on to a Bash tokenizer: `$( )` here drops a
trailing CRLF whole, and this host's sed reads in text mode so a CRLF
line ending loses its CR.
- The comparison-operand narrowing (behavior change, deliberate), in
three commits whose last supersedes the first two. A quoted git literal
is blanked before the probe only when BOTH hold. First, it is the
right-hand operand of a comparison operator (`-eq -ne -in -notin
-contains -notcontains -like -notlike -match -notmatch -lt -le -gt -ge`,
optional `c`/`i` prefix, reached through `(`, `@(` and `,`-separated
list elements; left-hand operands and `& 'git' -eq $x` are calls, not
comparisons). Second, the whole command is provably a read-only cmdlet
pipeline, judged as an allowlist (`ps::_is_readonly_cmdlet_pipeline`):
with every quoted string replaced by an opaque placeholder, the command
is refused outright on a surviving backtick, `<#`, `--%`, `::`, `[`, `&`
or a `.` before `(`, then walked token by token so that every token at a
command position (start of input, or after `|` `;` `{` `}` `(` `=`
newline) is an allowlisted interrogator: any `Get-*` verb, the read-only
aliases (`gps`, `ps`, `gcim`, `gwmi`, `gci`, `ls`, `dir`, `gi`, `gc`,
`cat`, `type`, `gsv`, `gcm`, `gmo`, `gv`, `gl`, `pwd`),
`Where-Object`/`?`, `Select-Object`, `ForEach-Object`/`%`,
`Sort-Object`, `Measure-Object`, `Group-Object`, `Format-*`, `Out-*`,
`Write-Output`/`Write-Host`/`echo`, `Select-String`/`sls`, the path
helpers, `Compare-Object`/`diff`, `ConvertTo-*`/`ConvertFrom-*`, and the
`if`/`else`/`elseif`/`in`/`return` keywords. Variable chains,
`-parameters`, placeholders, numbers and arguments are inert; any other
command word refuses, and a command past the scan's length ceiling
refuses rather than being scanned.

The first two narrowing commits gated the exemption on a blocklist of
known executors. A fresh-context security review reproduced 13
BLOCKED-to-ALLOWED flips through executors no list enumerates:
`[Diagnostics.Process]::Start`,
`$ExecutionContext.InvokeCommand.InvokeScript`,
`[scriptblock]::Create().Invoke()`, a run-time `Set-Alias`,
`iwmi`/`Invoke-CimMethod` `Win32_Process Create`, `nsv`/`New-Service`,
`schtasks`, `New-ScheduledTaskAction`, `wmic`, and every launcher on
PATH (`npx`, `dotnet`, `cscript`, `explorer`, `ssh`). A blocklist of
executors is structurally under-inclusive, so the third commit inverts
it: an unrecognized command word now costs an over-block, never a
bypass. All 13 shapes plus `Invoke-Item`/`ii`, `Start-Job`, `New-Object`
are counterexample tests, and seven predicate pins assert the gate
directly. The correction item's own proposed rule (exempt any quoted
literal not preceded by `&`, `.` or iex) was not implemented either: it
would have allowed `Start-Process 'git' reset --hard` and `cmd /c 'git
push --force'`.

A second fresh-context review of the allowlist commit found that an
expandable string is itself a command position: PowerShell evaluates `$(
… )` inside a double-quoted string at construction time, and the
blanked-text walk cannot see inside a string. `Write-Output ("x" -eq
"$(cmd /c git push --force)")` and 14 sibling shapes flipped BLOCKED to
ALLOWED. The fourth commit adds two disqualifiers: the exemption refuses
on any expandable `"…"` span (detected by the quote walk's own opener
flag, since the opaque pass does not distinguish `'git'` from `"git"`),
and the sink refuses when an expandable `@"…"@` here-string was blanked
at intake, because that blanking erases the git token before the probe
runs. Single-quoted (verbatim) operands stay exempt. All 15 reviewed
shapes plus the three pre-existing holes they exposed (`-eq "$(git push
--force)"`, a `schtasks` subexpression, and the here-string operand) are
blocked test cases.

Guardrails is bumped to 0.35.0 (a behavior change). No lib or dispatcher
file changes here.

## Verification

Census (exact job-object count of the harness's own `bash -c`
invocation, n=5, isolated p50), re-run on the final commit:

| Payload | Creations before | Creations after | p50 before | p50 after
|
|---|---|---|---|---|
| PowerShell `exit 0` | 80 | 3 | 2422 ms | 300 ms |
| PowerShell `git status \| Where-Object { $_ -match 'x' }` (reaches the
sink, rc 2) | 44 | 3 | 1433 ms | 299 ms |
| PowerShell `Get-Process \| Where-Object { $_.Name -eq 'git' }` (newly
allowed) | 44 | 3 | | 333 ms |
| Bash `true` | 3 | 3 | 288 ms | 274 ms |

3 is the harness floor (`bash -c`, `env`, bash). The PowerShell lane is
now at 1.1x Bash; the program asked for at most 2x and at most 6
creations.

Differentials (two-root, rc + stdout + stderr byte-for-byte):

- Fork-removal commit (baseline = #4185's tree): perf corpus 17 commands
x {Bash, PowerShell} 34 cases 0 mismatches; 653 harvested commands
PowerShell lane 0 mismatches; 653 Bash lane 0 mismatches. Two probes the
chain differentials cannot see: every converted helper over 697 inputs,
15334 rows 0 differences; each pure-bash substitution against the sed it
replaced, 5054 cases 0 differences. Those probes caught two bugs the
corpus could not (an ERE mistranslation and a CR left on
`PS_SAFE_COMMAND`), fixed before commit.
- Narrowing, final tree (baseline = the pre-narrowing tree, which has no
exemption at all): 34 cases 0 mismatches; 653 PowerShell lane 0 decision
changes, with 0 predicted beforehand (no harvested command has a git
literal in operand position, and none carries a double-quoted string
together with one); one run of that lane reported 6 mismatches on rows
whose baseline arm returned an abnormal rc 256 with empty output under
host contention, and re-running those 15 rows in isolation gave 0
mismatches; 653 Bash lane 0 mismatches. The designed corpus of 47 shapes
(the exempt forms, the original counterexamples and the 13 reviewed
executor shapes) flips exactly the 5 intended commands (2 to 0, sink
message gone), allows 6 on both roots and leaves 36 blocked
byte-for-byte; extended with the 18 reviewed expandable-string shapes,
65 commands, it flips 6: those same 5, plus the `@"…"@` here-string
operand ALLOWED to BLOCKED, with 6 allowed on both roots and 53 blocked
byte-for-byte.

Suites: block-dangerous-git 491 to 579 pass, 0 fail (the correction's 8
rows, the counterexamples, the executor shapes, the expandable-string
shapes, predicate pins); block-no-verify 256/0;
block-noncanonical-commit 227/0; block-hook-bypass 645 pass with the
same 2 pre-existing symlink failures; run-guards 216 pass with the same
11 pre-existing failures on this host. `check-changelog-parity.sh
--check-bump origin/main` passes; `shellcheck -x` clean.

Residuals: the CRLF fidelity is pinned to Windows Git Bash behavior; on
Linux the pre-change code would have chomped only the LF and sed would
have read bytes, so the new code makes the Windows behavior uniform. The
allowlist over-blocks by design: a read-only pipeline that uses a cmdlet
outside the list, exceeds the length ceiling, or carries any
double-quoted string next to a compared git literal keeps today's
fail-closed sink, and a read-only git command carrying both a balanced
`@"…"@` here-string and another sink trigger now blocks where it did not
before. Pre-existing and out of scope: a bare `Write-Output @"` / `$(git
push --force)` / `"@` is allowed on both roots because the here-string
body is blanked at intake and no sink trigger fires at all; closing it
means treating every expandable here-string as a sink trigger, filed as
a follow-up. `ForEach-Object -MemberName Kill` / `-ArgumentList` method
dispatch is pinned at its current (allowed) behavior as a recorded
residual.

## Related

- #4185 (item 1, base of this stack; its shellcheck fix df80eb8 is
merged in here so the lint lane matches).
- #4189 (item 3). Program item 20260915-153000; the item-4 hooks
(20260913-034032/034033/034034/034037) follow as their own PRs.


🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 16, 2026
)

No related issue: handoff-inbox item 20260913-034037 under program item
20260915-153000 (spawn budget per tool call); no GitHub issue was filed.

## Summary

autonomy's `Stop` hook runs on every interactive stop. Its payload-free
pre-filter, the path every stop outside a lane takes, asked `uname -s`
which platform's managed-settings path to test; on the Windows Git Bash
host this gate is tuned for, that one command substitution costs three
process creations (the `$( )` fork, then the fork and exec of uname). An
unanchored (`--plugin-dir`) install paid a jq read of the plugin
manifest on the same path, answering a question nothing above the
pre-filter asks.

## Fix

- The pre-filter tests the fixed primary managed-settings path of every
platform with `[[ -f ]]`, which is equivalent: a candidate belonging to
another platform does not exist. The scan only routes
(`gate_managed_candidates_load` fills its own array and leaves
`GATE_MANAGED_FILES` alone), so every managed value still comes from the
`uname`-selected, absoluteness-asserted list in
`gate_managed_settings_files_load` once a session is actually evaluated.
The one asymmetry, the cwd-relative Windows spelling on a POSIX host,
can only force an evaluation that a repository's own settings `env`
block can already force through the two `CLAUDE_PLUGIN_OPTION_*`
presence tests; documented in the lib header and the pre-filter comment.
- `gate_resolve_install` is split: `gate_resolve_anchor` (pure parameter
expansion) stays above the pre-filter because it sets the
`GATE_CONFIG_ROOT` the user-settings locator needs;
`gate_resolve_plugin_name` (the jq) moves below it.

autonomy is bumped 0.23.12 to 0.23.13 with the numbers and the
threat-model note in the CHANGELOG.

## Verification

Job-object census (n=5, identical every run; Stop row as `bash.exe -c
'${CLAUDE_PLUGIN_ROOT}/hooks/lane-stop-gate.sh'`, HOME and
`CLAUDE_PLUGIN_DATA` at temp dirs). The floor on this invocation shape
is 4, measured from a `#!/usr/bin/env bash` + `exit 0` script under the
identical call.

| Arm | Creations before | Creations after |
|---|---|---|
| Outside a lane, anchored install | 6 | 4 |
| Outside a lane, unanchored install | 8 | 4 |
| Inside a lane (settings.json enables the gate) | 19 | 19 |

The hook now equals the floor: it spawns nothing of its own outside a
lane. Spawn sites were attributed with xtrace and a `$BASHPID` prompt
before the change (`lane-stop-gate-lib.sh` line 210 `uname -s`, and line
95 jq for the unanchored case) and show a single PID after. A PATH-shim
case proves the discrimination out of band: the previous hook launches
`uname`, the patched one launches nothing. Inside a lane the
`{"decision":"block",…}` payload is byte-identical before and after.

Suites: `lane-stop-gate.test.sh` 99/1 to 103/1 (the one failure,
`LANE-STOP\r-OK authorized: LAST must preserve CR`, is pre-existing on
the unmodified tree and unrelated); `lane-notify.test.sh` 10/0 both
sides. Four new cases: the candidate scan never fills
`GATE_MANAGED_FILES` and emits only fixed-root paths; the PATH shim; an
enabled lane still blocks; the strace budget case is updated to 0
launches (it SKIPs on Windows and is verified on CI's Linux lane only).
shellcheck clean on all four hook scripts; `check-changelog-parity.sh
--check-bump origin/main` passes.

Residual, out of scope here: the lib spells the Windows managed path
`C:/Program Files/ClaudeCode/managed-settings.json`; the pre-filter now
shares that one constant with the authoritative branch so they cannot
drift. Whether the current Claude Code settings reference names that
directory or `ProgramData` is worth checking separately, since it
decides whether this gate has ever read Windows managed settings.

## Related

- Program item 20260915-153000; siblings #4185 (item 1), #4188 (item 2),
#4189 (item 3), and the context-guard, hook-failure-audit and
disk-hygiene hooks as their own PRs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 16, 2026
…anged (#4193)

No related issue: handoff-inbox item 20260913-034032 under program item
20260915-153000 (spawn budget per tool call); no GitHub issue was filed.

Stacked on #4185 (base branch
`guardrails-run-guards-in-process-no-subs`): the envelope parse rides on
that PR's `hook::jq_fields` builtin parser in the synced lib. GitHub
retargets it to main when #4185 merges.

## Summary

context-guard's `zone-crossing-inject.sh` fires on every `PostToolBatch`
and `UserPromptSubmit`. Each fire spawned a jq for the envelope and then
the zone resolver (its own bash plus a jq) even when nothing the
resolver reads had changed since the last fire, which is the common
case: the statusline snapshot is rewritten only when the statusline
renders.

## Fix

- An unchanged-input skip. A `$STATE_DIR/$SESSION.seen` mark records the
inputs behind the last completed resolve in two ways: its mtime, stamped
with a redirection and compared with `-nt`, and one flags line (`z=<0|1>
c=<0|1>`) recording whether `zones.json` and the compaction marker
existed when that resolve ran, read back with builtin `read`. The fire
exits before starting a process only when the per-session snapshot,
`zones.json` and the compaction marker are all no newer than the mark
AND both existence flags still match the current `-e` results; a mark
with no readable flags line never takes the skip. The mark moves only
after the resolve persisted its markers, so a resolver failure, an
`unknown` reading and a failed marker write are each retried next fire.
A missing snapshot is never skippable. Skipping can only choose silence:
no arrangement of timestamps or existence changes can manufacture an
injection the full path would not have made.
- The envelope parse is size-branched. Under 64 KiB it goes through
`hook::jq_fields`' builtin parser (zero spawns); above it keeps the
single here-string jq (2 creations), because the helper's oversize
fallback reads through a process substitution and measured 4. Plain
routing through the helper would have made the large-payload path 9 to
11 creations; the branch is what keeps every cell at or below before.
The `65536` literal mirrors the helper's private proof ceiling and is
documented at the site.
- `STATE_DIR` resolution moves ahead of the resolver, so a session with
no state root exits one process earlier.

context-guard is bumped 0.7.64 to 0.7.65 with the numbers in the
CHANGELOG, and the README gains a "Skipping the resolve when nothing
moved" subsection carrying the table.

## Verification

Job-object census (n=5, identical across reps; subject floor 3 = `bash
-c`, `env`, bash):

| Fire | Payload | Creations before | Creations after |
|---|---|---|---|
| first (resolves) | small | 11 | 9 |
| repeat, nothing moved | small | 9 | 3 |
| snapshot rewritten | small | 9 | 7 |
| first | 150 KB | 11 | 11 |
| repeat, nothing moved | 150 KB | 9 | 5 |
| snapshot rewritten | 150 KB | 9 | 9 |

No cell is worse than before. Small repeat fire wall p50 1448 ms to 237
ms (re-run 365 ms; the host is bimodal). On Windows a resolve costs 4
creations rather than 2 because `bash "$RESOLVER"` hits the
`bin\bash.exe` wrapper, which re-spawns `usr\bin\bash`.

Crossing messages are byte-identical, asserted in the suite against a
control session driven through the same zone sequence with no skipped
fire. Suite: `zone-crossing-inject.test.sh` 78/1 to 97/2, where both
failures are `strace: no usable trace` on this host (Git Bash's cygwin
strace rejects `-e trace=`; the second is the new strace block, not a
regression); `zone-gate` 26/0, `post-compact-mark` 18/0. shellcheck
clean; `check-changelog-parity.sh --check-bump origin/main` and
`--check` green.

Residuals: the strace pins (steady fire 0 creations / 1 execve;
resolving fire 2 / 3) are set by reasoning and verified only on CI's
Linux lane. The one miss window is the resolve itself: a snapshot
written between the resolver's read and the stamp is marked seen and its
crossing is reported one fire late, never lost, since the statusline
rewrites the snapshot on its next render; on a filesystem or bash build
that compares mtimes at whole-second granularity the window is up to one
second. The README still carries the older "0.7.49 brought it to 3"
paragraph; the new subsection supersedes it.

## Related

- #4185 (base of this stack). Program item 20260915-153000; siblings
#4188, #4189, #4190, #4191, #4192.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 16, 2026
#4192)

No related issue: handoff-inbox item 20260913-034033 under program item
20260915-153000 (spawn budget per tool call); no GitHub issue was filed.

Stacked on #4185 (base branch
`guardrails-run-guards-in-process-no-subs`): this hook rides on that
PR's `hook::buffer_stdin_to` field read in the synced lib. GitHub
retargets it to main when #4185 merges.

## Summary

claude-ops' `hook-failure-audit.sh` `Stop` hook set the per-turn wall:
it re-scanned the whole session transcript on every stop (a `wc`, a
`grep`, a jq over every candidate) to discover that no new
`hook_non_blocking_error` record had appeared, which is the common case.

## Fix

- A per-session cursor,
`${CLAUDE_PLUGIN_DATA}/hook-failure-audit/<session>.cursor`, holds the
count of complete lines already audited and the `transcript_path` it was
taken against, beside the existing warning marker and under the same
7-day prune.
- A warm `Stop` reads with `mapfile -s` from one line before the cursor
(that line is an anchor: the same read that fetches new lines proves the
file still has that many; an empty result means shrinkage), prefilters
candidates in bash with the same fixed string the grep used, and runs jq
only for a candidate line. `mapfile` runs without `-t` so joined
candidates are byte-identical to `$(grep …)`; a final line with no
newline is scanned but not counted, so it is re-read next turn.
- The cursor resets to 0 (a full rescan) on: no data home, malformed
cursor, different `transcript_path`, fewer lines than the cursor, pruned
cursor, or bash without `mapfile` (the old grep path). It advances only
at the four disposal points (no candidate, empty summary, nothing new,
after the system message); a jq failure leaves it put. Rescanning cannot
re-warn, because the marker still decides that.
- The cold scan keeps its tail cap; one `wc -lc` now answers both the
cap decision and the cursor's starting line count.
- Payload fields ride on `hook::buffer_stdin_to INPUT '.transcript_path'
'.session_id'`, which fuses the library's validation probe into the
builtin field read; `hook::require_jq` moves after it.

claude-ops is bumped 0.56.14 to 0.56.15 with the numbers in the
CHANGELOG. Sibling #4189 also claims 0.56.14 on its branch; whichever
merges second needs its version and CHANGELOG re-based (parity is
checked against origin/main).

## Verification

Job-object census (n=5, Windows Git Bash; floor 3 = `bash -c`, `env`,
`bash`):

| Arm | Creations before | Creations after |
|---|---|---|
| second Stop, 20 benign lines appended (the common turn) | 10 | 3 (=
floor) |
| first Stop, data dir exists | 10 | 5 |
| first Stop, fresh data dir | 10 | 7 |
| second Stop, a failure record appended | 30 | 21 |

Wall clock after 0.23 to 0.36 s where before ran 0.36 to 1.38 s across
two runs; the host drifts about 4x within an hour, so the creation count
is the record.

Byte identity, old versus new hook on fresh data dirs, identical stdout:
a single record; three mixed classes; multiple registrations plus both
false-positive shapes; an over-cap transcript; a last line without a
trailing newline; and a two-turn incremental sequence where turn 2
appends a failure past the cursor. The suite asserts the same in-tree:
incremental turn-2 output equals a full rescan against identical marker
state.

Suites: `hook-failure-audit.test.sh` 84/4 to 94/4 (10 new assertions;
the 4 failures are pre-existing on this host, a Windows CRLF artefact in
`jq … @tsv` on `HAS_COMPLETED`). Discrimination checked: the jq
PATH-shim test fails against the old hook, and the path-reset test fails
against a copy with the path check removed. Degraded paths: a
no-`mapfile` copy warns then dedups, an unwritable marker home still
warns, an empty transcript is silent. shellcheck, shfmt, markdownlint
and `check-changelog-parity.sh --check-bump origin/main` clean.

Unproven: the strace budget block runs only on CI's Linux lane; the warm
ceiling (0 creations, 0 execs) follows from the Windows census at floor,
the cold ceilings (6 creations, 2 execs) are conservative estimates and
are what to adjust if CI reports otherwise. The bash 3.2 fallback was
exercised by forcing `HAVE_MAPFILE=0`, not on a real bash 3.2. An
over-cap cold scan sets the cursor from the whole file's line count
while reading only the last 2 MB; pre-window lines were never read
before either, and the cursor makes that permanent for the session.

## Related

- #4185 (base of this stack), #4189 (item 3, the other claude-ops bump).
Program item 20260915-153000; siblings #4188, #4190, #4191 and the
context-guard hook as its own PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Sep 25, 2026
…p per-edit spawns (#4418)

No related issue: WSL hook-latency program, rounds 1 to 5 (Write/Edit
PostToolUse window); no GitHub issue was filed for it.

## Summary

Every Write and Edit waits on the PostToolUse hooks, and on WSL the
window was 85 ms p50 for a small `.txt` edit and 399 ms for a 33 KB
Markdown edit. Most of it was not the hooks' work. Round 1: the builtin
JSON parse ran under the UTF-8 locale, `hook::begin` re-piped its
buffered payload through a byte-at-a-time `read`, guardrails parsed and
resolved the same payload once per verifier, and `repo_relative_path_to`
probed every PATH directory for a Windows-only `cygpath`. Round 2: every
hook paid a `#!/usr/bin/env bash` hop, a `realpath` process, and
stale-path-verify ran up to nine grep/sed/sort/head processes per
Markdown edit. Round 4: the JSON skeleton spent about 3.6 ms per parse
compiling one regex per token. Round 5: a 33 KB Markdown edit still
started two jq processes that could be answered in the shell. No hook's
output changes.

## Fix

Merge with #4450 (a74f8e8, 1c5ebf1; merge 0d17cdd). #4450 moved
every row to `bash "${CLAUDE_PLUGIN_ROOT}"/hooks/<script>` with
`"shell": "bash"` and added exit-before-source on every event. Both hold
here. The rows this PR changed keep `exec` on top of #4450's shape:
bash-format (2), markdown-format (2), guardrails post-verify (5), and
context-guard zone-crossing (2). Claude Code runs a row through `sh -c`
(dash) on Linux, and dash forks for a trailing `bash ...`; `exec bash
...` does not (strace: 1 fork vs 0). The eol-normalizer and typos-format
rows are now byte-identical to main. In zone-crossing-inject.sh, this
PR's no-context-dir exit and #4450's hoisted kill switch both run before
the first `source`; both are silent `exit 0`. Every overlapping carrier
moves one patch above main again (claude-ops 0.61.1, source-control
0.58.1).

Round 5 (0c3960e, 65fab8c; merges a80cd01, eeae865):

- `lib/hook-utils.sh` (17 carriers): `hook::_fast_fields` also proves
`.key // false | tostring` and `.key.sub // false | tostring` (absent or
null -> `false`, a boolean -> its name, a string -> itself; anything
else -> jq). Fuzzed against jq on 6000 payloads under C.UTF-8, C and
OSTYPE=msys (1131 answered, 0 mismatches); the old filter shapes are
unchanged on the round-4 corpora.
- `guardrails/hooks/skill-reference-verify.sh`: the plugins-root gate
runs before the payload read. With the lib change, the post-verify
dispatcher starts no jq on a `.md` edit outside a marketplace repo (it
had one per edit, for the `structuredPatch` filter). Inside a
marketplace repo nothing changes: one jq, shared with stale-path-verify
through the cache. run-guards test covers both.
- `typos-format/hooks/typos-format.sh`: in report-only mode (the
default) the finding classifier is built with builtins, byte-identical
to the jq program's output, when every output line is a typo finding in
typos' field order with printable-ASCII text and no quote or backslash,
and the output is at most 16 KB. Anything else, write mode, and Windows
bash run jq. 0.2 ms vs 2.9 ms for jq (primitive). Fuzzed on 7000
synthetic typos outputs (1857 answered, 0 mismatches); parity cases in
typos-format.test.sh.

Round 4 (5db7897; merges 98209f4, 8c049af, e681934):

- `lib/hook-utils.sh` (17 carriers): `hook::_json_skeleton` checks the
grammar with whole-string rewrites instead of one anchored regex (a
regcomp) per token: one regex for whitespace inside a token or between
two scalars, one for every scalar token, then a reduction loop over a
grammar string (per pass `k:v`, `s:v`, `v,v`, `m,m`, `{m}`, `{}`, `[v]`,
`[]`, root wrapped as `<k:ROOT>`). Escapes and control bytes are one
search over the whole neutralized payload. A structure past 8192
characters outside strings goes to jq. Parse of the S1 payload 3.6 ->
0.64 ms, S3 5.0 -> 2.1 ms (primitive).
- `markdown-format.sh` (`normalize_path_entry`, once per PATH entry on
the missing-markdownlint notice) and `powershell-format.sh`
(`to_pwsh_path`): `cygpath` only on msys/cygwin/win32; both trust gates
test for a drive-letter path before the lookup.
- Not done: the S3 timed stdin read. bash 5.3 masks SIGCHLD around every
character of a timed `read` (`read.def` 731-742; `TMOUT` too), any
untimed phase blocks forever on a never-closed pipe, and no standard
tool implements the sliced idle bound, so no replacement keeps the stall
semantics identical.
- Every carrier re-bumped one patch above main's after #4394 took the
same numbers; CHANGELOG bullets under the new version.

Round 3 (cd1b9d9, 21c1ef5; merges b0a5009, 0eda101):

- `bash-format/hooks/bash-format.sh`: `cygpath` is looked up only on
msys/cygwin/win32, as in `hook::repo_relative_path_to`. Elsewhere the
miss probed every PATH directory, including 9 `/mnt/c` entries on WSL;
bash-format was the slowest `.sh` PostToolUse hook because of it.
- `guardrails/hooks/stale-path-verify.sh`: a word anchor's occurrence
count uses `spv_word_occurrences` (awk `split` on non-word bytes under
C) instead of a per-character walk (33 KB file: 14 -> 2 ms); a word
anchor with `.` or `-` skips the count, which the walk never matched.
Non-word anchors keep the walk. Parity test against the walk added.
- `typos-format/hooks/typos-format.sh`: the classifier returns counts
and arrays through `tostring`, so `hook::jq_fields` reads them with the
builtin parser instead of a second jq.
- `lib/hook-utils.sh` (17 carriers): the skeleton's control-byte and
escape checks are one regex search each on Linux and macOS (Windows bash
keeps the glob checks, its regex decodes UTF-8 under C), and the key
walks take a key's text from its split part when no escape was rewritten
in it, instead of slicing the whole payload per string.
- CHANGELOG bullets added to each carrier's existing entry.

Round 2 (2540560, 9cc5cdf, 99d6205):

- hooks.json (superseded in part by #4450, which shipped the `bash`
prefix and `"shell": "bash"` on every row and `exec bash` on the
typos-format and eol-normalizer rows): what remains of this PR is `exec`
in front of `bash` on the markdown-format (2), bash-format (2),
guardrails post-verify (5) and context-guard zone-crossing
(PostToolBatch, UserPromptSubmit) rows. bash-format's launch-gate test
accepts both forms.
- `lib/hook-utils.sh` (synced to all 17 carriers): new
`hook::_physical_builtin_to` answers realpath's question with `builtin
cd -P` in one subshell, Linux only, for absolute existing directories
and existing non-symlink files; every other path, and Git Bash and
macOS, still runs realpath. Used by `hook::_physical_prime` and
`hook::physical_path_to`.
- `guardrails/hooks/stale-path-verify.sh`: builtin twins of the
code-span scan, the reconstruction token set, its anchor lines and its
recovered context, used only for printable-ASCII text; any other byte
keeps the pipelines.
- `context-guard/hooks/zone-crossing-inject.sh`: exits before sourcing
its libraries when `~/.claude/context-guard/context` does not exist and
jq is on PATH (no snapshot and no compaction marker for any session, so
the hook could only reach a silent `unknown`).
- Tests: stale-path twins vs pipelines on 14 fixed texts plus the gate;
`_physical_builtin_to` vs realpath per shape and its declined shapes;
the unresolved-path probe disables the builtin with `enable -n cd`;
run-guards expects 0 realpath on Linux.
- CHANGELOG bullets added to each carrier's existing entry for this PR
(versions were bumped in round 1).

Round 1 (7630993 .. c8daa81): C-locale builtin JSON parse
(`hook::_c_locale`), `buffer_stdin_to` without `jq -e .` for a proven
object, `read_file_path_to`/`raw_file_path_to` on the buffered payload,
`cygpath` lookup only on Windows bashes, `_uncached` twins; run-guards
caches path, root and fields per event; eol-normalizer skips `git
rev-parse` with no sink; context-guard starts its resolver only with a
readable snapshot; instruction-placement SC2154 init.

## Final verdicts

- S1 `.txt`, S2 `.md`, S2 `.sh`: MET. Every exec-count target: MET.
- S3 (33 KB `.md`): NOT MET. 70 ms p50 against a 60 ms target, per the
independent verifier at 0d17cdd.
- The guardrails post-verify exec count of 3 on a `.md` edit holds only
outside a marketplace repo; inside one, the dispatcher still starts one
jq.
- Part of the before -> after speedup comes from main commits merged in
after base 3fa014b, not from this PR alone.
- The implementer's `windows.py` dropped the first call per log; the
verifier's own cuts, which keep it, gave the same verdicts.
- Modes not exercised: markdown-format lint output, bash-format
shfmt/shellcheck output, and the Windows `cygpath` branches (covered by
`test-windows` CI only).
- Head f155042 adds main merges and a test-only fix
(`secret-pattern-detection.test.sh`, guardrails CHANGELOG) after the
measured 0d17cdd. Main changed two runtime guardrails hook files in
those merges: `secret-pattern-detection.sh` (#4475, #4486; PreToolUse,
outside the measured window) and `guard-requires.sh` (#4480,
comment-only, sourced by run-guards). `lib/hook-utils.sh` is unchanged
by main. Not re-measured.

## Verification

**Post-merge re-measure (#4450)** (head 0d17cdd, main 0fb4e95).
Differential, 11068 cases, run twice. Against base 3fa014b: identical
except the go-format telemetry counter (`Edit m.go` and `Write m.go`,
Go's random `telemetry/local/weekends` under HOME, same output) and the
same 13 lib-level deltas as round 5. Against origin/main itself: the
same 15 and nothing else. Every `row-*` case is identical, so `exec
bash` vs main's `bash` changes no behavior, and #4450's zone-crossing
hoist behaves as on main. CI dispatched at 0d17cdd: `ci.yml` (lint,
lint-2, hook-utils, test-linux 0-3, ci-status) and `test-windows.yml`
green.

One interleaved run at 0d17cdd (5 rounds x 4 strata, base / after /
safe; no local load; host qualified before and after, spread 1.67x /
1.74x, floor 0.3 / 0.4 ms). Write/Edit PostToolUse window, ms p50 / p95
(n):

| Stratum | base | after | safe | target | status |
|---|---|---|---|---|---|
| S1 `.txt` | 89 / 106 (35) | 42 / 50 (35) | 13 / 28 (35) | p50 <= 50,
p95 <= 90 | met |
| S2 `.md` | 94 / 115 (30) | 43 / 61 (30) | 13 / 59 (30) | <= S1 + 20 =
62 | met |
| S2 `.sh` | 86 / 136 (30) | 43 / 53 (30) | 13 / 321 (30) | <= 62 | met
|
| S3 33 KB `.md` | 398 / 450 (30) | 70 / 87 (30) | 12 / 52 (30) | <= S1
+ 15 = 57 | missed by 13 |

Counters (strace census, W/E post Edit, medians) are identical to round
5. Successful execs: typos-format S1 4 and S3 4, guardrails post-verify
`.md` 3 and S3 5, eol-normalizer 3, markdown-format 3, bash-format 4,
context-guard 2. Processes per hook are also unchanged (bash-format 6,
markdown-format 7, post-verify 6 / 4, context-guard 1), so `exec` still
saves the fork. Failed execve: 46 per hook. Failed probes: typos S3 308,
post-verify `.md` 260, S3 368.

**Round 5, final** (head eeae865). Differential extended to 11068
cases (replace_all as absent, null, false, true, "x", "", 1, [], {} in a
marketplace repo and in a git repo without `plugins/`, a null
tool_input, MultiEdit, typos content with ambiguous corrections and
non-ASCII words): identical except the go-format telemetry counter and
the same 13 lib-level deltas as round 4. Run at 65fab8c; the merge to
eeae865 changes no file under `lib/` or any tested plugin (`git diff
--stat` empty). CI dispatched at eeae865: `ci.yml` (lint, lint-2,
hook-utils, test-linux 0-3, ci-status) and `test-windows.yml` green.

One interleaved run at eeae865 (5 rounds x 4 strata, base / after /
safe; no local load; host qualified before and after, spread 1.67x /
1.64x, floor 0.3 ms). Write/Edit PostToolUse window, ms p50 / p95 (n):

| Stratum | base | after | safe | target | status |
|---|---|---|---|---|---|
| S1 `.txt` | 85 / 100 (35) | 41 / 47 (35) | 13 / 31 (35) | p50 <= 50,
p95 <= 90 | met |
| S2 `.md` | 93 / 144 (30) | 43 / 53 (30) | 12 / 21 (30) | <= S1 + 20 =
61 | met |
| S2 `.sh` | 86 / 103 (30) | 44 / 79 (30) | 14 / 40 (30) | <= 61 | met |
| S3 33 KB `.md` | 414 / 1828 (30) | 68 / 103 (30) | 13 / 26 (30) | <=
S1 + 15 = 56 | missed by 12 |

Counters (strace census, W/E post Edit, medians): successful execs
typos-format S1 4, S3 5 -> 4; guardrails post-verify `.md` 4 -> 3, S3 6
-> 5; eol-normalizer 3; markdown-format 3; bash-format 4; context-guard
2. No jq in the typos or post-verify rows for Edit any more. Failed
execve 46 per hook, unchanged.

S3 floor (primitives, idle host): the S3 - S1 difference that no code
change here can remove is the timed stdin read on the 35 KB payload
(+7.5 ms, bash's per-character signal mask), the typos scan of 33 KB
against 2 bytes (13.1 vs 11.5 ms, +1.6), and a minimal file_path
extraction (+0.4): about 9.5 ms. So the S3 floor is about S1 + 10 and
the S1 + 15 target sits above it; the measured gap is 27 ms, so about 17
ms of it is still removable in principle (payload validation and copies
in `buffer_stdin_to`, about +3 ms; the rest of `hook::begin`, about +2.7
ms; guardrails' stale-path scan of the 33 KB file and its remaining
parse, now the last hook to finish at 52 ms native against typos' 48).

**Round 4** (head e681934). Grammar proof: all 2,396,744 token strings
up to 7 tokens over `{ } [ ] , : "x" 1` give the previous walk's verdict
exactly; 32k random payloads under C, C.UTF-8 and OSTYPE=msys give
identical verdicts, skeletons, offsets, file paths and fields.
Differential extended to 9114 cases (13 grammar payloads with a
file_path, a structure past the cap, markdown-format's PATH remediation
with and without an identity cygpath): identical except the go-format
telemetry counter and 13 lib-level deltas (round 3's 12 plus the cap's
builtin -> jq fallback, same answer); run at e681934. CI dispatched at
e681934: `ci.yml` and `test-windows.yml` green.

One interleaved run at e681934 (5 rounds x 4 strata, base / after /
safe; no local load; host qualified, spread 1.5x / 1.81x). Write/Edit
PostToolUse window, ms p50 / p95 (n):

| Stratum | base | after | safe | target | status |
|---|---|---|---|---|---|
| S1 `.txt` | 86 / 99 (35) | 41 / 50 (35) | 9 / 16 (35) | p50 <= 50, p95
<= 90 | met |
| S2 `.md` | 93 / 107 (30) | 42 / 49 (30) | 12 / 20 (30) | <= S1 + 20 =
61 | met |
| S2 `.sh` | 84 / 95 (30) | 41 / 51 (30) | 11 / 33 (30) | <= 61 | met |
| S3 33 KB `.md` | 402 / 1264 (24) | 71 / 81 (30) | 12 / 25 (30) | <= S1
+ 15 = 56 | missed by 15 |

The host ran faster than round 3 (safe 9-12 vs 14-15), so part of S1's
53 -> 41 is drift. Counters: exec counts unchanged from round 3; traced
walls down about 3 ms of CPU per parse (typos S1 68 -> 59, guardrails S3
130 -> 115, strace-inflated).

**Round 3** (head 0eda101). Differential extended to 7899 cases (word
anchors beside non-ASCII, invalid UTF-8, CR and a 40 KB line; typos
content with 0, ~30 and ~3000 findings incl. write mode; an identity
`cygpath` on PATH): identical except the same go-format telemetry
counter and the same 12 lib-level deltas as rounds 1-2. Fuzz: occurrence
counter 3000 cases, skeleton/fast-field parser 4000 payloads (C,
C.UTF-8, OSTYPE=msys), 0 mismatches. CI dispatched at 0eda101:
`ci.yml` and `test-windows.yml` green.

One interleaved run at 0eda101 (4 rounds x 4 strata, base / after /
safe; host qualified, spread 1.43x / 1.81x). Write/Edit PostToolUse
window, ms p50 / p95 (n):

| Stratum | base | after | safe | target | status |
|---|---|---|---|---|---|
| S1 `.txt` | 98 / 130 (28) | 53 / 76 (28) | 14 / 69 (28) | p50 <= 50,
p95 <= 90 | p50 missed by 3 (host ran slower: base +13, safe +3 vs round
2) |
| S2 `.md` | 93 / 151 (24) | 47 / 56 (24) | 15 / 41 (24) | <= S1 + 20 =
73 | met |
| S2 `.sh` | 87 / 128 (24) | 49 / 77 (24) | 15 / 24 (24) | <= 73 | met
(61 at round 2) |
| S3 33 KB `.md` | 449 / 723 (24) | 88 / 152 (24) | 15 / 30 (24) | <= S1
+ 15 = 68 | missed by 20 |

Counters (strace census): typos-format S3 execs 6 -> 5; bash-format
`/mnt` probes 9 -> 0. The remaining S3 cost includes a timed-`read`
floor (2 `rt_sigprocmask` per byte, ~7 ms per hook at 37 KB) that only a
stall-semantics change would remove; not done.

**Round 2 and earlier:**

**Behavior differential** (no timing; `diff.py`, base = origin/main
3fa014b vs this head, same fixture path, fresh HOME/data/TMPDIR per
run; stdout, stderr, exit code, fixture/HOME/data/TMPDIR/sink contents).
6102 cases: every changed hook and every `hook::begin` caller, the
guardrails dispatcher on its three lanes and the verifiers standalone,
context-guard, a lib-level probe, and now the hooks.json rows themselves
run as `sh -c <row>` from each arm (1905 row cases). Round 2 adds a
fixture history with deleted, moved and non-ASCII-named paths (so
stale-path-verify reports findings), 30 stale-path cases (spans,
unclosed and double backticks, CRLF, tabs, whitespace lines, fences,
root basenames, self-repeating anchors, replace_all with over 40 hits,
non-ASCII and ESC in new_string and in recovered lines, 40 KB ASCII, the
33 KB doc), context-guard with no context directory, and a HOME-unset
mode, on top of round 1's corpus and modes. Result: every hook-level
case identical except go-format Write/Edit `m.go`, whose only difference
is Go's own random telemetry counter under HOME. 12 lib-level deltas,
the same 12 as round 1 (base lacks round 1; none reachable as hook
output). Verdict: no behavior change.

Also: the twins were fuzzed on 3000 random texts under C and C.UTF-8
against this host's grep (ugrep 7.8.4); CI runs the fixed-case parity
test on GNU grep and Git Bash. `_physical_builtin_to` matched uutils and
GNU realpath on 29 path shapes wherever it answers.

**Suites** (local): hook-utils 513/0, typos-format 178/0, eol-normalizer
85/0, bash-format 55/0, run-guards 242/0, stale-path-verify 121/0,
skill-reference-verify 151/0, cli-flag-verify 96/0, zone-crossing-inject
103/0, claude-ops fleet-state pass; markdown-format 172/2 (the same 2
fail at c8daa81: this host has no jq in `/usr/bin`). shellcheck,
shfmt, killswitch-hoist, hook-exec-form, userconfig-argv,
shell-portability, hooks-description, silent-skips, wiring-liveness,
changelog-parity, `sync-hook-utils.sh --check` and `--check-bump
origin/main`, markdownlint: pass.

**CI** at 99d6205, dispatched: `ci.yml` (lint, lint-2, hook-utils,
test-linux 0-3, ci-status) and `test-windows.yml`: all green.

**Measurement** (WSL2, `claude -p` Haiku; one run at 99d6205, arms
base / after / `--safe-mode` interleaved, 4 rounds x 4 strata; host
qualified before and after, `is_measurable` True, spread 1.72x / 1.42x).
Write/Edit PostToolUse window, ms:

| Stratum | base p50 / p95 (n) | after p50 / p95 (n) | safe p50 / p95
(n) | target | status |
|---|---|---|---|---|---|
| S1 `.txt` | 85 / 98 (28) | 43 / 52 (28) | 11 / 17 (28) | p50 <= 50,
p95 <= 90 | met |
| S2 `.md` | 90 / 94 (24) | 44 / 46 (24) | 13 / 20 (24) | <= S1 + 20 |
met |
| S2 `.sh` | 83 / 87 (24) | 61 / 70 (24) | 10 / 16 (24) | <= S1 + 20 |
met |
| S3 33 KB `.md` | 399 / 444 (24) | 89 / 95 (24) | 11 / 28 (24) | <= S1
+ 15 | missed by 31 |

Counters per W/E post Edit call (strace census, 2 traced sessions per
arm per stratum), base -> after: successful execs typos-format 7 -> 4,
eol-normalizer 7 -> 3, markdown-format 6 -> 3, bash-format 7 -> 4,
guardrails post-verify `.md` 16 -> 4 (`.sh` 7 -> 3, 33 KB 23 -> 6),
context-guard PostToolBatch 4 -> 2. Failed file probes typos 580 -> 290,
eol 409 -> 212, post-verify `.md` 1060 -> 298. Failed execve stay 46 per
hook: dash's `exec bash` walks PATH with execve as `env` did; plain
`bash` would walk with stat but forks, and measured slower (1.5 vs 1.3
ms, hyperfine primitive).

The context-guard early exit applies only on a host without the
context-guard status line (this one); with the status line the directory
exists and the full path runs as before.

## Related

- Overlaps #4450 (row shape, exit-before-source); merged in and
reconciled as described under Fix.
- Pattern precedent: #4185 (dispatcher cache, `_uncached` twins,
17-carrier sync); row shape: #4421, #4422.
- Plugin README cost tables (guardrails, bash-format, typos-format,
eol-normalizer) are dated measurements of earlier versions and were not
re-measured here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant