You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Shell test suites pipe a payload into hooks that exit before reading stdin, so pipefail fails them intermittently #4458
Several shell test suites feed a hook's stdin through a pipe (printf ... | bash "$HOOK") while running under set -o pipefail, in cases where the hook exits before it reads stdin. When the hook exits first, the writer's write() hits a closed pipe. On the CI runner SIGPIPE is ignored, so printf prints write error: Broken pipe and returns 1 (with SIGPIPE at its default it dies with 141 instead). pipefail turns that into a nonzero pipeline status, so a case that asserts exit 0 fails even though the hook did exactly what it should. Whether the hook exits before the writer writes depends on scheduling, so the failure is intermittent and a rerun passes.
The hooks exit before reading stdin by design on these paths: the kill switch is hoisted above the library source (scripts/check-killswitch-hoist.sh, #3719), zone-gate's advisory mode returns before the read, statusline-shim exits when it finds nothing to run, and abort-boundary's forced-abort copies abort before the read. Hoisting the kill switch made that exit about 2 ms, which widens the window.
lib/hook-utils.test.sh#4357 is the same failure inside the library. hook::begin runs FILE=$(printf '%s' "$INPUT" | hook::read_file_path) || exit 0 (lib/hook-utils.sh:2461). The real hook::read_file_path reads to EOF. The BG_STUB_FILE stub in the test driver (lib/hook-utils.test.sh:3890) returns without reading, so the printf sometimes fails, the pipeline returns nonzero, and hook::begin exits 0 before the driver prints REACHED=1. That is the empty FILE_DIR/FILE_BASE signature. The Broken pipe line never shows in those logs because the driver's stderr is captured into the row output, and the fail message prints only the two fields. Run 35648159164 shows it directly: FAIL: begin escaped separator (rc=0): .../lib/hook-utils.sh: line 2449: printf: write error: Broken pipe. Production code is not affected, because the real reader drains.
A related form: a consumer that stops reading early, like printf | grep -q, can leave a still-writing producer with the same EPIPE. scripts/affected-tests.test.sh:1295 failed this way twice. The file's own comment at line 646 already warns about this pattern.
Reproduction (deterministic)
The producer is delayed 0.2 s so the hook always exits first. Run on Linux, bash 5.3, against the real hook:
That is the exact CI signature (line 116: printf: write error: Broken pipe / FAIL: kill switch failed (rc=1 out=)). For the hook-utils driver, a payload larger than the pipe buffer makes the stub race lose every time: REACHED is empty and rc is 0. With a draining stub, the same command prints REACHED=1 DIR=/a.
CI evidence
Taken from failed attempts of ci.yml in the last 1,000 runs (back to 2026-09-14) and the last 200 runs on main:
abort-boundary forced abort, on main: 35926275351 (block-exported-msys-pathconv: forced abort exits with the declared fail-open posture: expected exit 0, got 1).
hook-utils begin rows (lib/hook-utils.test.sh: 'a Windows backslash path' begin case is flaky on Linux CI #4357): 34627709404 (main), 34878040724, 35048389386, 35422052115, 35425847753, 35493773394, 35505472117, 35564735818, 35648159164, 35671551412, 35729512141 (attempt 2), 35784876592 (attempts 1 and 2), 35821568717, 35858998103, 35865887568, 35926874489, 35927528665, 35950573098, 36002361516. Five of the table's six rows have failed, "a file under the filesystem root" among them, and so has the escaped-separator case that uses the same stub.
I found the sites two ways. First, by reading every suite with a kill-switch case (47 files) and its run helper. Second, by running all 431 *.test.sh on Linux with a bash shim first on PATH. The shim runs the subject, then checks whether bytes were left unread on a piped stdin. It flagged every subject that exited without draining. I then classified each flagged site by how stdin was fed and whether the test observes the pipeline status or the producer's stderr. Suites whose tool was missing locally (actionlint, bash-format, go-format, ruff-format) were classified from source. Their helper is identical to the flagged copies.
Race: a producer process pipe, a subject that can exit first, and a status the test asserts. Line numbers are at 6cd0ae3.
Here-string or file feeds. This covers every guardrails suite through guard_invoke, claude-ops drive_with_sink/run_hook, hardcoded-path, secret-pattern, skill-reference, stale-path, cli-flag-verify, pr-linkage-mcp-gate, pr-ready-evidence(-mcp)-gate, and require-jq-posture. Bash writes a small here-string before the subject starts, so there is no writer left to fail.
The expected status is nonzero and the subject's own code wins under pipefail: worktree-create-gate.test.sh:186 (expects 1), goal-condition-length.test.sh:28 usage errors (expects 2), zone-crossing-inject.test.sh:1063 reserved destinations (expects 2).
The status is never read and the producer's stderr is not captured: index-drift.test.sh:165 and zone-gate.test.sh:71 (|| true).
The feed is a process substitution, so the producer is in no pipeline: lane-stop-gate.test.sh:143 (run_bare: with no settings file, gate_maybe_configured exits before the read). The directly piped cases at :223 and :471 set LANE_STOP_GATE_ENABLED, which carries the hook past that gate to hook::buffer_stdin_to, so they read first.
The pipe is held open on purpose: session-retention.test.sh:81 checks that the hook does not wait for EOF.
The subject inherits a while read loop's here-doc, with no producer process: check-shell-portability, check-silent-revert, check-contract-slice-prune, the code-metrics --version probes, and the stub CLIs in fetch-annotations, reap-project-plugin-records and fetch-all-pr-comments.
destructive-guard.test.sh:248 pipes into a script that reads all of stdin (INPUT="$(cat)") before its kill switch.
Fix shape
Feed the payload so that no writer process can outlive the reader. For the small JSON payloads here, use a here-string (the same shape guard_invoke already uses). Where a helper documents why it avoids here-strings (abort-boundary's pipe_run, which cites the 64 KiB here-string deadlock in lib/hook-utils.sh), feed from a file instead. Make the hook-utils stub drain stdin the way the real reader does. For the grep -q form, match on the captured string. Per docs/conventions/shell-test-helpers/README.md these helpers are duplicated per plugin on purpose, so the fix is one edit per copy, not a new shared helper.
Out of scope: #3504. It covers the production-side question of whether hooks should drain on early exit, and it stays a human decision. The session-event-log case (#4110) was fixed differently by #4322.
Problem
Several shell test suites feed a hook's stdin through a pipe (
printf ... | bash "$HOOK") while running underset -o pipefail, in cases where the hook exits before it reads stdin. When the hook exits first, the writer'swrite()hits a closed pipe. On the CI runner SIGPIPE is ignored, soprintfprintswrite error: Broken pipeand returns 1 (with SIGPIPE at its default it dies with 141 instead).pipefailturns that into a nonzero pipeline status, so a case that asserts exit 0 fails even though the hook did exactly what it should. Whether the hook exits before the writer writes depends on scheduling, so the failure is intermittent and a rerun passes.The hooks exit before reading stdin by design on these paths: the kill switch is hoisted above the library source (
scripts/check-killswitch-hoist.sh, #3719), zone-gate's advisory mode returns before the read, statusline-shim exits when it finds nothing to run, and abort-boundary's forced-abort copies abort before the read. Hoisting the kill switch made that exit about 2 ms, which widens the window.lib/hook-utils.test.sh#4357 is the same failure inside the library.hook::beginrunsFILE=$(printf '%s' "$INPUT" | hook::read_file_path) || exit 0(lib/hook-utils.sh:2461). The realhook::read_file_pathreads to EOF. TheBG_STUB_FILEstub in the test driver (lib/hook-utils.test.sh:3890) returns without reading, so theprintfsometimes fails, the pipeline returns nonzero, andhook::beginexits 0 before the driver printsREACHED=1. That is the emptyFILE_DIR/FILE_BASEsignature. TheBroken pipeline never shows in those logs because the driver's stderr is captured into the row output, and the fail message prints only the two fields. Run 35648159164 shows it directly:FAIL: begin escaped separator (rc=0): .../lib/hook-utils.sh: line 2449: printf: write error: Broken pipe. Production code is not affected, because the real reader drains.A related form: a consumer that stops reading early, like
printf | grep -q, can leave a still-writing producer with the same EPIPE.scripts/affected-tests.test.sh:1295failed this way twice. The file's own comment at line 646 already warns about this pattern.Reproduction (deterministic)
The producer is delayed 0.2 s so the hook always exits first. Run on Linux, bash 5.3, against the real hook:
That is the exact CI signature (
line 116: printf: write error: Broken pipe/FAIL: kill switch failed (rc=1 out=)). For the hook-utils driver, a payload larger than the pipe buffer makes the stub race lose every time:REACHEDis empty and rc is 0. With a draining stub, the same command printsREACHED=1 DIR=/a.CI evidence
Taken from failed attempts of
ci.ymlin the last 1,000 runs (back to 2026-09-14) and the last 200 runs onmain:line 58: printf: write error: Broken pipe,FAIL: [17] disabled hook exits 0 — exit expected 0 got 1).main: 35926275351 (block-exported-msys-pathconv: forced abort exits with the declared fail-open posture: expected exit 0, got 1).beginrows (lib/hook-utils.test.sh: 'a Windows backslash path' begin case is flaky on Linux CI #4357): 34627709404 (main), 34878040724, 35048389386, 35422052115, 35425847753, 35493773394, 35505472117, 35564735818, 35648159164, 35671551412, 35729512141 (attempt 2), 35784876592 (attempts 1 and 2), 35821568717, 35858998103, 35865887568, 35926874489, 35927528665, 35950573098, 36002361516. Five of the table's six rows have failed, "a file under the filesystem root" among them, and so has the escaped-separator case that uses the same stub.printf | grep -qF: 36010325374, 36015259142.Sweep
I found the sites two ways. First, by reading every suite with a kill-switch case (47 files) and its run helper. Second, by running all 431
*.test.shon Linux with abashshim first on PATH. The shim runs the subject, then checks whether bytes were left unread on a piped stdin. It flagged every subject that exited without draining. I then classified each flagged site by how stdin was fed and whether the test observes the pipeline status or the producer's stderr. Suites whose tool was missing locally (actionlint, bash-format, go-format, ruff-format) were classified from source. Their helper is identical to the flagged copies.Race: a producer process pipe, a subject that can exit first, and a status the test asserts. Line numbers are at 6cd0ae3.
run_hook_envin actionlint-check.test.sh:104, bash-format.test.sh:96, biome-format.test.sh:156, eol-normalizer.test.sh:210, go-format.test.sh:100, markdown-format.test.sh:179, powershell-format.test.sh:95, ruff-format.test.sh:116, typos-format.test.sh:108plugins/context-guard/hooks/post-compact-mark.test.sh:41(run)plugins/context-guard/hooks/zone-gate.test.sh:67,:116plugins/context-guard/hooks/zone-crossing-inject.test.sh:360plugins/desktop-notification/hooks/desktop-notification.test.sh:119(jq producer)plugins/rate-limit-guard/hooks/record-rate-limit-stop.test.sh:55(run)plugins/source-control/hooks/worktree-add-claim-gate.test.sh:56(run)plugins/source-control/hooks/worktree-add-containment-gate.test.sh:54(run)plugins/source-control/hooks/pr-body-linkage-gate.test.sh:638plugins/guardrails/hooks/abort-boundary.test.sh:52(pipe_run, used at :265 and :393)statusline-shim.test.shrun_envin context-guard (:115) and rate-limit-guard (:141)lib/hook-utils.test.sh:3890(stub underlib/hook-utils.sh:2461)scripts/affected-tests.test.sh:1295(printf | grep -qF)Safe (flagged or matched, but it cannot fail):
guard_invoke, claude-opsdrive_with_sink/run_hook, hardcoded-path, secret-pattern, skill-reference, stale-path, cli-flag-verify, pr-linkage-mcp-gate, pr-ready-evidence(-mcp)-gate, and require-jq-posture. Bash writes a small here-string before the subject starts, so there is no writer left to fail.|| true).run_bare: with no settings file,gate_maybe_configuredexits before the read). The directly piped cases at :223 and :471 setLANE_STOP_GATE_ENABLED, which carries the hook past that gate tohook::buffer_stdin_to, so they read first.while readloop's here-doc, with no producer process: check-shell-portability, check-silent-revert, check-contract-slice-prune, the code-metrics--versionprobes, and the stub CLIs in fetch-annotations, reap-project-plugin-records and fetch-all-pr-comments.INPUT="$(cat)") before its kill switch.Fix shape
Feed the payload so that no writer process can outlive the reader. For the small JSON payloads here, use a here-string (the same shape
guard_invokealready uses). Where a helper documents why it avoids here-strings (abort-boundary'spipe_run, which cites the 64 KiB here-string deadlock inlib/hook-utils.sh), feed from a file instead. Make the hook-utils stub drain stdin the way the real reader does. For thegrep -qform, match on the captured string. Perdocs/conventions/shell-test-helpers/README.mdthese helpers are duplicated per plugin on purpose, so the fix is one edit per copy, not a new shared helper.Out of scope: #3504. It covers the production-side question of whether hooks should drain on early exit, and it stays a human decision. The session-event-log case (#4110) was fixed differently by #4322.