shell: fix use-after-free when a pipe read fails inside the spawn call - #32754
Conversation
Cmd::buffered_output_close now applies the same null-interp gate that Cmd::on_exit already uses: if SubprocExec.interp has not been published yet (transition_to_exec does that only after spawn_async returns), mark the command Done and return Yield::Suspended instead of a runnable Yield::Next. transition_to_exec picks the Done state up right after the spawn. Without this, an eager pipe read failing inside ShellSubprocess::spawn_maybe_sync_impl (recv returning ENOMEM, EMFILE, etc.) made PipeReader::on_reader_error drive the trampoline into Cmd::deinit, freeing the Box<ShellSubprocess> that the spawn frame was still dereferencing.
|
Updated 8:22 AM PT - Jun 26th, 2026
❌ @robobun, your commit 0ec6037 has 1 failures in
🧪 To try this PR locally: bunx bun-pr 32754That installs a local version of the PR into your bun-32754 --bun |
|
Warning Review limit reached
More reviews will be available in 12 minutes and 37 seconds. Learn how PR review limits work. Your organization has used up its prepaid credits, and credit purchases are no longer available. Enable the review add-on in the billing tab to keep reviews running — you're only billed for reviews past your plan's rate limits ($0.25/file). ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based credits. 🚦 How do rate limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (3)
WalkthroughThe PR removes spawn-arena bookkeeping, adds a spawn-time guard for buffered subprocess completion, adjusts pipe-reader lifetime handling, and adds Linux regression coverage for synchronous pipe-read failures. ChangesShell subprocess fault handling
Suggested reviewers
🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
The second finding from the same campaign, the stale node-id panic, has the same root cause and is covered by the gate in this PR: plus the Independent reproduction on stock main, no fuzzing harness needed: an LD_PRELOAD shim that interposes libc The same fix plus that epoll variant of the regression test is on |
|
Hit this same crash from the inline repro in the report thread and ended up at the same fix (gating One thing worth folding in here: the comment block kept right below the new guard still says the pre-spawn completion arm is "unreachable in practice" because |
|
The same crash also reproduces through a second path that this diff does not cover: The gate stops the trampoline, but Branch Not opening a separate PR for this; folding those two pieces into this one keeps it to a single fix. |
A second in-spawn failure surfaces the same re-entrant teardown:
epoll_ctl failing inside PosixBufferedReader::register_poll (called
from PipeReader::start) fires on_reader_error synchronously. That
callback reaches Cmd::buffered_output_close -> close_io ->
Readable::finalize, which drops the Readable::Pipe slot's Arc, so the
callback's guard was the last strong ref and the PipeReader was freed
when the callback returned. PipeReader::start then kept dereferencing
self (subproc.rs:2015); PosixBufferedReader::start returns Ok after
register_poll, so the error never reaches the caller.
Readable::start_pipe_reader now holds an Arc clone across start() and
read_all(), the same keepalive PipeReader::guard_from_raw already
provides for the on_reader_{done,error} callbacks themselves.
Also remove Cmd::spawn_arena_freed: it was only ever written and then
discarded with `let _ =`, and the comment that justified keeping it
(initSubproc sets it before any pipe can close) states the premise
this PR disproves. The buffered_output_close gate on exec.interp now
encodes the pre/post-spawn split directly.
The regression test gains epoll_ctl variants alongside the recv ones
(one LD_PRELOAD shim selects the fault via SHELL_FAIL_RECV or
SHELL_FAIL_EPOLL) and the tests run concurrently.
|
#32759 fixes a third, independent use-after-free on this same error path: The freed objects differ (this PR: |
There was a problem hiding this comment.
Actionable comments posted: 3
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/runtime/shell/states/Cmd.rs (1)
573-584: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winCondense the spawn-time gate comments to the 3-line limit.
The invariant is useful, but both new comment blocks exceed the repository rule. Keep the local explanation short and leave the longer narrative in the PR/test context.
Suggested condensation
- // Stage the exec slot *before* spawning so PipeReader / process-exit - // callbacks (which deref `cmd_parent.exec`) see a populated `Subproc` - // with the correct `child` once `spawn_async` writes through - // `out_subproc`. `interp` is left null until `spawn_async` and the - // `did_exit_immediately` handling have returned: a synchronous - // `Cmd::on_exit` (process exit handler) or `Cmd::buffered_output_close` - // (an eager `read_all` on a pipe erroring inside the spawn) would - // otherwise drive the trampoline (`Yield::run(&*interp)`) while this - // frame still holds `&Interpreter`, tearing the Cmd down (and freeing - // `child`) underneath the live `subproc` borrow. With `interp` null, - // both record `exit_code`/`state = Done` and return; we resume via - // the Yield we hand back below. + // Stage `Exec::Subproc` before spawning so callbacks can record completion. + // Keep `interp` null during spawn-time callbacks; they must not run the + // trampoline until this frame unwinds and the `CmdState::Done` check resumes. ... - // Same gate as `on_exit`: `exec.interp` is null until - // `transition_to_exec` returns from `spawn_async`. A runnable - // Yield here would reach `Cmd::deinit` and free the - // `ShellSubprocess` the spawn frame still dereferences; - // `transition_to_exec` resumes via its `CmdState::Done` check. + // Same spawn-time gate as `on_exit`: before `exec.interp` is + // published, mark `Done` but let `transition_to_exec` resume.As per coding guidelines, "Keep code comments to 3 lines max."
Also applies to: 966-970
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/runtime/shell/states/Cmd.rs` around lines 573 - 584, Condense the spawn-time gate comments in Cmd::spawn_async-related logic to no more than 3 lines while preserving the key invariant: stage exec before spawn, keep interp null until spawn_async/did_exit_immediately finish, and avoid running the trampoline while the live subproc borrow exists. Keep the explanation brief in the Cmd state transition comments and remove the long narrative, including the duplicated explanation near the other commented block referenced by the review.Source: Coding guidelines
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/js/bun/shell/shell-pipe-read-fault.test.ts`:
- Around line 173-184: The crash-mode subprocess test is asserting parsed stdout
before checking the process result, which can hide the useful stderr/exit
diagnostics when the native crash happens. Update the shell-pipe-read-fault
tests around lastJsonLine and the shell survives... cases to keep stdout,
stderr, and exitCode asserted together in one failure path, and avoid parsing
stdout in a way that masks crash output; apply the same pattern to the related
concurrent tests referenced in the comment.
- Around line 1-22: The comment block in shell-pipe-read-fault.test.ts contains
detailed ASAN stack history and debug narrative that should not live in the
test. Shorten the header to a minimal regression note, or replace it with just
the issue reference if available, and keep only the durable invariant around the
fault scenarios covered by the test.
- Around line 158-162: The shell pipe fault test is inheriting the inactive shim
mode from bunEnv/process.env, so both fault shims can be enabled at once. Update
the env setup in shell-pipe-read-fault.test.ts to explicitly unset the
non-target shim variable when setting the active one in the mode-specific cases,
so each test only exercises the intended fault. Use the existing bunEnv,
LD_PRELOAD, and mode-specific env keys (SHELL_FAIL_RECV/SHELL_FAIL_EPOLL) to
keep the test hermetic.
---
Outside diff comments:
In `@src/runtime/shell/states/Cmd.rs`:
- Around line 573-584: Condense the spawn-time gate comments in
Cmd::spawn_async-related logic to no more than 3 lines while preserving the key
invariant: stage exec before spawn, keep interp null until
spawn_async/did_exit_immediately finish, and avoid running the trampoline while
the live subproc borrow exists. Keep the explanation brief in the Cmd state
transition comments and remove the long narrative, including the duplicated
explanation near the other commented block referenced by the review.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 42e86f2c-253b-483e-bb50-a707294cf31a
📒 Files selected for processing (3)
src/runtime/shell/states/Cmd.rssrc/runtime/shell/subproc.rstest/js/bun/shell/shell-pipe-read-fault.test.ts
Unset both shim mode variables before enabling the one under test so a
preset SHELL_FAIL_RECV / SHELL_FAIL_EPOLL in the outer environment
cannot turn on both faults at once, and fold each test's checks into
one combined { parsed, stderr, exitCode } assertion so a native crash
surfaces the child's stderr and exit code in the same failure diff.
Trim the test header and the buffered_output_close gate comment; the
ASan stacks and the longer narrative live in the PR description.
|
Update to my comment above: it was accurate against this PR's state at the time (only the Verified: with this branch's current Two small things from #32759's branch (
|
|
Builds 64945 and 64974 both go red on the same single job, the The Everything else flagged on build 64974 is a CI-tagged flaky auto-retry in code this PR does not touch: HMR races in The diff itself is verified on both ASAN-debug and release builds; the proof matrix is in the PR description. Getting this green needs someone with Buildkite access to retry that one darwin shard. |
There was a problem hiding this comment.
All three of my earlier concerns were addressed in 2767858/aa7fcade and the current pass found nothing new, but this is a memory-safety fix in core shell subprocess code (unsafe raw-pointer access, re-entrant Arc<PipeReader> teardown, a new Yield::suspended() gate on the completion path) that warrants a human look — also note the musl build failures on 2767858 and robobun's two optional hardening suggestions from #32759.
Extended reasoning...
Overview
The PR fixes two heap use-after-frees in the Bun Shell that occur when a pipe syscall (recv or epoll_ctl) fails synchronously from inside ShellSubprocess::spawn_async. It touches src/runtime/shell/states/Cmd.rs (adds an exec.interp.is_null() gate in buffered_output_close, removes the dead spawn_arena_freed field), src/runtime/shell/subproc.rs (adds an Arc::clone keepalive in Readable::start_pipe_reader so the PipeReader survives re-entrant close_io → Readable::finalize), and adds a Linux-only LD_PRELOAD fault-injection regression test.
I previously raised three inline concerns on this PR (the gate-only diff missing the Arc keepalive for the epoll_ctl path; the stale spawn_arena_freed field/comment; test.concurrent). All three were addressed in 2767858, and the CodeRabbit test-hygiene nits (hermetic env, condensed header comment, combined assertion) were addressed in aa7fcad. The bug hunting system found nothing new on the current revision.
Security risks
None in the conventional sense (no auth/crypto/permissions). But this is a memory-safety patch: it changes when Cmd returns a runnable Yield (which drives Cmd::deinit and frees the Box<ShellSubprocess>) and how long the Arc<PipeReader> is pinned across re-entrant vtable callbacks. The subproc.rs change rewrites the body of start_pipe_reader, deriving a &mut PipeReader from a raw pointer obtained via arc_as_mut_ptr on a cloned Arc rather than the slot's Arc — the SAFETY comment was updated, but the aliasing argument is subtle (the re-entrant callbacks also hold raw *mut PipeReader, and Readable::finalize may swap the slot to Closed while p is live). I believe the change is correct, but it is exactly the kind of unsafe-Rust lifetime reasoning that benefits from a maintainer's eyes.
Level of scrutiny
High. This is production-critical native code in the shell subprocess pipeline. The fix is small in line count but the correctness argument depends on understanding the interleaving of spawn_maybe_sync_impl → start_pipe_reader → PipeReader::start/read_all → vtable on_reader_error → Cmd::buffered_output_close → close_io → Readable::finalize, plus the existing transition_to_exec post-spawn CmdState::Done resume. Not a config tweak or mechanical refactor.
Other factors
- CI: robobun's status comment shows musl build failures (aarch64-musl, x64-musl, x64-musl-baseline) on 2767858. These may be infra-related, but should be confirmed green on aa7fcad before merge.
- Outstanding suggestions: robobun's last comment (12:37 UTC) lists two optional, non-blocking hardening items from the now-closed #32759 — an additional ENOSPC
epoll_ctlregression test, and an early-return inPipeReader::startoncestate == Err. Neither is required for memory safety with the keepalive in place, but a maintainer may want to weigh in on whether to fold them in here. - Test coverage: Strong. Four LD_PRELOAD fault-injection cases (recv × {stdout-only, both-pipes}, epoll × {stdout-only, both-pipes}), each independently load-bearing per the proof matrix in the description, plus 560 existing shell tests reported passing.
PosixBufferedReader::start reports a poll-registration failure through the on_reader_error vtable callback, not its return value, so PipeReader::start would continue into its Ok arm and set the Socket flags on a reader that on_reader_error had already torn down. That is harmless with the Arc keepalive in start_pipe_reader, but SubprocessPipeReader::start (the Bun.spawn twin of this function) pairs its keepalive with exactly this state check; the shell copy was missed when that was added. This restores the parity.
|
Checked everything raised above. musl: the "build failed on aarch64-musl / x64-musl / x64-musl-baseline" entries were on build 64966, which was canceled when aa7fcad superseded it; those are cancellation artifacts, not build errors. On build 64974 all three musl builds pass. CI on aa7fcad (build 64974): On the two suggestions carried over from #32759:
|
There was a problem hiding this comment.
All three of my earlier findings are addressed (the Arc<PipeReader> keepalive, the spawn_arena_freed cleanup, and test.concurrent), and the latest pass found nothing new — deferring only because a UAF fix touching unsafe pointer derivation and re-entrant interpreter teardown is worth a human sanity check.
Extended reasoning...
Overview
This PR fixes two heap use-after-frees in the Bun Shell where a pipe error (recv() ENOMEM or epoll_ctl() ENOMEM/ENOSPC) surfacing synchronously inside ShellSubprocess::spawn_async tears down state still in use by the spawn frame. Three files are touched:
src/runtime/shell/states/Cmd.rs: adds anexec.interp.is_null()gate inbuffered_output_close(mirroring the existingon_exitgate) so pipe-side completion suspends rather than driving the trampoline intoCmd::deinitwhile the spawn frame still holds theShellSubprocess. Also removes the now-deadspawn_arena_freedfield, its initializer, write site, and stale doc comment.src/runtime/shell/subproc.rs:Readable::start_pipe_readernow holds anArc::clonekeepalive acrossstart()/read_all()so a synchronouson_reader_error→close_io→Readable::finalizecannot drop the last strong ref whilePipeReader::startis still on the stack.PipeReader::startalso gained an early return whenstate == Err, matchingSubprocessPipeReader::start.test/js/bun/shell/shell-pipe-read-fault.test.ts: new Linux-only regression test with an LD_PRELOAD shim covering both fault modes × stdout-only / stdout+stderr (4 concurrent tests). The PR description includes a verified fail-before/pass-after matrix showing each half of the fix is independently load-bearing.
Prior review cycle
I left three inline comments on the gate-only revision (one 🔴 for the missing keepalive on the epoll path, two 🟡 nits). All three were addressed in 2767858 and are marked resolved. CodeRabbit's three follow-up comments (test header verbosity, env hermeticity, combined assertion) were addressed in aa7fcad. Commit 0ec6037 added the PipeReader::start early-return for parity with SubprocessPipeReader. The bug-hunting pass on the current head found nothing.
Security risks
None introduced. This hardens an error path against memory corruption; it does not touch auth, crypto, permissions, or any externally-reachable surface. The new test uses LD_PRELOAD only against a locally-compiled shim in a temp directory under the test harness's own bunEnv.
Level of scrutiny
High. The change is small and well-tested, but it edits memory-safety invariants in a critical native path: unsafe { &mut *arc_as_mut_ptr(...) } derivation, Arc strong-count reasoning under re-entrant callbacks, and the interpreter's Yield trampoline gating that decides when Cmd::deinit may run. The reasoning in the PR description and code comments is sound, and the existing 560-test shell suite plus the new fault-injection tests pass on all Linux lanes (the only CI red is darwin agent infrastructure unrelated to this change). Still, this is exactly the kind of lifetime/ownership change where a second pair of human eyes on the invariant is cheap insurance.
Other factors
CI on the latest commit passes everywhere the new test actually runs (including x64-asan and musl); the two red jobs are a darwin artifact-download timeout and a darwin agent uv_os_get_passwd failure, neither of which reaches a test. PR #32759 (which independently fixed the PipeReader half) has been closed in favor of this one and its remaining suggestion (the Err-state early return) was folded in.
|
Same result on the current head: build 64996 ( Nothing about the diff has changed since my comment above: it is complete, every reviewer finding is resolved, and the regression tests are CI-verified on every Linux lane including x64-asan and Alpine/musl. A push cannot fix this; it needs someone with Buildkite access to retry that shard (or look at that agent). |
|
Correction to my last comment now that build 64996 has finished: it ended with a second red job, the So the completed build is 284 of 286 jobs green, and both red jobs are the two darwin agents. Every Linux lane, including x64-asan and Alpine/musl, is green. |
…Reader (#32848) ## Problem `debug_assert!(process.is_some())` in `PipeReader::try_signal_done_to_cmd` fires in Bun Shell under syscall faults. Regression introduced by #32754 (reported by syscall-fault-injection fuzzing against `origin/main@df92f8f`, never seen on the previous build). ``` panic: assertion failed: process.is_some() at src/runtime/shell/subproc.rs:2189 <PipeReader>::try_signal_done_to_cmd shell/subproc.rs:2189 ``` ## Repro `sleep 5` stands in for any child whose stdout stays open: 1. `PipeReader::start` registers the poll (epoll_ctl succeeds), the eager `read_all()` begins. 2. `recv()` returns a chunk, then `EAGAIN`. 3. `read_with_fn`'s EAGAIN arm calls `register_poll()` BEFORE delivering the drained chunk. That `epoll_ctl` fails with `ENOMEM`. 4. `on_reader_error` -> `try_signal_done_to_cmd` -> `buffered_output_close` -> `close_io` -> `PipeReader::detach()` sets `process = None`. 5. `read_with_fn` then still delivers `on_read_chunk(chunk, Drained)`, which retries `register_poll()`; the resulting second `on_reader_error` re-enters `try_signal_done_to_cmd`. 6. `process` is `None`, the assert fires. Debug log from the repro: ``` [sys] register: FilePoll readable (13[socket]) <- second epoll_ctl, fails [shell_subproc] PipeReader(0x...250) onReaderError errno: 12 [shell_subproc] signalDoneToCmd (stdout) isDone=true [shell_subproc] PipeReader(0x...250, stdout) detach() <- process = None [shell_subproc] PipeReader(0x...250) onReadChunk(chunk_len=11, has_more=drained) [sys] register: FilePoll readable (13[socket]) <- third epoll_ctl, fails [shell_subproc] PipeReader(0x...250) onReaderError errno: 12 [shell_subproc] signalDoneToCmd (stdout) isDone=true panic: assertion failed: process.is_some() ``` ## Cause The `None` backref set by `PipeReader::detach()` is the intentional latch for late terminal callbacks (its own comment says so), and `try_signal_done_to_cmd` already handles it with the `if let Some(proc)` right below the assert. The assert claims a single-shot invariant that the buffered reader never provided. It was unreachable before #32754 only by accident: without that PR's `Arc` keepalive in `Readable::start_pipe_reader`, step 4 dropped the last reference to the `PipeReader`, so step 5 was a use-after-free (one of the two UAFs that PR fixed). Now that the reader correctly survives, the latent second callback reaches the assert. Same in the original Zig (`if (bun.Environment.allow_assert) assert(this.process != null);` followed by `if (this.process) |proc|`). ## Fix Delete the assert. The `if let Some(proc)` is the real guard; a detached reader yields `Yield::Suspended`, and the stream is already closed and accounted for by the first signal. ## Test New mode pair in `test/js/bun/shell/shell-pipe-read-fault.test.ts`'s `LD_PRELOAD` shim: - `SHELL_FAIL_EPOLL_AFTER`: the first `epoll_ctl` ADD on each AF_UNIX socket succeeds, every later ADD/MOD on that fd fails with `ENOMEM`. - `SHELL_RECV_ONE_CHUNK`: the first `recv()` on each AF_UNIX socket returns a fixed chunk, later ones return `EAGAIN`, so the "read some bytes, then stall" interleaving is deterministic instead of racing the child. Without the fix the new test fails in ~800ms with `panic: assertion failed: process.is_some()`; with it all 5 tests in the file pass. The rest of `test/js/bun/shell/` is unchanged (797 pass; the 21 failures in my container reproduce identically on the pre-#32754 binary and are root-user / slow-ASAN environment issues). Reported by a syscall-fault-injection fuzzing campaign as a regression between `c4d6844e7` and `df92f8fd`.
…ll-driven read (#32986) ## Problem Heap use-after-free in Bun Shell when `epoll_ctl` fails while re-registering a pipe's `FilePoll` from a poll-driven read. Found by syscall-fault-injection fuzzing against `origin/main`. Follow-up to #32754, which fixed the same failure on the eager spawn-time read path. ``` ERROR: AddressSanitizer: heap-use-after-free READ of size 1, thread T0 #0 <bun_io::pipe_reader::BufferedReaderVTable>::link io/PipeReader.rs:105 #1 <bun_io::pipe_reader::BufferedReaderVTable>::on_read_chunk io/PipeReader.rs:125 #2 <bun_io::pipe_reader::PosixBufferedReader>::read_with_fn io/PipeReader.rs:890 #3 <bun_io::pipe_reader::PosixBufferedReader>::read_socket io/PipeReader.rs:576 #4 <bun_io::pipe_reader::PosixBufferedReader>::on_poll io/PipeReader.rs:529 #5 __bun_run_file_poll runtime/dispatch.rs:677 freed by: <alloc::sync::Arc<bun_runtime::shell::subproc::PipeReader>>::drop ``` ## Repro 1. `PipeReader::start` registers the poll and the eager spawn-time `read_all()` hits `EAGAIN`, so `read_with_fn`'s `EAGAIN` arm re-registers the poll and the spawn returns. 2. The child writes to stdout and the poll fires. `__bun_run_file_poll`'s `BUFFERED_READER` arm dispatches straight into `PosixBufferedReader::on_poll` with a bare `&mut *h` and no keepalive. 3. `read_with_fn` drains the chunk, `recv()` returns a real `EAGAIN`, and `register_poll()` issues another `epoll_ctl`, which fails (`ENOMEM` in the repro). 4. `register_poll` dispatches `on_reader_error`. The shell `PipeReader::on_reader_error` signals the `Cmd`, the `Readable::Pipe` `Arc` is dropped, and the callback's own `guard_from_raw` keepalive becomes the last reference. The code already documents this: "Dropping `guard` is the matching `deref()`; may free `this`." 5. Back in `read_with_fn`, the `EAGAIN` arm still delivers the drained head: `parent.vtable.on_read_chunk(.., ReadState::Drained)` reads the freed vtable. Traced with the test's `LD_PRELOAD` shim: ``` [shim] epoll_ctl(ADD fd=13) unix call#1 -> ok PipeReader::start [shim] recv(fd=13) unix call#1 -> EAGAIN eager read, inside spawn [shim] epoll_ctl(MOD fd=13) unix call#2 -> ok re-register; spawn returns [shim] recv(fd=13) unix call#2 poll fired: the child's bytes [shim] recv(fd=13) unix call#3 real EAGAIN [shim] epoll_ctl(MOD fd=13) unix call#3 -> ENOMEM register_poll fails [shell_subproc] PipeReader(0x..250) onReaderError errno: 12 [shell_subproc] PipeReader(0x..250, stdout) detach() [shell_subproc] PipeReader(0x..250, stdout) deinit() ==ERROR: AddressSanitizer: heap-use-after-free ``` ## Cause `register_poll()`'s failure path dispatches `on_reader_error`, which the `BufferedReaderParent` contract explicitly allows to free the parent, but `register_poll` gave the caller no way to know that happened. `read_with_fn`'s `EAGAIN` arm is the only call site that touches the reader afterwards; every other `register_poll()` is in tail position. The `SAFETY` comment above the `parent` rebind claimed the parent is "never freed mid-call", which holds for `on_read_chunk` re-entry but not for `on_reader_error`. #32754 covered this exact sequence on the eager spawn-time entry by holding an `Arc<PipeReader>` across `start()` and `read_all()` in `Readable::start_pipe_reader`. The epoll dispatch has no equivalent keepalive, so the poll-driven entry was still exposed. ## Fix `PosixBufferedReader::register_poll()` now returns whether registration succeeded. `false` means `on_reader_error` was dispatched and `self` must not be touched again, so `read_with_fn`'s `EAGAIN` arm returns there instead of delivering the drained head to a possibly freed parent. The stream has already been completed with the registration error at that point, so nothing is lost. All other `register_poll()` call sites are tail calls and discard the result. ## Test Two new modes in `test/js/bun/shell/shell-pipe-read-fault.test.ts`'s `LD_PRELOAD` fault shim: - `SHELL_RECV_EAGAIN_FIRST=1`: the first `recv()` on each `AF_UNIX` socket returns `EAGAIN`, pushing the first successful read off the eager spawn-time `read_all()` and onto the epoll dispatch. - `SHELL_FAIL_EPOLL_FROM=N`: the Nth and later `epoll_ctl` `ADD`/`MOD` on each `AF_UNIX` socket fail with `ENOMEM`. `N=3` lets the initial registration and the eager read's re-registration succeed, then fails the first poll-driven one. The new test is `skipIf(!isASAN)` because the use-after-free is only reliably observable under ASAN. With `src/io/PipeReader.rs` reverted to `main` it fails in ~1.1s with the `heap-use-after-free` above; with the fix all 6 tests in the file pass. ## Out of scope Shell `PipeReader::on_read_chunk` also calls `self.reader.register_poll()` from a `&mut self` method whose stated contract is that it never frees `self`. If that inner registration fails, the same free can happen under `read_with_fn`'s mid-loop flush instead of its `EAGAIN` arm. Reaching it needs a large (>32 KB) burst in one poll wake; I have not reproduced it, so it is not changed here.
Problem
Two heap use-after-frees in the Bun Shell, both from the same root cause: a pipe error surfacing synchronously from inside the spawn call tears down state that the spawn frame is still using. Found by syscall-fault-injection fuzzing against
origin/main(release-asan).recv()failing on the eager pipe read frees theShellSubprocess:epoll_ctl()failing during poll registration frees thePipeReader:Cause
Cmd::transition_to_execstagesSubprocExec.interp = nullbefore callingspawn_async, specifically so a synchronousCmd::on_exitfired from inside the spawn cannot drive the trampoline intoCmd::deinitwhile the spawn frame still holds theShellSubprocess. The backref is published, and anyDonestate picked up, only after the spawn returns.The pipe callbacks bypass that gate in two ways:
PipeReadercarries its owninterpbackref, populated atPipeReader::createtime.spawn_maybe_sync_impldoes an eagerread_all()on the freshly created stdout/stderr pipes, so if thatrecv()fails (ENOMEM/EMFILE/ENFILE),on_reader_errorreachesCmd::buffered_output_close, which records the errno as the exit code, closes the stream, and, once every piped stream is closed, returnsYield::Next.PipeReader::finish_after_state_setruns that with its own ungatedinterp, landing inCmd::deinitand freeing theBox<ShellSubprocess>thatspawn_maybe_sync_implkeeps dereferencing. On a release build the follow-on access is a deterministicpanic: expected Node::Cmd at Node#2, got Free.The same
on_reader_errorcan also fire from insidePipeReader::start:PosixBufferedReader::register_pollreports anepoll_ctlfailure through the vtable andPosixBufferedReader::startthen returnsOk(())regardless.Cmd::buffered_output_closecallsclose_io, andReadable::finalizedrops theReadable::Pipeslot'sArc<PipeReader>from inside the callback, so the callback's own guard becomes the last reference and thePipeReaderis freed when it returns.PipeReader::startthen readsself.reader.handleand writesself.reader.flagson the freed object, andReadable::start_pipe_readerheld no keepalive of its own.Fix
Two pieces, one per victim:
Cmd::buffered_output_closenow applies the same null-interpgateCmd::on_exitalready has: ifexec.interpis still null, the command is markedDonebut the Yield is suspended, andtransition_to_exec's existingCmdState::Donecheck resumes it once the spawn has unwound. This is the single choke point for every pipe-side completion path (on_reader_error,on_reader_done,on_captured_writer_done,CapturedWriter::on_iowriter_chunk).Readable::start_pipe_readernow holds anArc::cloneof thePipeReaderacrossstart()andread_all(), since both can complete the reader synchronously and drop theReadable::Pipeslot's ref. This is the same keepalivePipeReader::guard_from_rawalready provides for theon_reader_{done,error}callbacks themselves.Also deleted in this PR:
Cmd::spawn_arena_freed. It was only ever written and then discarded withlet _ =, and the comment that justified keeping it ("initSubprocsetsspawn_arena_freed = truebefore any pipe can close") states exactly the premise this bug disproves. Theexec.interpgate now encodes the pre/post-spawn split directly.Test
test/js/bun/shell/shell-pipe-read-fault.test.tsuses anLD_PRELOADshim (same pattern asserve-epoll-add-fail.test.ts) with two env-selected fault modes:SHELL_FAIL_RECVmakesrecv()returnENOMEM, andSHELL_FAIL_EPOLLmakesepoll_ctl(ADD|MOD)on pipe-like fds returnENOMEM. The epoll mode interposessyscall(2)becauseFilePollregistration goes through the rawsyscall(SYS_epoll_ctl, ...)wrapper rather than the libcepoll_ctlsymbol. Four tests cover recv and epoll each with stdout-only and stdout+stderr capture pipes. Linux-only, skipped when no C compiler is present.Each half of the fix is independently load-bearing:
Existing shell suites (
bunshell.test.ts,bunshell-default.test.ts,shelloutput.test.ts,epipe.test.ts,exec.test.ts,file-io.test.ts,pipeline_stack.test.ts,yield.test.ts,shell-blocking-pipe.test.ts, and others, 560 tests) still pass.