Skip to content

http(h3): do not fail a retried request twice when the reconnect fails - #42900

Merged
Jarred-Sumner merged 2 commits into
mainfrom
robobun/406f1a4a/h3-retry-double-fail
Sep 17, 2026
Merged

Jarred-Sumner merged 2 commits into
mainfrom
robobun/406f1a4a/h3-retry-double-fail

Conversation

@robobun

@robobun robobun commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • An HTTP/3 fetch() whose QUIC connection dies before the response header can abort the process. ASan: heap-use-after-free READ of size 8 in HTTPClient::fail_from_h2 (src/http/lib.rs:2108), from ClientSession::retry_or_fail (src/http/h3_client/ClientSession.rs:288). Release builds panic: fetch on the HTTP thread holds a ticket.
  • The retry queues the request on a new session through ClientContext::connect. When no connection opens, connect fails that session with PendingConnect::fail_session, which fails every request queued on it. That dispatch frees the AsyncHTTP the client is part of. Then the retry fails the same client again.

Fix

  • connect takes the request back off the session before it fails that session. A false return leaves the request on no session, so the caller is its only failure path, which is what the other two callers assume.
  • Correct because the session is one call old: the request enqueue just queued is its only entry, so detach leaves fail_session nothing to fail. The teardown, the registry removal and the session's last reference do not change.
  • The retried request keeps the error of the stream that closed. start_ still reports ConnectionRefused for its own failed connect.
  • Verified: test/js/web/fetch/fetch-http3-client.test.ts, one new test (main aborts with an empty stdout). Also the three other fetch-http3-* suites, serve-http3 and serve-protocols.

Background

  • The h3 fetch client pools one QUIC connection per origin. retry_or_fail re-sends a stream that closed before any response header, once, on a fresh connection.
  • ClientContext::connect finds a pooled connection or opens one, and queues the request. enqueue binds a Stream to the request before the QUIC connect, because that stream has to exist when the handshake completes.
  • HTTPClient::start_ sets defer_terminal_dispatch_until_connecting_is_complete before its own connect call, so a failure inside that frame is recorded and dispatched later. That flag is why the two initial connect sites survived the double failure.
Notes

Fail-before. With src/ and packages/ back on 55c11065f2, the new test gives exitCode: 1 and an empty stdout. That run, the passing run and the suites above were on 55c11065f2 plus this change, built with LLVM 21. The branch has since merged main, which needs LLVM 23 (#42851). The build environment used here does not have it, so on the merged tree only cargo check and cargo clippy for bun_http were run locally, and CI is the test run for it. The three commits that merge brought in touch none of the files involved. The ASan frames are the report above:

READ of size 8 at 0x... thread T4 (HTTP Client)
  #2 <bun_http::HTTPClient>::fail_from_h2                src/http/lib.rs:2108
  #3 <ClientSession>::retry_or_fail                      src/http/h3_client/ClientSession.rs:288
  #4 h3_client::callbacks::on_conn_close                 src/http/h3_client/callbacks.rs:151
freed by thread T4 (HTTP Client) here:
  #7 <AsyncHTTP>::on_async_http_callback_raw             src/http/AsyncHTTP.rs:783
  #10 <bun_http::HTTPClient>::fail_from_h2               src/http/lib.rs:2122
  #11 <PendingConnect>::fail_session                     src/http/h3_client/PendingConnect.rs:149
  #12 <ClientContext>::connect                           src/http/h3_client/ClientContext.rs:179
  #13 <ClientSession>::retry_or_fail                     src/http/h3_client/ClientSession.rs:287

A release build aborts as well, so the fault is not an ASan artifact: on_async_http_callback_raw resets the client's stage before the dealloc, so the once-only guard in fail_from_h2 cannot stop the second dispatch. Making that guard survive the reset is a separate change.

How the test reaches it. A connect to a resolved hostname probes each address with a throwaway UDP connect(2), and gives up when no entry is reachable (packages/bun-usockets/src/quic.c, us_quic_connect_result). An LD_PRELOAD shim allows the first probe and refuses every later one, so the reconnect fails inside connect. rejectUnauthorized against the suite's self-signed certificate fails the handshake, which is what closes the stream before any header and starts the retry. localhost answers from is_localhost_name as [::1, 127.0.0.1] without the resolver, so no connect waits for DNS, and the shim refuses the IPv6 entry the way a host without an IPv6 route does, which pins both connects to the same address. Linux only, and only where a C compiler exists, like the DPLPMTUD shim test in fetch-http3-syscall-fault.test.ts. 5 runs, 5 passes, about 500 ms each on the debug ASan build.

Other ways to reach the same failure. Any synchronous failure of the QUIC connect does it: a cached resolver error, an IP literal whose family the shared client endpoint cannot serve, lsquic_engine_connect returning NULL, or the shared client UDP endpoint dying on a hard recvmsg error and the poll registration for its replacement failing. The last one needs no resolver, so it reaches this path for an IP-literal origin too. One test is enough: all of them end in the same return false, and the endpoint-replacement route needs several iterations of a loop to line up.

Earlier shape. The first version of this PR removed the retry's failure call instead, and documented connect as owning the request. Review pushed back: it left both if !connect { self.fail(..) } arms in start_ dead, it made the bool unusable by every caller, and it set the opposite contract from #40385, which removes the same double failure from the callee side. This version fixes the callee, which also keeps the closed stream's error in the rejection instead of replacing it with ECONNREFUSED.

Scope. retry_or_fail is also edited by #41564 (a retry budget) and #42579 (no replay of a non-idempotent request), and #40598 changes which pre-header closes retry. None of them touch this branch, so this applies on top of any of them, and #40385 keeps the same contract.


no test proof · iteration 1 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/web/fetch/fetch-http3-client.test.ts

@coderabbitai

coderabbitai Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 5ef38307-7c05-4f61-a974-fbb28f476212

📥 Commits

Reviewing files that changed from the base of the PR and between b111217 and a4e0919.

📒 Files selected for processing (2)
  • src/http/h3_client/ClientContext.rs
  • test/js/web/fetch/fetch-http3-client.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 0 remain after this review.


Walkthrough

ClientContext::connect now detaches failed H3 streams before session failure. The fetch tests add gated coverage for failed HTTP/3 reconnects using a compiled UDP probe shim.

Changes

HTTP/3 retry failure handling

Layer / File(s) Summary
Connection ownership and failure handling
src/http/h3_client/ClientContext.rs
ClientContext::connect documents its ownership and failure semantics. Failed connection setup detaches the H3 stream before session cleanup and reports ConnectionRefused.
Failed reconnect test coverage
test/js/web/fetch/fetch-http3-client.test.ts
A Linux/compiler-gated subprocess test uses an embedded C shim to control UDP probes and verifies distinct handshake and connection-refused errors, clean exit, and no signal termination.

Suggested reviewers: jarred-sumner

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to a4e09

The HTTP/3 reconnect failure path now rejects each affected request once without retaining invalid stream ownership. No actionable merge-blocking risk remains.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: preventing a retried HTTP/3 request from being failed twice when reconnecting fails.
Description check ✅ Passed The description explains the problem, fix, behavior changes, verification, test coverage, and known environment limits. It does not use the exact template headings, but it provides the required inform…

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline note, I also checked the two other ways connect could bite the retry caller: the true outcomes (pooled-session enqueue and the sync ConnectResult::Socket arm) do not dispatch a terminal result before returning, and the ConnectResult::Err arm's unregister is index-and-pointer checked, so the old session's own teardown in on_conn_close unregistering again is a no-op rather than a registry corruption.

Extended reasoning...

The core change is a three-line removal of the redundant fail_from_h2 after a false return from ClientContext::connect, which already fails the request via PendingConnect::fail_session and frees the owning AsyncHTTP. I read connect in ClientContext.rs:116-193 to confirm the ownership claim: the pooled-reuse path only enqueues, the sync socket path only stores the socket pointer, and only the Err arm dispatches a terminal failure. I also confirmed unregister (ClientContext.rs:195-206) tolerates a second call for an already-removed session. Not approving because an inline finding is posted and a further verified finding was dropped from the posted set; a human should weigh the ECONNREFUSED-vs-original-error reporting choice and the LD_PRELOAD-based test's CI robustness.

2 verified lower-impact observations (convention, logging or cleanup points) were not posted.

Comment thread test/js/web/fetch/fetch-http3-client.test.ts Outdated
@robobun

robobun commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 10:56 AM PT - Sep 16th, 2026

✅ @robobun, your commit a4e091918f42c9b332eb73ae539b7e763520f440 passed in Build #116626! 🎉


🧪   To try this PR locally:

bunx bun-pr 42900

That installs a local version of the PR into your bun-42900 executable, so you can run:

bun-42900 --bun

@robobun
robobun force-pushed the robobun/406f1a4a/h3-retry-double-fail branch from b111217 to f6909ee Compare September 16, 2026 17:15
Comment thread src/http/h3_client/ClientContext.rs Outdated
Comment thread src/http/h3_client/ClientContext.rs Outdated
A stream that closes before any response header is retried on a fresh
connection. When ClientContext::connect could not open one it tore the
new session down through PendingConnect::fail_session, which calls
fail_from_h2 on every request queued on it. That dispatch frees the
AsyncHTTP the client is part of, and ClientSession::retry_or_fail then
failed the same client again through the freed memory.

connect now takes the request back off the session it queued it on
before it fails that session, so a false return leaves the request on no
session and the caller is the only one to fail it. The request keeps the
error of the stream that closed, and the initial connect sites in start_
keep reporting ConnectionRefused.
@robobun
robobun force-pushed the robobun/406f1a4a/h3-retry-double-fail branch from f6909ee to bc7d5ab Compare September 16, 2026 17:17
Comment thread src/http/h3_client/ClientContext.rs
@robobun

robobun commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator Author

Reworked after review: the fix moved from the caller to the callee.

ClientContext::connect now takes the request back off the session it queued it on before it fails that session, so a false return leaves the request on no session. The retry keeps its own failure call, and the request keeps the error of the stream that closed (HTTP3HandshakeFailed in the test) instead of the connect's ConnectionRefused. The first version removed the retry's failure call instead, which left both if !connect { self.fail(..) } arms in start_ dead and set the opposite contract from #40385.

How it reproduces, in one test on the debug ASan build:

  1. An in-process Bun.serve({ http3: true }) origin with the suite's self-signed certificate.
  2. A child process fetches it by hostname with protocol: "http3" and rejectUnauthorized: true. The handshake fails, so the request's stream closes before any response header and the retry runs.
  3. An LD_PRELOAD shim allows the first throwaway UDP connect(2) route probe and refuses every later one, so the reconnect fails inside connect().

With src/ and packages/ on this PR's base the child aborts with AddressSanitizer: heap-use-after-free and prints nothing. With the fix both requests reject once and the child exits 0. 5 runs, 5 passes, about 500 ms each.

Those runs were on 55c11065f2 plus this change, built with LLVM 21. The branch has since merged main, which needs LLVM 23 (#42851). The build environment used here does not have it, so on the merged tree only cargo check and cargo clippy for bun_http were run locally, and CI is the test run for it. The three commits that merge brought in touch none of the files involved.

The binary-size failure on the earlier builds came from that gap and not from this change: every target differed from main by hundreds of KB, in both directions (darwin-aarch64 -404 KB, freebsd-aarch64 +1.78 MB), with identical sizes on both builds. I did not use [skip size check], because the increase is not intentional and the marker would also hide a real one.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the updated push and didn't find any bugs; the earlier note about the dead DNS TTL knob is addressed in this version. Because the fix sits on the raw-pointer/refcount lifecycle of the h3 client, a maintainer familiar with that code should still take a look.

What was reviewed:

  • Session refcount across the new ConnectResult::Err arm: ClientSession::new (1) + enqueue (2) → detach (1) → fail_session's final deref (0), so the session is freed exactly once and detach's "never the last ref" precondition still holds.
  • detach_with on a stream that has no qstream yet (connect never returned a socket) — only unlinks from pending, nulls client.h3, and frees the Stream.
  • The three connect() callers (start_ x2 in src/http/lib.rs, retry_or_fail) all still fail the client themselves on false, and fail_session now finds pending empty, so no second fail_from_h2 on a freed AsyncHTTP.
  • Test follows the existing LD_PRELOAD shim precedent in fetch-http3-syscall-fault.test.ts (Linux + cc gated, bunEnv.LD_PRELOAD chained, pipes drained concurrently, signalCode === null asserted).
Extended reasoning...

Overview

The PR changes one arm of ClientContext::connect() in /home/claude/bun/src/http/h3_client/ClientContext.rs: when the QUIC connect fails synchronously, the request's just-enqueued Stream is detached from the fresh session before unregister and PendingConnect::fail_session. Previously fail_session drained pending and called fail_from_h2 on the queued client, which dispatches and frees the AsyncHTTP; then every caller of connect() failed the same client again on false. The retry caller (ClientSession::retry_or_fail) has no defer_terminal_dispatch guard, so that was a genuine UAF. The test adds a Linux-only LD_PRELOAD shim that refuses the second UDP connect(2) probe so the retry's reconnect fails inside connect(), and asserts the fixture prints HTTP3HandshakeFailed then ECONNREFUSED and exits 0 without a signal.

Security risks

None specific to this change. It does not touch TLS validation, credential handling, or input parsing; it only reorders teardown so a request is failed once. The rejectUnauthorized: true in the test is used as a handshake-failure trigger, not weakened.

Level of scrutiny

Medium-high: the diff is five lines but sits in raw-pointer, intrusive-refcount lifecycle code on the HTTP thread. I traced the refcounts (new + enqueue → detach → fail_session's final deref) and confirmed detach_with is safe on a stream with no qstream. The &mut HTTPClient held by connect() across detach's write to client.h3 through the stored backref is the same shape the pre-existing start_ callers already have and the backref is derived from that same &mut, so it is not a new aliasing problem. Still, this is the kind of code where a maintainer who owns the h3 client should confirm the ownership story matches their model (including the note in the description that #40385 fixes the same double-failure from the other side).

Other factors

The prior review's only inline note (the unused BUN_CONFIG_DNS_TIME_TO_LIVE_SECONDS env and its comment) is gone in this push. Two github-actions[bot] inline comments landed on the changed lines just before the commit; their content is not visible here, so I cannot confirm whether they were addressed, which is another reason for a human to glance rather than auto-approve. The bug hunt exited on a dry streak with no findings. The test reuses the compiled-shim pattern from fetch-http3-syscall-fault.test.ts, drains stdout/stderr concurrently, and asserts output before exit code as the harness conventions require; I did not run it locally.

@robobun

robobun commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator Author

A note for a reader who arrives from the automated reviews above.

The second automated review could not see whether the two inline notes on ClientContext.rs were handled. They were: the doc comment on connect is one line, the comment in the Err arm is two lines, and all four review threads have a reply and are resolved.

An earlier version of this comment said that the CodeRabbit walkthrough described the first version of this PR. CodeRabbit has since reviewed the current head, and its walkthrough is accurate now.

@Jarred-Sumner
Jarred-Sumner merged commit b27413b into main Sep 17, 2026
10 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the robobun/406f1a4a/h3-retry-double-fail branch September 17, 2026 00:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants