GH-4167: one BatchingProcessor per batched message type, not one per racing caller - #4168
Merged
Merged
Conversation
…racing caller BatchingOptions.BuildHandler was an unguarded `if (_handler != null)` lazy init, and it runs on the message-handling path rather than at bootstrap. Callers reach it through HandlerPipeline's LightweightCache<Type, IExecutor>, whose indexer does not lock: two concurrent misses each invoke the factory AND each returns its own instance. That is harmless for a stateless executor, which is why the cache is fine everywhere else. A BatchingProcessor is not stateless -- each instance owns a separate BatchingChannel buffer, its own flush Timer, and two Blocks with live worker tasks. Two instances means members of one logical batch are split across two buffers and flushed as separate batches, and the losing instance is never handed back to anyone, so nothing ever disposes it: its timer and worker tasks leak for the life of the process. Instrumented, two instances were built on two threads and a two-member batch arrived as [one] | [two]. Fixed with a double-checked lock; _handler is now volatile for the unlocked read. concurrent_BuildHandler_yields_one_shared_processor asserts reference equality across 8 racing callers and fails on every run without this change, against the currently referenced JasperFx 2.56.0 -- this race is reachable today, it is not introduced by the JasperFx fix. What the JasperFx fix (#714) changes is only how easy it is to hit: its inline Block continuations were running the first messages on the publisher's thread and serializing them, which is why the end-to-end batch test still passes on 2.56.0 and only becomes a real guard after the bump. Also fixes Bug_3399's wait condition, which was a latent test race the same inline execution was hiding. A WaitFor* condition REPLACES quiescence rather than adding to it, so WaitForMessageToBeReceivedAt released the session at receipt, before any handler ran. It now waits for execution -- of BOTH separated chains for the batched array, since waiting on one just trades a receive/execute race for a which-chain-won race. Verified: CoreTests 2664 passed / 2 skipped / 0 failed against a locally built JasperFx carrying #714, and green on stock 2.56.0. Full wolverine.slnx builds clean under -c Release -f net9.0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
jeremydmiller
added a commit
that referenced
this pull request
Aug 27, 2026
Picks up #714: a Block no longer runs its action on the publisher's thread. Block built its channel with AllowSynchronousContinuations = true, so a reader parked in WaitToReadAsync was resumed by TryWrite on the publishing thread, and Post() executed the action inline instead of enqueuing it. For Wolverine that meant a buffered local queue got no parallelism at all after its workers went idle: a burst of 20 published messages all ran inline and serialized on the publishing thread, and that thread -- a broker listener loop, an HTTP request -- stalled for the full duration of each handler. Adds the local-queue reproduction from GH-4167 as this bump's acceptance test. It was deliberately held out of #4168 because it cannot pass against 2.56.0 and would have made CI red; it passes from 2.57.0 onward. Note the reported ".NET 10 regression" framing is only half right, and the test comments record why: an UNBOUNDED channel (what a buffered local queue uses, GH-3287) ran continuations inline on every runtime. dotnet/runtime#116021 changed the BOUNDED case, which is what broker-backed BufferedReceivers and DurableReceiver use. The companion Wolverine fix for the batching race this exposes is already on main (29d49a6), so this bump is safe to take. Verified against the published 2.57.0 packages, not a local build: CoreTests 2664 passed / 2 skipped / 0 failed, and wolverine.slnx clean under -c Release -f net9.0 after a cleared HTTP cache and full restore. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This was referenced Aug 28, 2026
This was referenced Sep 4, 2026
This was referenced Sep 11, 2026
This was referenced Sep 14, 2026
This was referenced Sep 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Found while verifying the JasperFx fix for #4167 (JasperFx/jasperfx#714). This is a live bug on
maintoday — it does not depend on that fix.The bug
BatchingOptions.BuildHandleris an unguarded lazy init, and it runs on the message-handling path, not at bootstrap:Callers reach it through
HandlerPipeline'sLightweightCache<Type, IExecutor>, whose indexer does not lock — two concurrent misses each invoke the factory and each returns its own instance.For a stateless executor that is harmless, which is why the cache is fine everywhere else. A
BatchingProcessoris not stateless: each instance owns a separateBatchingChannelbuffer, its own flushTimer, and twoBlocks with live worker tasks.Instrumented — two instances on two threads, and a two-member batch arrives as two batches:
Two consequences:
Timerand worker tasks live for the process lifetime.The fix
Double-checked lock, with
_handlermadevolatilefor the unlocked read.Tests
concurrent_BuildHandler_yields_one_shared_processorraces 8 callers intoBuildHandlerand asserts reference equality. It fails on every run without the fix, against the currently referenced JasperFx 2.56.0 — deliberately written to exercise the race directly rather than through message publishing, so it does not depend on the JasperFx release to be meaningful.The end-to-end companion (
concurrent_first_messages_still_assemble_a_single_batch) still passes on 2.56.0 even unfixed, because inlineBlockcontinuations serialize the first messages. It becomes a real guard after the JasperFx bump; the comment says so.Also:
Bug_3399's wait conditionA latent test race the same inline execution was hiding. A
WaitFor*condition replaces quiescence rather than adding to it, soWaitForMessageToBeReceivedAtreleased the session at receipt, before any handler ran, and the assertion raced it. Probed:Batched=0immediately,Batched=1after 1s — the handler runs, the test just asked the wrong question.Now waits for execution, of both separated chains for the batched array; waiting on one only trades a receive/execute race for a which-chain-won race.
Verification
CoreTests2664 passed / 2 skipped / 0 failed against a locally built JasperFx carrying sp_MSdropconstraints is not available on Azure SQL Database #714CoreTestsgreen on stock JasperFx 2.56.0wolverine.slnxclean under-c Release -f net9.0Follow-up
The GH-4167 local-queue reproduction test is not in this PR — it asserts the publisher thread never runs the handler, which cannot pass until JasperFx ships #714. It should land with the JasperFx version bump, as that bump's acceptance test.
🤖 Generated with Claude Code