Repository navigation
fix(#4654): stop exiting a thread per burst — IoPool lanes keep their threads; ThreadPool workers kept alive - #6259
Conversation
… threads; ThreadPool workers kept alive A thread EXIT triggers a CoreCLR defect: a thread that touched a collectible [ThreadStatic] frees, on exit, a loader handle in an unrelated LIVE collectible context, leaving its GC static dangling (Doc/Architecture/CollectibleThreadStaticHandleReuse). The runtime fix is upstream; this removes the gratuitous exits this process produces. - LimitedConcurrencyLevelTaskScheduler: a dedicated lane thread now parks when its queue is empty and is woken by the next leaf, instead of exiting per drain; IoPool's disposal releases the threads via Complete(). - Chart: DOTNET_ThreadPool_ThreadsToKeepAlive=-1 so pool workers stop retiring after 20 s idle (measured: 18 workers retire within 3 s at a 500 ms timeout without it, all 18 remain with it). - BlockingLaneThreadsAreKeptTest: 20 sequential bursts on one thread (negative control: the exit-per-drain shape uses 20), a warm lane still runs its full cap, disposal releases the kept threads. - Doc: the 2026-10-07 09:53Z memex-cloud SIGSEGV read from createdump output (crashing thread young, no managed frames, addr 0x8; the dump died with its emptyDir on the 13:09Z roll); pearl's restarts are managed OOM. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
🟡 Changes recommended
Disposal reports completion before retained lane threads have exited, violating teardown safety guarantees.
4 open findings
What changed in this PR
Introduces a stopgap for #4654 by reducing thread exits that can trigger the CoreCLR collectible thread-static defect.
Changes:
- Keeps IoPool lane threads parked between bursts and ThreadPool workers alive.
- Adds lifecycle/concurrency tests and gate exemptions.
- Updates crash analysis and pooling documentation.
| File | Description |
|---|---|
BlockingLaneThreadsAreKeptTest.cs |
Tests thread reuse, concurrency, and disposal. |
HandWovenGateRatchetGuard.cs |
Sanctions the IoPool scheduler’s monitor. |
LimitedConcurrencyLevelTaskScheduler.cs |
Parks and wakes dedicated lane threads. |
IoPool.cs |
Completes the scheduler during disposal. |
ControlledIoPooling.md |
Documents retained lane threads. |
CollectibleThreadStaticHandleReuse.md |
Records remedies and production evidence. |
deployment.yaml |
Prevents ThreadPool worker retirement. |
🧠 Review effort: Balanced
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| // The lane threads were kept parked between bursts; no blocking leaf can run any more | ||
| // (_gateUsers, _inFlight and _blockingInFlight are all zero), so let them exit. Non-blocking: | ||
| // it pulses and returns, and is safe even when this runs ON a lane thread. | ||
| _blockingScheduler.Complete(); |
There was a problem hiding this comment.
Real defect, fixed in 23fd328. Complete() now takes a drained callback: the scheduler invokes it once, outside its lock, from the last lane thread right after it takes the live drain-loop count to zero (or inline from Complete() when no lane thread is alive). IoPool publishes Disposed only from that callback, so the signal again means no pool thread remains, and nothing blocks or joins.
| [Fact(Timeout = 60_000)] | ||
| public async Task Disposal_ReleasesTheKeptThreads() | ||
| { | ||
| var pool = new IoPool(2); |
There was a problem hiding this comment.
Fixed in 23fd328: the pool in Disposal_ReleasesTheKeptThreads is now a using local; the explicit Dispose stays valid (idempotent).
| - **ThreadPool workers retire** after about 20 s idle. Those are the threads hub turns and | ||
| deserialisation run on, so they touch NodeType-typed `ArrayPool<T>` constantly. | ||
| - **Every blocking lane of the pools `IoPoolRegistry` registers starts a fresh thread per burst, and | ||
| - **Until #4654, every blocking lane of the pools `IoPoolRegistry` registers started a fresh thread per burst (it now keeps them — see Remedies), and |
There was a problem hiding this comment.
Fixed in 23fd328: the whole bullet is rewritten in historical tense, including the new Thread(_ => DrainQueue()) continuation.
| | **Sync-blocking / CPU** (`InvokeBlocking`) | Dedicated `LimitedConcurrencyLevelTaskScheduler` | Blocking work holds a real thread for its whole duration. The limited-concurrency scheduler runs at most *cap* leaves at a time, **on threads it starts itself** (`mw-io-lane`; the CPU lane's are `mw-cpu-lane`), never ThreadPool workers. It used to BORROW pool workers, and a cap never protected the pool: with caps up to 256 over a pool whose minimum is the core count, a burst of blocking reads held every worker the grain turns need — see [Blocking Leaves Off the ThreadPool](../BlockingLeavesOffTheThreadPool). | | ||
| | **Sync-blocking / CPU** (`InvokeBlocking`) | Dedicated `LimitedConcurrencyLevelTaskScheduler` | Blocking work holds a real thread for its whole duration. The limited-concurrency scheduler runs at most *cap* leaves at a time, **on threads it starts itself** (`mw-io-lane`; the CPU lane's are `mw-cpu-lane`), never ThreadPool workers, and KEEPS them parked between bursts until the pool is disposed — a thread exit per burst was the trigger of a CoreCLR defect, see [Collectible Thread-Static Handle Reuse](../CollectibleThreadStaticHandleReuse). It used to BORROW pool workers, and a cap never protected the pool: with caps up to 256 over a pool whose minimum is the core count, a burst of blocking reads held every worker the grain turns need — see [Blocking Leaves Off the ThreadPool](../BlockingLeavesOffTheThreadPool). | | ||
|
|
||
| > The two shapes use threads differently, on purpose. **Async leaves reuse the ThreadPool the framework already uses**, with a governor in front of it — they hold a worker only between awaits, which is what the pool is built for. **Blocking leaves do not touch the ThreadPool at all**: they run on short-lived threads the pool's scheduler starts (one per concurrent drain loop, exiting when its queue is empty), because a blocked thread is exactly what Orleans cannot coordinate with when it is one of its own workers. The thread COUNT is the same either way — a burst of N blocking leaves needs N threads — but a borrowed worker is taken from the grain turns and replaced only by the pool's slow injection, while a started thread is taken from no one. |
There was a problem hiding this comment.
Fixed in 23fd328: the following paragraph now says blocking leaves run on dedicated threads kept parked between bursts and exiting only when the pool is disposed.
Test Results (shard 3) 4 files ± 0 4 suites ±0 7m 25s ⏱️ - 2m 24s Results for commit 23fd328. ± Comparison against base commit d1a2a9d. This pull request removes 72 and adds 110 tests. Note that renamed tests count towards both.♻️ This comment has been updated with latest results. |
…review follow-ups on docs and test Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Test Results (shard 4) 4 files ± 0 4 suites ±0 16m 41s ⏱️ -1s Results for commit 23fd328. ± Comparison against base commit d1a2a9d. This pull request removes 148 and adds 88 tests. Note that renamed tests count towards both. |
Test Results 22 files ± 0 22 suites ±0 49m 9s ⏱️ - 2m 47s Results for commit 23fd328. ± Comparison against base commit d1a2a9d. This pull request removes 40 and adds 31 tests. Note that renamed tests count towards both. |



Refs #4654. This is a stopgap against a runtime defect, not the fix. The root cause is in CoreCLR and has to be fixed upstream (
FreeTLSIndicesForLoaderAllocatornever scrubs threads'pLoaderHandles). Full write-up:Doc/Architecture/CollectibleThreadStaticHandleReuse. This PR removes the thread exits this process produces itself. The defect needs a thread exit to fire.What changes
LimitedConcurrencyLevelTaskScheduler. This is the internal scheduler behind everyIoPoolblocking lane.QueueTaskwakes it; the waker counts down, so N queued tasks wake N parked threads. A thread is started only when none is parked and the lane is below its cap.IoPool.TryFinishDisposalcalls the newComplete()once no blocking leaf can run, and the kept threads exit then.Complete()does not block.Chart.
DOTNET_ThreadPool_ThreadsToKeepAlive=-1, so ThreadPool workers stop retiring after about 20 s idle. I verified the runtime honours the setting on .NET 10, with a 500 ms thread timeout and a 40-item burst:This takes effect on the next roll of a deployment that renders this chart.
HandWovenGateRatchetGuard. The scheduler file joinsIoPool.csin the sanctioned register, with a stated reason. It is IoPool's internal half and is constructed only byIoPool. Its oneMonitor.Waitis an idle lane thread, which the scheduler owns, waiting for work. It is never a hub turn, a grain turn or a ThreadPool worker.Docs.
CollectibleThreadStaticHandleReusenow records both remedies as applied.createdumpoutput in Loki:45dehad no managed frames and was among the newest threads;signo 11 code 1 addr 0x8;memex-dumpsemptyDir, and the 13:09Z roll replaced the pod.signo 6.ControlledIoPoolingis updated to match.Verification
dotnet build -c Release -warnaserrorgives 0 warnings and 0 errors on each of:MeshWeaver.Mesh.Contract,MeshWeaver.Hosting.Test,MeshWeaver.Documentation.Test,MeshWeaver.Documentation.BlockingLaneThreadsAreKeptTest, 3/3 pass:IsAlive == false).LaneLoopreverted to exit on an empty queue, all 3 fail. The first fails with "found 20" threads. With the code restored, all 3 pass.FullyQualifiedName~Blocking: 21/21 pass;FullyQualifiedName~IoPool: 68/68 pass;HandWovenGate: 4/4 pass.helm template deploy/helmrenders the new env var.Not established
For the maintainer
NT_SIGINFOon both memex and memex-cloud./data/dumpsoff the emptyDir before rolls, or moving dumps to a volume that survives the pod. Otherwise every production dump dies with its pod.No recycle is needed: this is process-level, and a roll applies it.
Pairs-with: none — no public type or member removed; only internal members added.
Implementers: none — no interface member added.
🤖 Generated with Claude Code