Skip to content

fix(#4654): stop exiting a thread per burst — IoPool lanes keep their threads; ThreadPool workers kept alive - #6259

Merged
rbuergi merged 2 commits into
mainfrom
fix/4654-tls-handle-reuse
Oct 7, 2026
Merged

rbuergi merged 2 commits into
mainfrom
fix/4654-tls-handle-reuse

Conversation

@rbuergi

@rbuergi rbuergi commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Refs #4654. This is a stopgap against a runtime defect, not the fix. The root cause is in CoreCLR and has to be fixed upstream (FreeTLSIndicesForLoaderAllocator never scrubs threads' pLoaderHandles). Full write-up: Doc/Architecture/CollectibleThreadStaticHandleReuse. This PR removes the thread exits this process produces itself. The defect needs a thread exit to fire.

What changes

  1. LimitedConcurrencyLevelTaskScheduler. This is the internal scheduler behind every IoPool blocking lane.

    • Before: it started a thread whenever a leaf reached an idle lane, and that thread exited when the queue drained. That meant one OS thread created and destroyed per burst.
    • After: a lane thread parks on the queue monitor when idle. The next QueueTask wakes it; the waker counts down, so N queued tasks wake N parked threads. A thread is started only when none is parked and the lane is below its cap.
    • IoPool.TryFinishDisposal calls the new Complete() once no blocking leaf can run, and the kept threads exit then. Complete() does not block.
    • The borrowing mode is unchanged: it never parks a ThreadPool worker.
  2. Chart. DOTNET_ThreadPool_ThreadsToKeepAlive=-1, so ThreadPool workers stop retiring after about 20 s idle. I verified the runtime honours the setting on .NET 10, with a 500 ms thread timeout and a 40-item burst:

    • without it, 18 workers were left after the burst and 0 after 3 s idle;
    • with it, 18 were still there after 3 s.

    This takes effect on the next roll of a deployment that renders this chart.

  3. HandWovenGateRatchetGuard. The scheduler file joins IoPool.cs in the sanctioned register, with a stated reason. It is IoPool's internal half and is constructed only by IoPool. Its one Monitor.Wait is an idle lane thread, which the scheduler owns, waiting for work. It is never a hub turn, a grain turn or a ThreadPool worker.

  4. Docs.

    • CollectibleThreadStaticHandleReuse now records both remedies as applied.
    • It also records the 2026-10-07 09:53Z memex-cloud SIGSEGV, read from the createdump output in Loki:
      • the crashing thread 45de had no managed frames and was among the newest threads;
      • signo 11 code 1 addr 0x8;
      • no managed exception type;
      • 54 min uptime.
    • The dump is lost: it was written to the memex-dumps emptyDir, and the 13:09Z roll replaced the pod.
    • The same doc records that pearl's restarts are a different defect: managed OOM, signo 6.
    • ControlledIoPooling is updated to match.

Verification

  • dotnet build -c Release -warnaserror gives 0 warnings and 0 errors on each of: MeshWeaver.Mesh.Contract, MeshWeaver.Hosting.Test, MeshWeaver.Documentation.Test, MeshWeaver.Documentation.
  • New test, BlockingLaneThreadsAreKeptTest, 3/3 pass:
    • 20 sequential bursts, each waiting for the lane to go idle, run on 1 thread;
    • a warm lane still runs its full cap at once and starts no new thread;
    • disposal releases every kept thread (IsAlive == false).
  • Negative control: with LaneLoop reverted to exit on an empty queue, all 3 fail. The first fails with "found 20" threads. With the code restored, all 3 pass.
  • Regression runs:
    • FullyQualifiedName~Blocking: 21/21 pass;
    • FullyQualifiedName~IoPool: 68/68 pass;
    • HandWovenGate: 4/4 pass.
  • helm template deploy/helm renders the new env var.

Not established

  • That this mechanism caused any production crash. No production dump has been read. The 09:53Z one is gone.
  • How many thread exits the portal actually produced per hour. The tid span in the dump is not a thread count, because child processes share the pid space.

For the maintainer

  • File the CoreCLR repro upstream. It is a public act, and the repro is in the doc.
  • After the next roll, verify a ≥ 24 h single-generation window with zero NT_SIGINFO on both memex and memex-cloud.
  • Consider copying /data/dumps off the emptyDir before rolls, or moving dumps to a volume that survives the pod. Otherwise every production dump dies with its pod.

No recycle is needed: this is process-level, and a roll applies it.

Pairs-with: none — no public type or member removed; only internal members added.
Implementers: none — no interface member added.

🤖 Generated with Claude Code

… threads; ThreadPool workers kept alive

A thread EXIT triggers a CoreCLR defect: a thread that touched a collectible
[ThreadStatic] frees, on exit, a loader handle in an unrelated LIVE
collectible context, leaving its GC static dangling
(Doc/Architecture/CollectibleThreadStaticHandleReuse). The runtime fix is
upstream; this removes the gratuitous exits this process produces.

- LimitedConcurrencyLevelTaskScheduler: a dedicated lane thread now parks
  when its queue is empty and is woken by the next leaf, instead of exiting
  per drain; IoPool's disposal releases the threads via Complete().
- Chart: DOTNET_ThreadPool_ThreadsToKeepAlive=-1 so pool workers stop
  retiring after 20 s idle (measured: 18 workers retire within 3 s at a
  500 ms timeout without it, all 18 remain with it).
- BlockingLaneThreadsAreKeptTest: 20 sequential bursts on one thread
  (negative control: the exit-per-drain shape uses 20), a warm lane still
  runs its full cap, disposal releases the kept threads.
- Doc: the 2026-10-07 09:53Z memex-cloud SIGSEGV read from createdump
  output (crashing thread young, no managed frames, addr 0x8; the dump died
  with its emptyDir on the 13:09Z roll); pearl's restarts are managed OOM.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Disposal reports completion before retained lane threads have exited, violating teardown safety guarantees.

4 open findings
What changed in this PR

Introduces a stopgap for #4654 by reducing thread exits that can trigger the CoreCLR collectible thread-static defect.

Changes:

  • Keeps IoPool lane threads parked between bursts and ThreadPool workers alive.
  • Adds lifecycle/concurrency tests and gate exemptions.
  • Updates crash analysis and pooling documentation.
File Description
BlockingLaneThreadsAreKeptTest.cs Tests thread reuse, concurrency, and disposal.
HandWovenGateRatchetGuard.cs Sanctions the IoPool scheduler’s monitor.
LimitedConcurrencyLevelTaskScheduler.cs Parks and wakes dedicated lane threads.
IoPool.cs Completes the scheduler during disposal.
ControlledIoPooling.md Documents retained lane threads.
CollectibleThreadStaticHandleReuse.md Records remedies and production evidence.
deployment.yaml Prevents ThreadPool worker retirement.

🧠 Review effort: Balanced


💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

// The lane threads were kept parked between bursts; no blocking leaf can run any more
// (_gateUsers, _inFlight and _blockingInFlight are all zero), so let them exit. Non-blocking:
// it pulses and returns, and is safe even when this runs ON a lane thread.
_blockingScheduler.Complete();

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Real defect, fixed in 23fd328. Complete() now takes a drained callback: the scheduler invokes it once, outside its lock, from the last lane thread right after it takes the live drain-loop count to zero (or inline from Complete() when no lane thread is alive). IoPool publishes Disposed only from that callback, so the signal again means no pool thread remains, and nothing blocks or joins.

[Fact(Timeout = 60_000)]
public async Task Disposal_ReleasesTheKeptThreads()
{
var pool = new IoPool(2);

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 23fd328: the pool in Disposal_ReleasesTheKeptThreads is now a using local; the explicit Dispose stays valid (idempotent).

- **ThreadPool workers retire** after about 20 s idle. Those are the threads hub turns and
deserialisation run on, so they touch NodeType-typed `ArrayPool<T>` constantly.
- **Every blocking lane of the pools `IoPoolRegistry` registers starts a fresh thread per burst, and
- **Until #4654, every blocking lane of the pools `IoPoolRegistry` registers started a fresh thread per burst (it now keeps them — see Remedies), and

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 23fd328: the whole bullet is rewritten in historical tense, including the new Thread(_ => DrainQueue()) continuation.

| **Sync-blocking / CPU** (`InvokeBlocking`) | Dedicated `LimitedConcurrencyLevelTaskScheduler` | Blocking work holds a real thread for its whole duration. The limited-concurrency scheduler runs at most *cap* leaves at a time, **on threads it starts itself** (`mw-io-lane`; the CPU lane's are `mw-cpu-lane`), never ThreadPool workers. It used to BORROW pool workers, and a cap never protected the pool: with caps up to 256 over a pool whose minimum is the core count, a burst of blocking reads held every worker the grain turns need — see [Blocking Leaves Off the ThreadPool](../BlockingLeavesOffTheThreadPool). |
| **Sync-blocking / CPU** (`InvokeBlocking`) | Dedicated `LimitedConcurrencyLevelTaskScheduler` | Blocking work holds a real thread for its whole duration. The limited-concurrency scheduler runs at most *cap* leaves at a time, **on threads it starts itself** (`mw-io-lane`; the CPU lane's are `mw-cpu-lane`), never ThreadPool workers, and KEEPS them parked between bursts until the pool is disposed — a thread exit per burst was the trigger of a CoreCLR defect, see [Collectible Thread-Static Handle Reuse](../CollectibleThreadStaticHandleReuse). It used to BORROW pool workers, and a cap never protected the pool: with caps up to 256 over a pool whose minimum is the core count, a burst of blocking reads held every worker the grain turns need — see [Blocking Leaves Off the ThreadPool](../BlockingLeavesOffTheThreadPool). |

> The two shapes use threads differently, on purpose. **Async leaves reuse the ThreadPool the framework already uses**, with a governor in front of it — they hold a worker only between awaits, which is what the pool is built for. **Blocking leaves do not touch the ThreadPool at all**: they run on short-lived threads the pool's scheduler starts (one per concurrent drain loop, exiting when its queue is empty), because a blocked thread is exactly what Orleans cannot coordinate with when it is one of its own workers. The thread COUNT is the same either way — a burst of N blocking leaves needs N threads — but a borrowed worker is taken from the grain turns and replaced only by the pool's slow injection, while a started thread is taken from no one.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 23fd328: the following paragraph now says blocking leaves run on dedicated threads kept parked between bursts and exiting only when the pool is disposed.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 1)

1 846 tests  +2   1 844 ✅ +2   4m 50s ⏱️ -8s
    4 suites ±0       2 💤 ±0 
    4 files   ±0       0 ❌ ±0 

Results for commit 23fd328. ± Comparison against base commit d1a2a9d.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 2)

    5 files  ±0      5 suites  ±0   3m 56s ⏱️ -18s
1 230 tests ±0  1 230 ✅ ±0  0 💤 ±0  0 ❌ ±0 
1 231 runs  ±0  1 231 ✅ ±0  0 💤 ±0  0 ❌ ±0 

Results for commit 23fd328. ± Comparison against base commit d1a2a9d.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 0)

  3 files  ±0    3 suites  ±0   7m 38s ⏱️ +8s
579 tests +8  388 ✅ +8  191 💤 ±0  0 ❌ ±0 
583 runs  +8  392 ✅ +8  191 💤 ±0  0 ❌ ±0 

Results for commit 23fd328. ± Comparison against base commit d1a2a9d.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 3)

    4 files  ± 0      4 suites  ±0   7m 25s ⏱️ - 2m 24s
2 654 tests +38  2 654 ✅ +41  0 💤 ±0  0 ❌  - 3 
2 658 runs  +38  2 658 ✅ +41  0 💤 ±0  0 ❌  - 3 

Results for commit 23fd328. ± Comparison against base commit d1a2a9d.

This pull request removes 72 and adds 110 tests. Note that renamed tests count towards both.
Memex.Portal.Shared.Test.InstanceIdRulesMatchTheRegistryTest ‑ TheSetupHostAgreesWithTheRegistry(candidate: "512fb709-b343-484b-a98a-868a01c1bcd6")
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ ALegacyRegistryToken_IsOperatorProvisioning
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ ANamedRegistrysOwnToken_IsOperatorProvisioning
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ ARegistrationKey_IsOperatorProvisioning
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ NeitherKeyNorToken_IsTheOpenLane
Memex.Portal.Shared.Test.TokenValidationIsIssuedOffTheRouterTest ‑ TheCapture_RecordsAViolation_WhenTheRouterGenuinelyPostsWork
Memex.Portal.Shared.Test.TokenValidationIsIssuedOffTheRouterTest ‑ TokenValidationIssuedFromTheRootMeshHub_NeverPutsTheRouterOnEitherEnd
Memex.Portal.Shared.Test.TokenValidationReadinessTest ‑ FewerFailuresThanTheThreshold_StayReady
Memex.Portal.Shared.Test.TokenValidationReadinessTest ‑ GoodSample_IsReady
Memex.Portal.Shared.Test.TokenValidationReadinessTest ‑ NoStore_IsReady
…
Memex.Portal.Shared.Test.InstanceIdRulesMatchTheRegistryTest ‑ TheSetupHostAgreesWithTheRegistry(candidate: "2b367cf1-b275-4321-9bb6-bb627c322789")
Memex.Portal.Shared.Test.TokenMintedOnOneReplicaValidatesOnAnotherTest ‑ AFreshTokenIsValidOnTheReplicaThatDidNotMintIt
Memex.Portal.Shared.Test.TokenMintedOnOneReplicaValidatesOnAnotherTest ‑ AnUnknownTokenIsADefinitiveNegativeOnEitherReplica
Memex.Portal.Shared.Test.TokenValidationHotPathTest ‑ RealToken_ValidatesFromTheStore_AndIsCached
Memex.Portal.Shared.Test.TokenValidationHotPathTest ‑ UnknownToken_IsADefinitiveNotFound_NotUnavailable
Memex.Portal.Shared.Test.TokenValidationNegativeStagesTest ‑ ARowThatExistsButCarriesNoTokenRecord_IsLoggedUnreadable_NotNotFound
Memex.Portal.Shared.Test.TokenValidationNegativeStagesTest ‑ AnAbsentRow_IsLoggedNotFound_NeverUnreadable
Memex.Portal.Shared.Test.TokenVerdictAndCacheTest ‑ Cache_IsBounded
Memex.Portal.Shared.Test.TokenVerdictAndCacheTest ‑ Cache_NeverRemembersFailures
Memex.Portal.Shared.Test.TokenVerdictAndCacheTest ‑ Cache_RemembersSuccesses_UntilTheTtl
…

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 5)

    2 files  ±0      2 suites  ±0   8m 36s ⏱️ -4s
1 122 tests +3  1 122 ✅ +3  0 💤 ±0  0 ❌ ±0 
1 123 runs  +3  1 123 ✅ +3  0 💤 ±0  0 ❌ ±0 

Results for commit 23fd328. ± Comparison against base commit d1a2a9d.

♻️ This comment has been updated with latest results.

…review follow-ups on docs and test

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@meshweaver-cloud
meshweaver-cloud Bot enabled auto-merge October 7, 2026 16:42
@systemorph-com
systemorph-com Bot disabled auto-merge October 7, 2026 16:44
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Test Results (shard 4)

    4 files  ± 0      4 suites  ±0   16m 41s ⏱️ -1s
3 568 tests  - 36  3 568 ✅  - 36  0 💤 ±0  0 ❌ ±0 
3 571 runs   - 36  3 571 ✅  - 36  0 💤 ±0  0 ❌ ±0 

Results for commit 23fd328. ± Comparison against base commit d1a2a9d.

This pull request removes 148 and adds 88 tests. Note that renamed tests count towards both.

   --- End of inner exception stack trace ---
   --- End of inner exception stack trace ---, expected: True)
   --- End of inner exception stack trace ---, isDenial: True)
 ---> (Inner Exception #1) MeshWeaver.Mesh.QueryProviderStalledException: Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.<---
 ---> (Inner Exception #1) MeshWeaver.Messaging.Hub.Test.InfrastructureFaultTest+ProviderException (0x80004005): Failed to connect to 10.42.18.4:5432<---
 ---> (Inner Exception #1) System.ArgumentException: Value does not fall within the expected range.<---
 ---> (Inner Exception #1) System.InvalidOperationException: boom<---
 ---> (Inner Exception #1) System.InvalidOperationException: source B is misconfigured<---
 ---> (Inner Exception #1) System.Net.Sockets.SocketException (0xFFFDFFFF): Name or service not known<---
…
Memex.Portal.Shared.Test.SessionDenialIsAnAnswerTest ‑ OnlyAVerdictReadsAsADenial(shape: "the same verdict nested, as a late denial dispatch"···, failure: System.InvalidOperationException: write failed
 ---> System.UnauthorizedAccessException: Access denied
   --- End of inner exception stack trace ---, isDenial: True)
Memex.Portal.Shared.Test.TieredPgoIsOffInEveryCompilingHostTest ‑ TheRuntimeConfigOnDiskCarriesIt
Memex.Portal.Shared.Test.TieredPgoIsOffInEveryCompilingHostTest ‑ TheRuntimeIsHandedTieredPgoOff
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ ALegacyRegistryToken_IsOperatorProvisioning
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ ANamedRegistrysOwnToken_IsOperatorProvisioning
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ ARegistrationKey_IsOperatorProvisioning
Memex.Portal.Shared.Test.TokenProvisionedInstanceIsKeyedTest ‑ NeitherKeyNorToken_IsTheOpenLane
Memex.Portal.Shared.Test.TokenValidationIsIssuedOffTheRouterTest ‑ TheCapture_RecordsAViolation_WhenTheRouterGenuinelyPostsWork
Memex.Portal.Shared.Test.TokenValidationIsIssuedOffTheRouterTest ‑ TokenValidationIssuedFromTheRootMeshHub_NeverPutsTheRouterOnEitherEnd
Memex.Portal.Shared.Test.TokenValidationReadinessTest ‑ FewerFailuresThanTheThreshold_StayReady
…

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Test Results

    22 files  ± 0      22 suites  ±0   49m 9s ⏱️ - 2m 47s
10 999 tests +15  10 806 ✅ +18  193 💤 ±0  0 ❌  - 3 
11 012 runs  +15  10 819 ✅ +18  193 💤 ±0  0 ❌  - 3 

Results for commit 23fd328. ± Comparison against base commit d1a2a9d.

This pull request removes 40 and adds 31 tests. Note that renamed tests count towards both.

   --- End of inner exception stack trace ---
   --- End of inner exception stack trace ---, expected: True)
   --- End of inner exception stack trace ---, isDenial: True)
 ---> (Inner Exception #1) MeshWeaver.Mesh.QueryProviderStalledException: Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.<---
 ---> (Inner Exception #1) MeshWeaver.Messaging.Hub.Test.InfrastructureFaultTest+ProviderException (0x80004005): Failed to connect to 10.42.18.4:5432<---
 ---> (Inner Exception #1) System.ArgumentException: Value does not fall within the expected range.<---
 ---> (Inner Exception #1) System.InvalidOperationException: boom<---
 ---> (Inner Exception #1) System.InvalidOperationException: source B is misconfigured<---
 ---> (Inner Exception #1) System.Net.Sockets.SocketException (0xFFFDFFFF): Name or service not known<---
…
Memex.Portal.Shared.Test.InstanceIdRulesMatchTheRegistryTest ‑ TheSetupHostAgreesWithTheRegistry(candidate: "2b367cf1-b275-4321-9bb6-bb627c322789")
Memex.Portal.Shared.Test.SessionDenialIsAnAnswerTest ‑ OnlyAVerdictReadsAsADenial(shape: "the same verdict nested, as a late denial dispatch"···, failure: System.InvalidOperationException: write failed
 ---> System.UnauthorizedAccessException: Access denied
   --- End of inner exception stack trace ---, isDenial: True)
Memex.Portal.Shared.Test.TieredPgoIsOffInEveryCompilingHostTest ‑ TheRuntimeConfigOnDiskCarriesIt
Memex.Portal.Shared.Test.TieredPgoIsOffInEveryCompilingHostTest ‑ TheRuntimeIsHandedTieredPgoOff
MeshWeaver.Compiler.Pipeline.Test.SourcesWatcherStallBackoffTest ‑ The_production_classifier_recognises_a_stall_through_every_wrapping(because: "a stall as an aggregate's FIRST member", fault: System.AggregateException: One or more errors occurred. (Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.)
 ---> MeshWeaver.Mesh.QueryProviderStalledException: Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.
   --- End of inner exception stack trace ---, expected: True)
MeshWeaver.Compiler.Pipeline.Test.SourcesWatcherStallBackoffTest ‑ The_production_classifier_recognises_a_stall_through_every_wrapping(because: "a stall as an aggregate's SECOND member", fault: System.AggregateException: One or more errors occurred. (The operation has timed out.) (Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.)
 ---> System.TimeoutException: The operation has timed out.
   --- End of inner exception stack trace ---
 ---> (Inner Exception #1) MeshWeaver.Mesh.QueryProviderStalledException: Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.<---
, expected: True)
MeshWeaver.Compiler.Pipeline.Test.SourcesWatcherStallBackoffTest ‑ The_production_classifier_recognises_a_stall_through_every_wrapping(because: "a stall nested in an aggregate inside a wrapper", fault: System.InvalidOperationException: outer
 ---> System.AggregateException: One or more errors occurred. (Value does not fall within the expected range.) (Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.)
 ---> System.ArgumentException: Value does not fall within the expected range.
   --- End of inner exception stack trace ---
 ---> (Inner Exception #1) MeshWeaver.Mesh.QueryProviderStalledException: Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.<---

   --- End of inner exception stack trace ---, expected: True)
MeshWeaver.Compiler.Pipeline.Test.SourcesWatcherStallBackoffTest ‑ The_production_classifier_recognises_a_stall_through_every_wrapping(because: "a stall wrapped as InnerException", fault: System.InvalidOperationException: outer
 ---> MeshWeaver.Mesh.QueryProviderStalledException: Query provider(s) [pg] did not emit an Initial within the query fan-in's 16s bound for query 'nodeType:NodeType' (user 'system'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.
   --- End of inner exception stack trace ---, expected: True)
MeshWeaver.Compiler.Pipeline.Test.SourcesWatcherStallBackoffTest ‑ The_production_classifier_recognises_a_stall_through_every_wrapping(because: "an aggregate of unrelated faults", fault: System.AggregateException: One or more errors occurred. (The operation has timed out.) (Value does not fall within the expected range.)
 ---> System.TimeoutException: The operation has timed out.
   --- End of inner exception stack trace ---
 ---> (Inner Exception #1) System.ArgumentException: Value does not fall within the expected range.<---
, expected: False)
MeshWeaver.FaultInjection.Test.AChangeFeedGapReachesEveryConsumerTest ‑ OwnNodeMirror_ReReadsTheStore_AfterACommitToItDuringTheGap
…

@rbuergi
rbuergi merged commit f0336ab into main Oct 7, 2026
63 checks passed
rbuergi added a commit that referenced this pull request Oct 8, 2026
docs(sigsegv): sightings #19/#20 — FutuRe dumps in the unload drain's GC.Collect; #6259/#6254 not shown to explain them
@rbuergi
rbuergi deleted the fix/4654-tls-handle-reuse branch October 10, 2026 13:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants