2026-09-19 09:42:47Z memex-portal-deployment-69956b6dbc-s246c crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
[ROUTE] Routing back-pressure [38381b86#19 started 2026-09-19T09:42:47.4863252Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 0, deepest per-destination queue 0, routing pool subscribing 0. Latest dispatch target Collaboration/_Issue/2026 — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:31Z memex-portal-deployment-69956b6dbc-s246c crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
[ROUTE] Routing back-pressure [38381b86#20 started 2026-09-19T09:44:31.8935946Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 0, deepest per-destination queue 0, routing pool subscribing 0. Latest dispatch target Store/Core — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:31Z memex-portal-deployment-69956b6dbc-ndnxt crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
[ROUTE] Routing back-pressure [cbc2edb8#4 started 2026-09-19T09:44:31.9065662Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 1, deepest per-destination queue 62, routing pool subscribing 0. Latest dispatch target cache/YORFhwiyqEyTP9ESF-NUNg — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:32Z memex-portal-deployment-69956b6dbc-s246c crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
[ROUTE] Routing back-pressure [38381b86#21 started 2026-09-19T09:44:32.1153739Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 0, deepest per-destination queue 0, routing pool subscribing 1. Latest dispatch target Reinsurance/ExposureProfile — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:32Z memex-portal-deployment-69956b6dbc-gx6z6 crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
[ROUTE] Routing back-pressure [d50bcf55#13 started 2026-09-19T09:44:32.1252680Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 1, deepest per-destination queue 62, routing pool subscribing 0. Latest dispatch target cache/YORFhwiyqEyTP9ESF-NUNg — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
What is failing
RoutingGrain(MeshWeaver.Hosting.Orleans) repeatedly hit its 64-dispatch back-pressure threshold on thememex-portaldeployment, across all four pods within a four-minute window. This is one of dozens of identical RoutingGrain back-pressure incidents opened across the fleet today, so treat it as a systemic routing saturation event, not a single stuck grain.Probable cause
The samples split into two distinct shapes:
s246c, activation38381b86):stream destinations queued 0, deepest per-destination queue 0. Nothing is waiting on anything; all 64 slots are simply held by in-flight legs. Per the log's own episode semantics, episodes Update release-packages.yml #19→Update release-packages.yml #21 on the same activation each drained, so no slot leak. This points at general dispatch volume or ThreadPool starvation (slots are held through the unbounded wait for a ThreadPool thread before leg timeouts even start).ndnxtandgx6z6):stream destinations queued 1, deepest per-destination queue 62, latest dispatch targetcache/YORFhwiyqEyTP9ESF-NUNgon both pods. One cache destination leg appears blocked and ~62 legs are queued behind it — replicated across two pods, which suggests the destination's own leg (e.g. a hub activation or subscription) is stalled rather than a pod-local hiccup.Confidence: moderate. Shape 1 is well-supported by the episode stamps and zero queue depths; shape 2 rests on two samples naming the same target, and a queue depth sampled at the crossing is partly a threshold artefact.
Impact
14 critical occurrences in ~4 minutes (09:40–09:44 UTC) on 4/4 pods, plus dozens of sibling incidents fleet-wide over the following hours. Routing throughput on the affected namespace (
memex-cloud) is degraded or saturating; any client relying on route dispatch during these windows sees increased latency. If thecache/YORFhwiyqEyTP9ESF-NUNgbacklog persists without a higher episode following on its activations, the blocked leg's slots leak and routing capacity permanently shrinks on those grains.Where to look
MeshWeaver.Hosting.Orleans.RoutingGrain— the dispatch slot accounting, in particular the slot-release path and the unbounded ThreadPool wait before leg timeouts start (shape 1), and the per-destination stream queue feedingcache/YORFhwiyqEyTP9ESF-NUNg(shape 2).Doc/Architecture/OrleansTaskSchedulerandDoc/WhatsNew/2026-08-10-router-stays-a-router(background services were moved off the mesh router — check nothing regressed there).Duplicates: this log site has produced many near-identical incidents today; if an umbrella issue already exists for the routing saturation event, close this in favour of it and link the samples there.
Evidence
e2be47a4a07c0f8bMeshWeaver.Hosting.Orleans.RoutingGrainmemex-cloudmemex-portal-deployment-69956b6dbc-g6bbb,memex-portal-deployment-69956b6dbc-gx6z6,memex-portal-deployment-69956b6dbc-s246c,memex-portal-deployment-69956b6dbc-ndnxtRecent log lines
Re-addressed by the current identity function: this incident inherited
e4a97855ab595beb(Systemorph/MeshWeaver.Plugins#1796). Those nodes are superseded and will not fold, file or comment again.Opened automatically from
Admin/_LogIncident/e2be47a4a07c0f8b. Recurrences are folded into this issue rather than opening new ones.