Skip to content

RoutingGrain dispatch saturation across memex-portal: ThreadPool-bound load plus head-of-line blocking on a cache destination leg #4893

Description

@systemorph-com

What is failing

RoutingGrain (MeshWeaver.Hosting.Orleans) repeatedly hit its 64-dispatch back-pressure threshold on the memex-portal deployment, across all four pods within a four-minute window. This is one of dozens of identical RoutingGrain back-pressure incidents opened across the fleet today, so treat it as a systemic routing saturation event, not a single stuck grain.

Probable cause

The samples split into two distinct shapes:

  1. Pure load (majority, e.g. pod s246c, activation 38381b86): stream destinations queued 0, deepest per-destination queue 0. Nothing is waiting on anything; all 64 slots are simply held by in-flight legs. Per the log's own episode semantics, episodes Update release-packages.yml #19→Update release-packages.yml #21 on the same activation each drained, so no slot leak. This points at general dispatch volume or ThreadPool starvation (slots are held through the unbounded wait for a ThreadPool thread before leg timeouts even start).
  2. Head-of-line blocking on one destination (pods ndnxt and gx6z6): stream destinations queued 1, deepest per-destination queue 62, latest dispatch target cache/YORFhwiyqEyTP9ESF-NUNg on both pods. One cache destination leg appears blocked and ~62 legs are queued behind it — replicated across two pods, which suggests the destination's own leg (e.g. a hub activation or subscription) is stalled rather than a pod-local hiccup.

Confidence: moderate. Shape 1 is well-supported by the episode stamps and zero queue depths; shape 2 rests on two samples naming the same target, and a queue depth sampled at the crossing is partly a threshold artefact.

Impact

14 critical occurrences in ~4 minutes (09:40–09:44 UTC) on 4/4 pods, plus dozens of sibling incidents fleet-wide over the following hours. Routing throughput on the affected namespace (memex-cloud) is degraded or saturating; any client relying on route dispatch during these windows sees increased latency. If the cache/YORFhwiyqEyTP9ESF-NUNg backlog persists without a higher episode following on its activations, the blocked leg's slots leak and routing capacity permanently shrinks on those grains.

Where to look

  • MeshWeaver.Hosting.Orleans.RoutingGrain — the dispatch slot accounting, in particular the slot-release path and the unbounded ThreadPool wait before leg timeouts start (shape 1), and the per-destination stream queue feeding cache/YORFhwiyqEyTP9ESF-NUNg (shape 2).
  • Determine whether the two shapes share a root cause (e.g. ThreadPool starvation cascading into the cache destination's leg) before fixing one in isolation.
  • Related background: Doc/Architecture/OrleansTaskScheduler and Doc/WhatsNew/2026-08-10-router-stays-a-router (background services were moved off the mesh router — check nothing regressed there).

Duplicates: this log site has produced many near-identical incidents today; if an umbrella issue already exists for the routing saturation event, close this in favour of it and link the samples there.


Evidence

Fingerprint e2be47a4a07c0f8b
Category MeshWeaver.Hosting.Orleans.RoutingGrain
Severity Critical
Namespace memex-cloud
Pods memex-portal-deployment-69956b6dbc-g6bbb, memex-portal-deployment-69956b6dbc-gx6z6, memex-portal-deployment-69956b6dbc-s246c, memex-portal-deployment-69956b6dbc-ndnxt
Occurrences 14
First seen 2026-09-19 09:40:35Z
Last seen 2026-09-19 09:44:32Z
Recent log lines
2026-09-19 09:42:47Z memex-portal-deployment-69956b6dbc-s246c crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
      [ROUTE] Routing back-pressure [38381b86#19 started 2026-09-19T09:42:47.4863252Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 0, deepest per-destination queue 0, routing pool subscribing 0. Latest dispatch target Collaboration/_Issue/2026 — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:31Z memex-portal-deployment-69956b6dbc-s246c crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
      [ROUTE] Routing back-pressure [38381b86#20 started 2026-09-19T09:44:31.8935946Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 0, deepest per-destination queue 0, routing pool subscribing 0. Latest dispatch target Store/Core — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:31Z memex-portal-deployment-69956b6dbc-ndnxt crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
      [ROUTE] Routing back-pressure [cbc2edb8#4 started 2026-09-19T09:44:31.9065662Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 1, deepest per-destination queue 62, routing pool subscribing 0. Latest dispatch target cache/YORFhwiyqEyTP9ESF-NUNg — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:32Z memex-portal-deployment-69956b6dbc-s246c crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
      [ROUTE] Routing back-pressure [38381b86#21 started 2026-09-19T09:44:32.1153739Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 0, deepest per-destination queue 0, routing pool subscribing 1. Latest dispatch target Reinsurance/ExposureProfile — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.
2026-09-19 09:44:32Z memex-portal-deployment-69956b6dbc-gx6z6 crit: MeshWeaver.Hosting.Orleans.RoutingGrain[0]
      [ROUTE] Routing back-pressure [d50bcf55#13 started 2026-09-19T09:44:32.1252680Z]: 64 route dispatches in flight (reporting threshold 64); stream destinations queued 1, deepest per-destination queue 62, routing pool subscribing 0. Latest dispatch target cache/YORFhwiyqEyTP9ESF-NUNg — the address that happened to cross the threshold, NOT a diagnosis. A slot is held from dispatch until the leg terminates, INCLUDING the unbounded wait for a ThreadPool thread before the leg's own timeouts start, so a CPU-starved silo raises this with nothing stuck. A deepest queue of 1 or more means legs are blocked behind a leg (head-of-line on one destination); 0 means nothing is waiting on anything, so read it as load. 🚨 Deepest is sampled AT THE CROSSING, so like the in-flight count it is partly an artefact of the threshold: with N destinations sharing the backlog it is ~InFlight/N whatever is wrong. READ THE EPISODE STAMP, not the depth: a later line with a HIGHER episode on this activation means this episode drained; a line with a DIFFERENT activation id means the grain was recycled; and if neither a clear nor a higher episode ever follows, the in-flight count never fell below half the threshold — which means a leg never terminated and its slot leaked, not that the silo was busy.

Re-addressed by the current identity function: this incident inherited e4a97855ab595beb (Systemorph/MeshWeaver.Plugins#1796). Those nodes are superseded and will not fold, file or comment again.

Opened automatically from Admin/_LogIncident/e2be47a4a07c0f8b. Recurrences are folded into this issue rather than opening new ones.

Activity

  1. rbuergi commented on Sep 19, 2026

    @rbuergi
    Contributor

    Closing as a duplicate of #4798.

    This issue and 46 others were filed today by one log site — MeshWeaver.Hosting.Orleans.RoutingGrain's back-pressure line — because the Orleans activation id in that line was not masked before the fingerprint was taken. LogLineParser.HexBlob required sixteen hex characters; an activation id is eight, so it fell through to the number masker, which stripped its digits and left its letters behind as literal text. An activation id changes on every grain recycle and every pod roll, and memex-cloud rolled repeatedly today, so each activation minted a new fingerprint, a new incident and a new issue.

    These are not 47 defects. Under correct masking all 47 normalize to one message and would have been one incident. Verified two of them against the control instance's own records: Admin/_LogIncident/5b6a59648c91fd85 and eb29857024b3f0ed differ only in cfb5d6 vs c0cc8 — the dispatch target is already {path} in both, and every count is already {n}.

    Fixed in Systemorph/MeshWeaver.Plugins#2165 (merged), which masks a run of ≥ 8 hex characters carrying both a digit and an a–f letter. The fault itself — routing back-pressure at the 64-dispatch threshold — remains open and is tracked on #4798, which carries the longest occurrence history.

    Reopen if you believe this one is a distinct fault rather than another activation of the same line.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions