Skip to content

Add Durability.MessagingEnabled for event-subscription-only nodes (GH-3746) - #3747

Closed
erdtsieck wants to merge 1 commit into
JasperFx:mainfrom
erdtsieck:gh-3746/messaging-disabled-node
Closed

erdtsieck wants to merge 1 commit into
JasperFx:mainfrom
erdtsieck:gh-3746/messaging-disabled-node

Conversation

@erdtsieck

Copy link
Copy Markdown
Contributor

Implements the ask in #3746. Happy to reshape any of it — the naming, the surface, or the placement — this is a first cut to make the discussion concrete rather than a take-it-or-leave-it.

The problem

There is no supported way to run a node that takes part in event subscription (projection) agent assignment while doing no message handling at all.

The use case is warming a read model before it serves traffic. When a projection version is bumped, its tables start empty and are rebuilt from the beginning of the event store; until that finishes, a node running the new version serves incomplete read models. Standing up a separate set of nodes to build the new version off to one side is the natural fix — and it works, because the projection version is part of the agent identity, so old and new nodes advertise disjoint projection agents.

But those warming nodes must be full Wolverine nodes to be assigned the new version's projection agents, and that also hands them message handling. Listeners we can stop via IEndpointCollection.StartListenerAsync/StopListenerAsync. Durability agents we cannot: they are assigned per message store, are version-independent, and a durability agent is not a listener, so it recovers and executes persisted envelopes on a node whose read models are half-built. NodeAgentController.DisableAgentsAsync is documented "STRICTLY FOR TESTING", and Solo/Serverless/MediatorOnly are not usable — Balanced is what enables the leader election and control queue the distribution needs.

For scale context: our deployment is 512 shard databases and ~854 tenants, so about 5 100 projection agents plus 512 durability agents.

What this adds

DurabilitySettings.MessagingEnabled (default true), plus a fluent services.RunWolverineForEventSubscriptionDistributionOnly() alongside the existing DisableAllWolverineMessagePersistence / UseWolverineSoloMode.

When false, the node:

  • registers no listener agent families, so the leader cannot pin an exclusive or leader-pinned listener there;
  • registers no durability agent family;
  • registers no transport-owned agent families;
  • starts no listeners;
  • runs no in-memory scheduled jobs.

It stays a full cluster member otherwise — registers, takes part in leader election, uses the control queue, and still advertises and runs the event subscription agents it was stood up for.

The gating follows the pattern DurabilityAgentEnabled already uses in the NodeAgentController constructor, so the mechanism should look familiar.

One thing worth your eye

A node that registers no durability family also cannot assign those agents while it is the leader, because EvaluateAssignmentsAsync iterates the registered families. That is pre-existing behaviour of DurabilityAgentEnabled which this flag inherits rather than introduces — but it does suggest such a node should not be leader-eligible.

I deliberately did not touch leader election in the same PR. If you would rather this flag also made the node ineligible, or rather the families stayed registered with SupportedAgentsAsync() returning empty (so a leader could still assign them to others while never taking them itself), say which and I will rework it. The second shape looks cleaner to me but is a bigger behavioural change and I would rather not guess.

Testing

Three unit tests in CoreTests.Runtime.Agents.messaging_disabled_node covering the listener families being absent when off, present by default, and the injected event subscription family still being registered when off. The neighbouring agent tests still pass.

Default is true, so nothing changes for existing hosts.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds an opt-out switch (DurabilitySettings.MessagingEnabled) to allow running a Wolverine node that participates in event-subscription/projection agent distribution while intentionally avoiding normal messaging responsibilities (listeners, durability agents, transport-owned agents, in-memory scheduled jobs). This fits into Wolverine’s clustered agent-assignment/runtime startup pipeline for DurabilityMode.Balanced.

Changes:

  • Introduces DurabilitySettings.MessagingEnabled (default true) and includes it in option descriptions.
  • Gates startup behaviors (in-memory scheduled jobs, endpoint listener startup) and agent-family registration to support “projection-only” nodes.
  • Adds a RunWolverineForEventSubscriptionDistributionOnly() IServiceCollection helper plus unit tests for the agent-family wiring.

Reviewed changes

Copilot reviewed 5 out of 5 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
src/Wolverine/Runtime/WolverineRuntime.HostService.cs Gates in-memory scheduled jobs; skips starting endpoint listeners when MessagingEnabled is false; adds informational logging.
src/Wolverine/Runtime/Agents/NodeAgentController.cs Suppresses registration of listener/durability/transport agent families when MessagingEnabled is false; adds HasFamily test hook.
src/Wolverine/HostBuilderExtensions.cs Adds fluent helper to configure MessagingEnabled = false via an internal extension.
src/Wolverine/DurabilitySettings.cs Adds the MessagingEnabled setting + XML docs and includes it in ToDescription().
src/Testing/CoreTests/Runtime/Agents/messaging_disabled_node.cs Unit tests asserting which agent families are (not) registered when messaging is disabled.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/Wolverine/DurabilitySettings.cs Outdated
Comment on lines +137 to +141
/// <summary>
/// Should the message durability agent be enabled during execution.
/// The default is true.
/// </summary>
public bool DurabilityAgentEnabled { get; set; } = true;
Comment on lines +547 to +551
// This node exists to run event subscription agents only. See
// DurabilitySettings.MessagingEnabled.
Logger.LogInformation(
"All endpoint listeners are disabled because Durability.MessagingEnabled is false");
}
Comment on lines +86 to 90
// A node with messaging disabled runs event subscription agents and nothing else, so it
// advertises no listener agents. See DurabilitySettings.MessagingEnabled.
if (runtime.Options.Durability.Mode == DurabilityMode.Balanced && runtime.Options.Durability.MessagingEnabled)
{
_agentFamilies[ExclusiveListenerFamily.SchemeName] = new ExclusiveListenerFamily(runtime);
@erdtsieck
erdtsieck force-pushed the gh-3746/messaging-disabled-node branch from 20217eb to 2db8707 Compare July 31, 2026 13:19
@erdtsieck

Copy link
Copy Markdown
Contributor Author

Pushed two follow-ups after a review question about scheduled messaging.

Scheduled messaging was only half off. I had gated startInMemoryScheduledJobs() but not startDurableScheduledJobs(), so a node with messaging disabled would still have processed durable scheduled messages — exactly the thing it must not do. Both are gated now.

ScheduleLocalExecutionInMemory already guarded against a null scheduler, but attributed it to Durability.Mode; it now names MessagingEnabled when that is the real cause, so the exception points at the setting that actually caused it.

One deliberate exception, worth stating explicitly: sending agents stay enabled. Wolverine's own control queue rides on them, so a node that could not send could not take part in agent assignment at all — which would defeat the purpose. So the boundary this PR draws is inbound and scheduled work off, outbound kept, rather than literally everything off. If you would rather see that expressed differently, I am happy to change it.

Also force-pushed a line-ending fix: my editor had rewritten four files to CRLF, which buried ~500 lines of real churn in whitespace. The diff is 169/9 now and every line in it is a real change.

Closes the gap described in JasperFxGH-3746. There is currently no supported way to
run a node that takes part in event subscription (projection) agent
assignment while doing no message handling at all.

The use case is warming a read model before it serves traffic. When a
projection version is bumped its tables start empty and are rebuilt from the
beginning of the event store; until that finishes, a node running the new
version serves incomplete read models. Standing up a separate set of nodes to
build the new version off to one side solves that - but only if those nodes
do not also pick up messages and run handlers against half-built state.

DurabilitySettings.MessagingEnabled = false (or the fluent
RunWolverineForEventSubscriptionDistributionOnly) makes a node:

- register no listener agent families, so the leader cannot pin an exclusive
  or leader-pinned listener there;
- register no durability agent family;
- register no transport-owned agent families;
- start no listeners;
- start neither the in-memory nor the durable scheduled job processor.

Sending agents deliberately stay enabled: Wolverine's own control queue rides
on them, so a node that could not send could not take part in agent
assignment at all.

The node stays a full cluster member otherwise: it registers, takes part in
leader election, uses the control queue, and still advertises and runs the
event subscription agents it was stood up for.

ScheduleLocalExecutionInMemory already guarded against a null scheduler but
attributed it to Durability.Mode; it now names MessagingEnabled when that is
the actual cause.

Default is true, so nothing changes for existing hosts.

The gating follows the pattern DurabilityAgentEnabled already uses in the
NodeAgentController constructor. Worth a maintainer's eye: a node that
registers no durability family also cannot *assign* those agents while it is
the leader, since EvaluateAssignmentsAsync iterates the registered families.
That is pre-existing behaviour of DurabilityAgentEnabled which this flag
inherits rather than introduces, but it means such a node arguably should not
be leader-eligible. Left alone here rather than changing leader election in
the same PR.
@erdtsieck
erdtsieck force-pushed the gh-3746/messaging-disabled-node branch from 2db8707 to 9e0cc94 Compare July 31, 2026 13:34
@erdtsieck

Copy link
Copy Markdown
Contributor Author

Thanks — all three land. Two are fixed, the third I would like your call on.

Stray carriage returns. Correct, and worse than it looked: my editor had rewritten several files to CRLF, which was burying ~500 lines of whitespace churn. Normalised every touched file back to the repo's LF and force-pushed. The diff is 169/9 now, all real changes, 0 stray CRs.

Log wording. Also correct — StartListenersAsync only covers external endpoints, and local queues keep running (which they must; agent commands ride on them). Reworded to say so explicitly rather than claiming all listeners are off.

The leader problem. You reached the same conclusion I flagged in the description, independently, which I take as a strong signal it should not ship as a documented caveat. Of your two options I also prefer the second — keep the families registered so a leader can still evaluate and reassign them, and suppress only this node's advertised capabilities.

Where I got stuck, and why I would rather ask than guess: MessageStoreCollection is a plain IAgentFamily, not IStaticAgentFamily, so its agents are placed by the leader through EvaluateAssignmentsAsync rather than from the persisted node capabilities that SupportedAgentsAsync feeds. Returning empty from SupportedAgentsAsync therefore fixes the static listener families cleanly, but I could not convince myself it keeps the durability agents off a messaging-disabled node — that seems to depend on whether the AssignmentGrid filters candidate nodes by capability for non-static families too.

If it does, a small decorator that delegates everything except SupportedAgentsAsync (empty) and BuildAgentAsync (throws) covers all three schemes and I will push it. If it does not, the durability family needs something else and I would rather you point me at the right seam than have me invent one in assignment code I have only just read.

I have not run this on a real cluster — the tests cover the wiring, not the assignment behaviour under leadership change.

@erdtsieck

Copy link
Copy Markdown
Contributor Author

The one red check, CIKafka, does not look related to this change:

  • it also fails on the current main HEAD (d49a1f5b), while passing on the three commits before it, so it appeared on main independently of this branch;
  • the failure itself is a DLQ assertion (Shouldly.ShouldAssertException : dlq) preceded by a long run of Database is not ready (Exception while reading from stream) — the container not coming up rather than anything MessagingEnabled touches.

Everything else is green (build, test, and the other 29 broker/persistence matrices).

Context for why I care about this one landing: we spent today failing to get a release with a projection version bump through our canary, and the root of it is that a bumped shard takes minutes to start rather than milliseconds (JasperFx/jasperfx#594), which then breaks three separate things in the distribution layer (#3748, #3749, #3750). An event-subscription-only node is what would let us warm a new projection version on pods that serve no traffic and hand the caught-up state to the serving fleet — i.e. it takes the replay out of the deploy path entirely instead of trying to make the deploy survive it. Happy to test whatever shape you prefer against our 512-database, ~6,500-agent cluster.

@jeremydmiller

Copy link
Copy Markdown
Member

Answering the direct question first, because it has a definite answer and it isn't the one you were hoping for.

No — the grid does not filter candidate nodes by capability on the paths these three families actually use. So the decorator shape doesn't work, and for durability it can't be made to work by that route at all. Here's the whole picture, since you were right to not want to guess at it.

Capabilities only ever come from static families

NodeAgentController.StartLocalProcessing.cs:10-13:

foreach (var controller in _agentFamilies.Values.OfType<IStaticAgentFamily>())
{
    current.Capabilities.AddRange(await controller.SupportedAgentsAsync());
}

You'd spotted that MessageStoreCollection is a plain IAgentFamily (MessageStoreCollection.cs:13). The consequence is stronger than "its agents are placed differently": durability agent Uris are never in any node's Capabilities, on any node, ever. There is no capability for a decorator to suppress.

Capability-aware distribution exists — but not on these paths

The grid does have capability matching, in MatchAgentsToCapableNodesFor (AssignmentGrid.cs:137-146), and two distribution methods consult it — but only when nodes actually differ:

  • DistributeEvenlyWithBlueGreenSemantics (AssignmentGrid.Distribution.cs:330-338)
  • DistributeByGroupAffinity (AssignmentGrid.Distribution.cs:140-145)

Those are the event-subscription paths. All three messaging families take the capability-blind ones:

family placement capability-aware?
MessageStoreCollection (durability, registered at NodeAgentController.cs:103) DistributeEvenly(Scheme) — MessageStoreCollection.cs:318-321 no
ExclusiveListenerFamily DistributeEvenly(SchemeName) — ExclusiveListenerFamily.cs:117-120 no
LeaderPinnedListenerFamily RunOnLeader(uri) — LeaderPinnedAgentFamily.cs:117-121 no

DistributeEvenly walks _nodes with no capability check anywhere in it (AssignmentGrid.Distribution.cs:25-100), Node.Assign has none, and RunOnLeader is _nodes.FirstOrDefault(x => x.IsLeader)?.Assign(agentUri) (AssignmentGrid.cs:244-248) with no condition at all.

So an empty SupportedAgentsAsync does drop the listener families out of Capabilities — and then changes nothing, because ExclusiveListenerFamily never reads capabilities when it distributes. Your decorator's BuildAgentAsync throw would convert a silent misassignment into a loud one, which beats the status quo, but you'd have a node permanently assigned agents it refuses to build.

And switching MessageStoreCollection to one of the capability-aware methods would be worse, not better: with the durability Uri in nobody's capability list, every node would be a non-candidate and durability would stop being assigned cluster-wide.

The seam

Both halves are needed, and only the second is new work:

  1. Keep the families registered. You'd already concluded this and it's right — EvaluateAssignments builds the grid by iterating the leader's own _agentFamilies (NodeAgentController.EvaluateAssignments.cs:79-96), so a leader missing the durability family stops assigning durability agents for everybody, not just itself. This is the pre-existing DurabilityAgentEnabled hazard you flagged in the description, and it's real.

  2. Exclude the node inside distribution, not via capabilities. MessagingEnabled should travel node→leader as a first-class attribute on WolverineNode and AssignmentGrid.Node — it isn't an agent Uri and doesn't belong in the capability list, which is also the blue/green matching set. Then DistributeEvenly skips excluded nodes.

The exclusion has to be per scheme, not global — the entire point of these nodes is that event-subscription agents still land on them. DistributeEvenly(scheme, filter) already exists, so the natural shape is for the grid to answer "which nodes are eligible for this scheme" rather than each family re-deriving it.

On the leader question

Pinned agents fall back; the node stays leader-eligible. Leader-ineligibility deadlocks a cluster that is entirely warming nodes, and keeping that shape working is the point.

So RunOnLeader pins to the leader when the leader can host messaging, and otherwise to the lowest-numbered messaging-enabled node. That preserves the single-instance guarantee and weakens "leader-pinned" to "prefer the leader" — the weaker promise, but the one that survives your deployment.

Push that if you'd like; happy to review. Two things worth knowing before you do:

  • Nothing in the suite would catch this today. Your three unit tests cover the wiring, and the assignment behaviour under leadership change is precisely what they don't reach. A test in CoreTests.Runtime.Agents.assigning_agent_logic that builds a grid containing a messaging-disabled node and asserts the durability agent lands elsewhere is the cheapest real proof, and it should fail before your change.
  • Ignore MessageDatabase.Agents.cs:42-46. MessageDatabase<T> also implements IAgentFamily with its own RunOnLeader(_defaultAgent), but it is not the registered family — _runtime.Stores is — so it isn't on this path. Mentioning it only so you don't chase it.

@jeremydmiller

Copy link
Copy Markdown
Member

Following up on my own comment above, because I've changed my mind about the shape rather than the analysis — please don't start on that RunOnLeader work.

I'm taking #3746 out of this release. Everything I wrote about the assignment internals still holds, and that's precisely the problem: making this work correctly means a new node-level attribute travelling node→leader, per-scheme eligibility inside DistributeEvenly, and a redefinition of what "leader-pinned" promises. That's a structural change to the distribution layer, landing under deadline pressure, to serve a deployment shape we cannot reproduce outside your cluster. That's the wrong trade right now.

What I'd do instead, and it should work today

Run the warming fleet as its own Wolverine cluster: a different ServiceName, and — the part that actually does the work — a separate Wolverine durability schema.

opts.ServiceName = "trips-projection-warmer";

opts.PersistMessagesWithPostgresql(connectionString, schemaName: "warmer_wolverine");
// or, on the Marten/Polecat integration:
//   .IntegrateWithWolverine(o =>
//   {
//       o.MessageStorageSchemaName = "warmer_wolverine";
//       o.TransportSchemaName      = "warmer_wolverine";
//   });

The durability schema is the cluster boundary. wolverine_nodes, wolverine_node_assignments, and wolverine_agent_restrictions all live in it, so two fleets on two schemas are two independent clusters: separate leader elections, separate assignment grids, and no node in one is ever a candidate in the other's DistributeEvenly.

That gets you what MessagingEnabled was reaching for, from the other direction. The warming nodes still run a durability agent — but only over their own schema, which no application traffic is ever routed to, so it has nothing to recover and nothing to execute. You don't need to suppress the agent; you need to make sure it's pointed somewhere harmless, and a dedicated schema does that.

Your projection agents are unaffected, because they're distributed per-cluster off the projections each host registers. The warming fleet registers V2, the serving fleet registers V1, and — as you noted in the description — the version is part of the agent identity, so the two sets are disjoint by construction. That disjointness is what makes the isolation safe, and it's worth checking deliberately rather than assuming: two independent clusters registering the same projection version against one event store would both try to assign it, which is the failure mode this arrangement has instead of the one it fixes.

Practical notes: point listeners/queues at the warming fleet only if you mean to, since anything delivered there stays there; and the new schema needs creating, which the usual resource application will do.

This is the official guidance for now

I'll document it as the supported way to warm a projection version off to one side. If it turns out to have a sharp edge in practice on your cluster, that's a much better-grounded case for a framework-level flag than the one we have today, and I'd rather add MessagingEnabled knowing which specific thing it has to fix.

What isn't dropped

The bugs you found underneath this are the real prize and none of them are affected:

Thank you for the PR, genuinely — the description and the follow-ups were unusually careful, and finding the leader-cannot-assign-what-it-doesn't-register hazard was worth the exercise on its own. I'll leave it open rather than close it out from under you; it just isn't tracking for this release. Once #598 lands and we can see what's actually left of the problem, it's worth picking back up.

What I'd still like most is the pinned-prerelease-plus-measurement arrangement we discussed. Nothing in this chain reproduces at dev scale, and your 512-database canary is the only real feedback loop any of it has.

@jeremydmiller

Copy link
Copy Markdown
Member

We're going to try to solve the real issue a different way. This is generating way too much complexity that will likely be a problem later.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants