Conversation
1e0ddd3 to
20217eb
Compare
There was a problem hiding this comment.
Pull request overview
Adds an opt-out switch (DurabilitySettings.MessagingEnabled) to allow running a Wolverine node that participates in event-subscription/projection agent distribution while intentionally avoiding normal messaging responsibilities (listeners, durability agents, transport-owned agents, in-memory scheduled jobs). This fits into Wolverine’s clustered agent-assignment/runtime startup pipeline for DurabilityMode.Balanced.
Changes:
- Introduces
DurabilitySettings.MessagingEnabled(defaulttrue) and includes it in option descriptions. - Gates startup behaviors (in-memory scheduled jobs, endpoint listener startup) and agent-family registration to support “projection-only” nodes.
- Adds a
RunWolverineForEventSubscriptionDistributionOnly()IServiceCollection helper plus unit tests for the agent-family wiring.
Reviewed changes
Copilot reviewed 5 out of 5 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| src/Wolverine/Runtime/WolverineRuntime.HostService.cs | Gates in-memory scheduled jobs; skips starting endpoint listeners when MessagingEnabled is false; adds informational logging. |
| src/Wolverine/Runtime/Agents/NodeAgentController.cs | Suppresses registration of listener/durability/transport agent families when MessagingEnabled is false; adds HasFamily test hook. |
| src/Wolverine/HostBuilderExtensions.cs | Adds fluent helper to configure MessagingEnabled = false via an internal extension. |
| src/Wolverine/DurabilitySettings.cs | Adds the MessagingEnabled setting + XML docs and includes it in ToDescription(). |
| src/Testing/CoreTests/Runtime/Agents/messaging_disabled_node.cs | Unit tests asserting which agent families are (not) registered when messaging is disabled. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| /// <summary> | ||
| /// Should the message durability agent be enabled during execution. | ||
| /// The default is true. | ||
| /// </summary> | ||
| public bool DurabilityAgentEnabled { get; set; } = true; |
| // This node exists to run event subscription agents only. See | ||
| // DurabilitySettings.MessagingEnabled. | ||
| Logger.LogInformation( | ||
| "All endpoint listeners are disabled because Durability.MessagingEnabled is false"); | ||
| } |
| // A node with messaging disabled runs event subscription agents and nothing else, so it | ||
| // advertises no listener agents. See DurabilitySettings.MessagingEnabled. | ||
| if (runtime.Options.Durability.Mode == DurabilityMode.Balanced && runtime.Options.Durability.MessagingEnabled) | ||
| { | ||
| _agentFamilies[ExclusiveListenerFamily.SchemeName] = new ExclusiveListenerFamily(runtime); |
20217eb to
2db8707
Compare
|
Pushed two follow-ups after a review question about scheduled messaging. Scheduled messaging was only half off. I had gated
One deliberate exception, worth stating explicitly: sending agents stay enabled. Wolverine's own control queue rides on them, so a node that could not send could not take part in agent assignment at all — which would defeat the purpose. So the boundary this PR draws is inbound and scheduled work off, outbound kept, rather than literally everything off. If you would rather see that expressed differently, I am happy to change it. Also force-pushed a line-ending fix: my editor had rewritten four files to CRLF, which buried ~500 lines of real churn in whitespace. The diff is 169/9 now and every line in it is a real change. |
Closes the gap described in JasperFxGH-3746. There is currently no supported way to run a node that takes part in event subscription (projection) agent assignment while doing no message handling at all. The use case is warming a read model before it serves traffic. When a projection version is bumped its tables start empty and are rebuilt from the beginning of the event store; until that finishes, a node running the new version serves incomplete read models. Standing up a separate set of nodes to build the new version off to one side solves that - but only if those nodes do not also pick up messages and run handlers against half-built state. DurabilitySettings.MessagingEnabled = false (or the fluent RunWolverineForEventSubscriptionDistributionOnly) makes a node: - register no listener agent families, so the leader cannot pin an exclusive or leader-pinned listener there; - register no durability agent family; - register no transport-owned agent families; - start no listeners; - start neither the in-memory nor the durable scheduled job processor. Sending agents deliberately stay enabled: Wolverine's own control queue rides on them, so a node that could not send could not take part in agent assignment at all. The node stays a full cluster member otherwise: it registers, takes part in leader election, uses the control queue, and still advertises and runs the event subscription agents it was stood up for. ScheduleLocalExecutionInMemory already guarded against a null scheduler but attributed it to Durability.Mode; it now names MessagingEnabled when that is the actual cause. Default is true, so nothing changes for existing hosts. The gating follows the pattern DurabilityAgentEnabled already uses in the NodeAgentController constructor. Worth a maintainer's eye: a node that registers no durability family also cannot *assign* those agents while it is the leader, since EvaluateAssignmentsAsync iterates the registered families. That is pre-existing behaviour of DurabilityAgentEnabled which this flag inherits rather than introduces, but it means such a node arguably should not be leader-eligible. Left alone here rather than changing leader election in the same PR.
2db8707 to
9e0cc94
Compare
|
Thanks — all three land. Two are fixed, the third I would like your call on. Stray carriage returns. Correct, and worse than it looked: my editor had rewritten several files to CRLF, which was burying ~500 lines of whitespace churn. Normalised every touched file back to the repo's LF and force-pushed. The diff is 169/9 now, all real changes, 0 stray CRs. Log wording. Also correct — The leader problem. You reached the same conclusion I flagged in the description, independently, which I take as a strong signal it should not ship as a documented caveat. Of your two options I also prefer the second — keep the families registered so a leader can still evaluate and reassign them, and suppress only this node's advertised capabilities. Where I got stuck, and why I would rather ask than guess: If it does, a small decorator that delegates everything except I have not run this on a real cluster — the tests cover the wiring, not the assignment behaviour under leadership change. |
|
The one red check,
Everything else is green (build, test, and the other 29 broker/persistence matrices). Context for why I care about this one landing: we spent today failing to get a release with a projection version bump through our canary, and the root of it is that a bumped shard takes minutes to start rather than milliseconds (JasperFx/jasperfx#594), which then breaks three separate things in the distribution layer (#3748, #3749, #3750). An event-subscription-only node is what would let us warm a new projection version on pods that serve no traffic and hand the caught-up state to the serving fleet — i.e. it takes the replay out of the deploy path entirely instead of trying to make the deploy survive it. Happy to test whatever shape you prefer against our 512-database, ~6,500-agent cluster. |
|
Answering the direct question first, because it has a definite answer and it isn't the one you were hoping for. No — the grid does not filter candidate nodes by capability on the paths these three families actually use. So the decorator shape doesn't work, and for durability it can't be made to work by that route at all. Here's the whole picture, since you were right to not want to guess at it. Capabilities only ever come from static families
foreach (var controller in _agentFamilies.Values.OfType<IStaticAgentFamily>())
{
current.Capabilities.AddRange(await controller.SupportedAgentsAsync());
}You'd spotted that Capability-aware distribution exists — but not on these pathsThe grid does have capability matching, in
Those are the event-subscription paths. All three messaging families take the capability-blind ones:
So an empty And switching The seamBoth halves are needed, and only the second is new work:
The exclusion has to be per scheme, not global — the entire point of these nodes is that event-subscription agents still land on them. On the leader questionPinned agents fall back; the node stays leader-eligible. Leader-ineligibility deadlocks a cluster that is entirely warming nodes, and keeping that shape working is the point. So Push that if you'd like; happy to review. Two things worth knowing before you do:
|
|
Following up on my own comment above, because I've changed my mind about the shape rather than the analysis — please don't start on that I'm taking #3746 out of this release. Everything I wrote about the assignment internals still holds, and that's precisely the problem: making this work correctly means a new node-level attribute travelling node→leader, per-scheme eligibility inside What I'd do instead, and it should work todayRun the warming fleet as its own Wolverine cluster: a different opts.ServiceName = "trips-projection-warmer";
opts.PersistMessagesWithPostgresql(connectionString, schemaName: "warmer_wolverine");
// or, on the Marten/Polecat integration:
// .IntegrateWithWolverine(o =>
// {
// o.MessageStorageSchemaName = "warmer_wolverine";
// o.TransportSchemaName = "warmer_wolverine";
// });The durability schema is the cluster boundary. That gets you what Your projection agents are unaffected, because they're distributed per-cluster off the projections each host registers. The warming fleet registers V2, the serving fleet registers V1, and — as you noted in the description — the version is part of the agent identity, so the two sets are disjoint by construction. That disjointness is what makes the isolation safe, and it's worth checking deliberately rather than assuming: two independent clusters registering the same projection version against one event store would both try to assign it, which is the failure mode this arrangement has instead of the one it fixes. Practical notes: point listeners/queues at the warming fleet only if you mean to, since anything delivered there stays there; and the new schema needs creating, which the usual resource application will do. This is the official guidance for nowI'll document it as the supported way to warm a projection version off to one side. If it turns out to have a sharp edge in practice on your cluster, that's a much better-grounded case for a framework-level flag than the one we have today, and I'd rather add What isn't droppedThe bugs you found underneath this are the real prize and none of them are affected:
Thank you for the PR, genuinely — the description and the follow-ups were unusually careful, and finding the leader-cannot-assign-what-it-doesn't-register hazard was worth the exercise on its own. I'll leave it open rather than close it out from under you; it just isn't tracking for this release. Once #598 lands and we can see what's actually left of the problem, it's worth picking back up. What I'd still like most is the pinned-prerelease-plus-measurement arrangement we discussed. Nothing in this chain reproduces at dev scale, and your 512-database canary is the only real feedback loop any of it has. |
|
We're going to try to solve the real issue a different way. This is generating way too much complexity that will likely be a problem later. |
Implements the ask in #3746. Happy to reshape any of it — the naming, the surface, or the placement — this is a first cut to make the discussion concrete rather than a take-it-or-leave-it.
The problem
There is no supported way to run a node that takes part in event subscription (projection) agent assignment while doing no message handling at all.
The use case is warming a read model before it serves traffic. When a projection version is bumped, its tables start empty and are rebuilt from the beginning of the event store; until that finishes, a node running the new version serves incomplete read models. Standing up a separate set of nodes to build the new version off to one side is the natural fix — and it works, because the projection version is part of the agent identity, so old and new nodes advertise disjoint projection agents.
But those warming nodes must be full Wolverine nodes to be assigned the new version's projection agents, and that also hands them message handling. Listeners we can stop via
IEndpointCollection.StartListenerAsync/StopListenerAsync. Durability agents we cannot: they are assigned per message store, are version-independent, and a durability agent is not a listener, so it recovers and executes persisted envelopes on a node whose read models are half-built.NodeAgentController.DisableAgentsAsyncis documented "STRICTLY FOR TESTING", andSolo/Serverless/MediatorOnlyare not usable —Balancedis what enables the leader election and control queue the distribution needs.For scale context: our deployment is 512 shard databases and ~854 tenants, so about 5 100 projection agents plus 512 durability agents.
What this adds
DurabilitySettings.MessagingEnabled(defaulttrue), plus a fluentservices.RunWolverineForEventSubscriptionDistributionOnly()alongside the existingDisableAllWolverineMessagePersistence/UseWolverineSoloMode.When false, the node:
It stays a full cluster member otherwise — registers, takes part in leader election, uses the control queue, and still advertises and runs the event subscription agents it was stood up for.
The gating follows the pattern
DurabilityAgentEnabledalready uses in theNodeAgentControllerconstructor, so the mechanism should look familiar.One thing worth your eye
A node that registers no durability family also cannot assign those agents while it is the leader, because
EvaluateAssignmentsAsynciterates the registered families. That is pre-existing behaviour ofDurabilityAgentEnabledwhich this flag inherits rather than introduces — but it does suggest such a node should not be leader-eligible.I deliberately did not touch leader election in the same PR. If you would rather this flag also made the node ineligible, or rather the families stayed registered with
SupportedAgentsAsync()returning empty (so a leader could still assign them to others while never taking them itself), say which and I will rework it. The second shape looks cleaner to me but is a bigger behavioural change and I would rather not guess.Testing
Three unit tests in
CoreTests.Runtime.Agents.messaging_disabled_nodecovering the listener families being absent when off, present by default, and the injected event subscription family still being registered when off. The neighbouring agent tests still pass.Default is
true, so nothing changes for existing hosts.