You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Leader assigns the wolverinedb durability agent to nodes with DurabilityAgentEnabled = false — assignment never converges and outgoing recovery never runs #3954
In a Balanced cluster, the leader assigns the wolverinedb://… durability agent to nodes running Durability.DurabilityAgentEnabled = false. Such a node cannot start the agent (ArgumentOutOfRangeException: Unrecognized agent scheme 'wolverinedb'), the leader re-issues the identical assignment forever, and — because no durability agent ends up running anywhere for that store — owner_id = 0 outgoing envelopes are never recovered. All queue tables read zero the whole time, so the failure is silent. We found a 9-day backlog of undelivered envelopes in a staging environment this way.
Observed on 6.17.3 and 6.24.4; from reading main, the relevant path looks unchanged.
Setup
Two-node cluster on PostgreSQL persistence + the PostgreSQL queue transport:
Node A (worker/consumer): durability enabled, ListenToPostgresqlQueue(...).ExclusiveNodeWithParallelism(...).
Node B (web producer): stages envelopes inside caller-owned EF transactions (so they land in wolverine_outgoing_envelopes with owner_id = 0 and rely on durability-agent recovery), no listeners, conventional discovery disabled, and Durability.DurabilityAgentEnabled = false — the intent being "this node stages but runs no recovery/scheduled agents".
Observed behavior
Node A happened to hold leadership. The leader assigned the wolverinedb://postgresql/<host>/<db>/wolverine agent to node B.
Node B throws on the StartAgent command:
System.ArgumentOutOfRangeException: Unrecognized agent scheme 'wolverinedb' (Parameter 'uri')
at Wolverine.Runtime.Agents.NodeAgentController.startWithRetriesAsync(Uri agentUri)
at Wolverine.Runtime.Agents.NodeAgentController.StartAgentAsync(Uri agentUri)
at Wolverine.Runtime.Agents.StartAgent.ExecuteAsync(...)
and the envelope moves to the error queue; the leader logs Unable to confirm that agent wolverinedb://… started on node … and Giving up on a batched agent command after no progress for 00:05:00.
The leader re-evaluates every ~5 minutes and makes the same assignment again (wolverine_node_records shows an AssignmentChanged row every ~5m10s, indefinitely). It never converges to the capable node.
Result: no durability agent runs for the store → CheckRecoverableOutgoingMessagesOperation never executes → producer-staged envelopes sit in wolverine_outgoing_envelopes (owner_id = 0) forever, while the queue tables stay empty.
Why (from source)
A node with DurabilityAgentEnabled = false never registers the store as an agent family, so it genuinely cannot start the agent (NodeAgentController ctor only does _agentFamilies[_runtime.Stores.Scheme] = _runtime.Stores when the flag is true).
Node capabilities correctly exclude the scheme in that case (StartLocalProcessing populates WolverineNode.Capabilities from the registered families' SupportedAgentsAsync), so the leader has the data to avoid the node.
But the assignment path for the store family doesn't consult capabilities: MessageStoreCollection.EvaluateAssignmentsAsync uses DistributeEvenly (6.x) / DistributeEvenlyWithAffinity (main, Database affinity is per agent family, not per database: 73% of shard databases have durability and projections on different nodes #3785, falling back to the even spread when there are no projection agents), and both spread across all grid nodes. The remainder pass even prefers non-leader nodes (_nodes.FirstOrDefault(x => !x.IsLeader …)), which in a two-node cluster deterministically targets the incapable producer whenever the consumer is leader.
Capability matching (AllNodesHaveSameCapabilities / MatchAgentsToCapableNodesFor) exists in DistributeByGroupAffinity and DistributeEvenlyWithBlueGreenSemantics, but not in the plain even spread used for wolverinedb agents.
Repro sketch (deterministic, single node)
A single-node Balanced cluster with DurabilityAgentEnabled = false reproduces it without any second node: the node elects itself leader, assigns the durability agent to itself, throws Unrecognized agent scheme, and an envelope staged into wolverine_outgoing_envelopes with owner_id = 0 is never recovered into the queue. (That's how we pinned it in a regression test.)
Or DurabilityAgentEnabled = false is documented as unsupported for nodes participating in a Balanced cluster whose store has durability work.
Workaround
We removed DurabilityAgentEnabled = false from the producer nodes, so any node the leader picks can run the agent. That's safe in practice because a Main-role store's cluster-assigned agent doesn't auto-start scheduled-job polling (GH-3376/GH-3439) and incoming recovery no-ops for endpoints without a local listener — but it does mean the durability settings (e.g. KeepAfterMessageHandling) must be kept identical across all nodes, since the agent runs with whichever node hosts it.
Thanks for Wolverine — happy to provide more logs or turn the single-node repro into a PR-able failing test if useful.
Summary
In a Balanced cluster, the leader assigns the
wolverinedb://…durability agent to nodes runningDurability.DurabilityAgentEnabled = false. Such a node cannot start the agent (ArgumentOutOfRangeException: Unrecognized agent scheme 'wolverinedb'), the leader re-issues the identical assignment forever, and — because no durability agent ends up running anywhere for that store —owner_id = 0outgoing envelopes are never recovered. All queue tables read zero the whole time, so the failure is silent. We found a 9-day backlog of undelivered envelopes in a staging environment this way.Observed on 6.17.3 and 6.24.4; from reading
main, the relevant path looks unchanged.Setup
Two-node cluster on PostgreSQL persistence + the PostgreSQL queue transport:
ListenToPostgresqlQueue(...).ExclusiveNodeWithParallelism(...).wolverine_outgoing_envelopeswithowner_id = 0and rely on durability-agent recovery), no listeners, conventional discovery disabled, andDurability.DurabilityAgentEnabled = false— the intent being "this node stages but runs no recovery/scheduled agents".Observed behavior
wolverinedb://postgresql/<host>/<db>/wolverineagent to node B.StartAgentcommand:Unable to confirm that agent wolverinedb://… started on node …andGiving up on a batched agent command after no progress for 00:05:00.wolverine_node_recordsshows anAssignmentChangedrow every ~5m10s, indefinitely). It never converges to the capable node.CheckRecoverableOutgoingMessagesOperationnever executes → producer-staged envelopes sit inwolverine_outgoing_envelopes(owner_id = 0) forever, while the queue tables stay empty.Why (from source)
DurabilityAgentEnabled = falsenever registers the store as an agent family, so it genuinely cannot start the agent (NodeAgentControllerctor only does_agentFamilies[_runtime.Stores.Scheme] = _runtime.Storeswhen the flag is true).StartLocalProcessingpopulatesWolverineNode.Capabilitiesfrom the registered families'SupportedAgentsAsync), so the leader has the data to avoid the node.MessageStoreCollection.EvaluateAssignmentsAsyncusesDistributeEvenly(6.x) /DistributeEvenlyWithAffinity(main, Database affinity is per agent family, not per database: 73% of shard databases have durability and projections on different nodes #3785, falling back to the even spread when there are no projection agents), and both spread across all grid nodes. The remainder pass even prefers non-leader nodes (_nodes.FirstOrDefault(x => !x.IsLeader …)), which in a two-node cluster deterministically targets the incapable producer whenever the consumer is leader.AllNodesHaveSameCapabilities/MatchAgentsToCapableNodesFor) exists inDistributeByGroupAffinityandDistributeEvenlyWithBlueGreenSemantics, but not in the plain even spread used forwolverinedbagents.Repro sketch (deterministic, single node)
A single-node Balanced cluster with
DurabilityAgentEnabled = falsereproduces it without any second node: the node elects itself leader, assigns the durability agent to itself, throwsUnrecognized agent scheme, and an envelope staged intowolverine_outgoing_envelopeswithowner_id = 0is never recovered into the queue. (That's how we pinned it in a regression test.)Expected
Either of these would resolve it:
wolverinedbagents (mirroring what the blue/green and group-affinity paths already do) — related in spirit to Can a node take event-subscription agents but not durability agents? (warming a new projection version before it serves traffic) #3746 / Add Durability.MessagingEnabled for event-subscription-only nodes (GH-3746) #3747, which scoped the inverse case.DurabilityAgentEnabled = falseis documented as unsupported for nodes participating in a Balanced cluster whose store has durability work.Workaround
We removed
DurabilityAgentEnabled = falsefrom the producer nodes, so any node the leader picks can run the agent. That's safe in practice because a Main-role store's cluster-assigned agent doesn't auto-start scheduled-job polling (GH-3376/GH-3439) and incoming recovery no-ops for endpoints without a local listener — but it does mean the durability settings (e.g.KeepAfterMessageHandling) must be kept identical across all nodes, since the agent runs with whichever node hosts it.Thanks for Wolverine — happy to provide more logs or turn the single-node repro into a PR-able failing test if useful.