Skip to content

Leader assigns the wolverinedb durability agent to nodes with DurabilityAgentEnabled = false — assignment never converges and outgoing recovery never runs #3954

Description

@beriniwlew

Summary

In a Balanced cluster, the leader assigns the wolverinedb://… durability agent to nodes running Durability.DurabilityAgentEnabled = false. Such a node cannot start the agent (ArgumentOutOfRangeException: Unrecognized agent scheme 'wolverinedb'), the leader re-issues the identical assignment forever, and — because no durability agent ends up running anywhere for that store — owner_id = 0 outgoing envelopes are never recovered. All queue tables read zero the whole time, so the failure is silent. We found a 9-day backlog of undelivered envelopes in a staging environment this way.

Observed on 6.17.3 and 6.24.4; from reading main, the relevant path looks unchanged.

Setup

Two-node cluster on PostgreSQL persistence + the PostgreSQL queue transport:

  • Node A (worker/consumer): durability enabled, ListenToPostgresqlQueue(...).ExclusiveNodeWithParallelism(...).
  • Node B (web producer): stages envelopes inside caller-owned EF transactions (so they land in wolverine_outgoing_envelopes with owner_id = 0 and rely on durability-agent recovery), no listeners, conventional discovery disabled, and Durability.DurabilityAgentEnabled = false — the intent being "this node stages but runs no recovery/scheduled agents".

Observed behavior

  1. Node A happened to hold leadership. The leader assigned the wolverinedb://postgresql/<host>/<db>/wolverine agent to node B.
  2. Node B throws on the StartAgent command:
    System.ArgumentOutOfRangeException: Unrecognized agent scheme 'wolverinedb' (Parameter 'uri')
       at Wolverine.Runtime.Agents.NodeAgentController.startWithRetriesAsync(Uri agentUri)
       at Wolverine.Runtime.Agents.NodeAgentController.StartAgentAsync(Uri agentUri)
       at Wolverine.Runtime.Agents.StartAgent.ExecuteAsync(...)
    
    and the envelope moves to the error queue; the leader logs Unable to confirm that agent wolverinedb://… started on node … and Giving up on a batched agent command after no progress for 00:05:00.
  3. The leader re-evaluates every ~5 minutes and makes the same assignment again (wolverine_node_records shows an AssignmentChanged row every ~5m10s, indefinitely). It never converges to the capable node.
  4. Result: no durability agent runs for the store → CheckRecoverableOutgoingMessagesOperation never executes → producer-staged envelopes sit in wolverine_outgoing_envelopes (owner_id = 0) forever, while the queue tables stay empty.

Why (from source)

  • A node with DurabilityAgentEnabled = false never registers the store as an agent family, so it genuinely cannot start the agent (NodeAgentController ctor only does _agentFamilies[_runtime.Stores.Scheme] = _runtime.Stores when the flag is true).
  • Node capabilities correctly exclude the scheme in that case (StartLocalProcessing populates WolverineNode.Capabilities from the registered families' SupportedAgentsAsync), so the leader has the data to avoid the node.
  • But the assignment path for the store family doesn't consult capabilities: MessageStoreCollection.EvaluateAssignmentsAsync uses DistributeEvenly (6.x) / DistributeEvenlyWithAffinity (main, Database affinity is per agent family, not per database: 73% of shard databases have durability and projections on different nodes #3785, falling back to the even spread when there are no projection agents), and both spread across all grid nodes. The remainder pass even prefers non-leader nodes (_nodes.FirstOrDefault(x => !x.IsLeader …)), which in a two-node cluster deterministically targets the incapable producer whenever the consumer is leader.
  • Capability matching (AllNodesHaveSameCapabilities / MatchAgentsToCapableNodesFor) exists in DistributeByGroupAffinity and DistributeEvenlyWithBlueGreenSemantics, but not in the plain even spread used for wolverinedb agents.

Repro sketch (deterministic, single node)

A single-node Balanced cluster with DurabilityAgentEnabled = false reproduces it without any second node: the node elects itself leader, assigns the durability agent to itself, throws Unrecognized agent scheme, and an envelope staged into wolverine_outgoing_envelopes with owner_id = 0 is never recovered into the queue. (That's how we pinned it in a regression test.)

Expected

Either of these would resolve it:

Workaround

We removed DurabilityAgentEnabled = false from the producer nodes, so any node the leader picks can run the agent. That's safe in practice because a Main-role store's cluster-assigned agent doesn't auto-start scheduled-job polling (GH-3376/GH-3439) and incoming recovery no-ops for endpoints without a local listener — but it does mean the durability settings (e.g. KeepAfterMessageHandling) must be kept identical across all nodes, since the agent runs with whichever node hosts it.

Thanks for Wolverine — happy to provide more logs or turn the single-node repro into a PR-able failing test if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions