Skip to content

Document capacity-aware agent assignment (GH-3959 / #4297) #4594

Description

@jeremydmiller

Follow-up to #4297 (GH-3959), which merged DurabilitySettings.CapacityAwareAssignment, NodeOverloadThreshold, OverloadShedBatchSize and INodeLoadMonitor with no user-facing documentation.

docs/guide/durability/leadership-and-troubleshooting.md is the natural home, next to the existing Solo Mode and metrics sections.

Text should be in Jeremy's voice — flagging this as an issue rather than drafting the prose.

What the page has to cover, because each of these is a way to get burned:

  • It is opt-in and it needs a schema change. Enabling it provisions load_factor on wolverine_nodes. Under AutoCreate.None, or in a process without DDL rights, apply the migration before turning the flag on.
  • It is PostgreSQL-only today, and silently inert elsewhere (GH-3959 follow-up: advertise node load from the remaining message stores #4593).
  • It only affects even distribution — not group affinity, not blue/green, not durability affinity (GH-3959 follow-up: capacity awareness for group-affinity and blue/green distribution #4592). Which means it does not yet help the multi-database async-daemon shape that motivated it.
  • What INodeLoadMonitor is for, and how to write one. Pending GH-3959 follow-up: the default node load monitor reads the wrong denominator #4589, this is likely to become required rather than defaulted, so the page should lead with "you supply this" rather than treating the built-in as the normal path.
  • The two thresholds and the band between them. NodeOverloadThreshold is the shed line; the receive line sits 10 points below it; a node between the two neither sheds nor receives. Worth a sentence on why, since a single threshold is what everyone expects.
  • What "no node has headroom" looks like in production — agents deliberately left unassigned, which is a quieter symptom than a crash loop and needs an operator to know it can happen. See the CritterWatch alert issue.

Larger framing, and probably a second page eventually: this is the first piece of what could be real dynamic-scaling support in Wolverine — the cluster having an opinion about how much work a node should take, rather than dividing by node count and hoping. Worth saying where this is going, so the feature doesn't read as a lone knob.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions