You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up to #4297 (GH-3959), which merged DurabilitySettings.CapacityAwareAssignment, NodeOverloadThreshold, OverloadShedBatchSize and INodeLoadMonitor with no user-facing documentation.
docs/guide/durability/leadership-and-troubleshooting.md is the natural home, next to the existing Solo Mode and metrics sections.
Text should be in Jeremy's voice — flagging this as an issue rather than drafting the prose.
What the page has to cover, because each of these is a way to get burned:
It is opt-in and it needs a schema change. Enabling it provisions load_factor on wolverine_nodes. Under AutoCreate.None, or in a process without DDL rights, apply the migration before turning the flag on.
The two thresholds and the band between them.NodeOverloadThreshold is the shed line; the receive line sits 10 points below it; a node between the two neither sheds nor receives. Worth a sentence on why, since a single threshold is what everyone expects.
What "no node has headroom" looks like in production — agents deliberately left unassigned, which is a quieter symptom than a crash loop and needs an operator to know it can happen. See the CritterWatch alert issue.
Larger framing, and probably a second page eventually: this is the first piece of what could be real dynamic-scaling support in Wolverine — the cluster having an opinion about how much work a node should take, rather than dividing by node count and hoping. Worth saying where this is going, so the feature doesn't read as a lone knob.
Follow-up to #4297 (GH-3959), which merged
DurabilitySettings.CapacityAwareAssignment,NodeOverloadThreshold,OverloadShedBatchSizeandINodeLoadMonitorwith no user-facing documentation.docs/guide/durability/leadership-and-troubleshooting.mdis the natural home, next to the existing Solo Mode and metrics sections.Text should be in Jeremy's voice — flagging this as an issue rather than drafting the prose.
What the page has to cover, because each of these is a way to get burned:
load_factoronwolverine_nodes. UnderAutoCreate.None, or in a process without DDL rights, apply the migration before turning the flag on.INodeLoadMonitoris for, and how to write one. Pending GH-3959 follow-up: the default node load monitor reads the wrong denominator #4589, this is likely to become required rather than defaulted, so the page should lead with "you supply this" rather than treating the built-in as the normal path.NodeOverloadThresholdis the shed line; the receive line sits 10 points below it; a node between the two neither sheds nor receives. Worth a sentence on why, since a single threshold is what everyone expects.Larger framing, and probably a second page eventually: this is the first piece of what could be real dynamic-scaling support in Wolverine — the cluster having an opinion about how much work a node should take, rather than dividing by node count and hoping. Worth saying where this is going, so the feature doesn't read as a lone knob.