Repository navigation
fix(storage): stop failed partition consumption before releasing capacity - #1500
Conversation
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: CurateLabs/graphforge/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Description
A failed partition-consume callback released its dispatch slot before its error reached the coordinator stop path. With three workers, failure at partition 1 could therefore admit partition 4 and produce three abandoned loads instead of the valid maximum of two.
Propagate consume failure before releasing capacity or waking workers. Successful consumption retains the existing window behavior; ordered primary errors, stop signaling, and joining remain unchanged. The private release helper makes this ordering invariant directly testable without relying on thread timing.
This is the verified validation blocker for #1416, originating in #1459. It does not include timestamp semantics or the other session's partition-cap work.
Validation
released=1, expected0, in 0.00s) and passes with the fix. The existing threaded abandonment assertion remains unchanged.cargo test --release -p graphforge-storage --lib: 1,213 passed, no failures, six existing ignored tests.make pre-push-fast, andmake gate-registry-check: passed.Closes #1499
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.