Skip to content

fix(storage): stop failed partition consumption before releasing dispatch capacity #1499

Description

@DecisionNerd

Problem

consume_in_partition_order releases a consumed partition's dispatch slot and notifies workers before propagating the consume callback's error. A worker can claim replacement work during that interval even though consumption has failed.

Observed while validating #1416 against current main ea717991: cargo test --release -p graphforge-storage -p graphforge-value --lib reported 1214 passed, one failed, six existing ignored. consume_error_stops_the_pool_and_abandoned_loads_are_never_observed failed with abandoned=3 against its valid upper bound of two. The unchanged scheduler originates in #1459 (76462543). Independent source review confirms that the failed consume at index 1 increments released from 1 to 2, admitting partition 4 alongside abandoned partitions 2 and 3 before stop is set.

Scope

Propagate consume failure before releasing its slot or notifying workers. Only successful consumption admits replacement work. Preserve ordered error selection, worker joining, cancellation, and the existing bounded-window assertion. This is a verified validation blocker for #1416; it does not take ownership of the other session's #1439 or #1464 work.

Acceptance

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions