Skip to content

fix: retry a shuffle page allocation that Spark failed after dropping the task's entry - #6403

Open
dwsmith1983 wants to merge 17 commits into
apache:mainfrom
dwsmith1983:fix/6304-allocator-parked-acquire
Open

dwsmith1983 wants to merge 17 commits into
apache:mainfrom
dwsmith1983:fix/6304-allocator-parked-acquire

Conversation

@dwsmith1983

@dwsmith1983 dwsmith1983 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Closes #6304.

Rationale for this change

Spark's ExecutionMemoryPool.acquireMemory registers the task's memoryForTask entry once, before its wait loop, and reads it with memoryForTask(taskAttemptId) on every pass. releaseMemory removes the entry when the task's balance reaches zero and wakes the waiters. A caller that wakes after the removal throws NoSuchElementException: key not found: <taskAttemptId>. The code is the same from Spark 3.4 through 4.2.

TaskMemoryManager.releaseExecutionMemory takes no monitor, so a release can run while an acquire of the same task is parked. When CometUnifiedShuffleMemoryAllocator has a page or pointer array request parked below the task's minimum share and Comet's native consumer releases the task's last bytes, the parked request fails with that exception. The shuffle writers only expect SparkOutOfMemoryError from the allocator, so the task fails instead of spilling.

The root cause is in Spark, filed as SPARK-59827 with a fix in apache/spark#59103. Until that lands in the versions Comet supports, Comet can only guard its own callers, and Spark's own operators in the task stay exposed.

What changes are included in this PR?

CometUnifiedShuffleMemoryAllocator.allocate and a new allocateArray override run the Spark call through a helper that catches this one exception, identified by its message prefix (key not found: ) and a frame in ExecutionMemoryPool, and calls Spark again. The retry registers the task's entry again, checks its share and parks again if it has to, so the caller ends up with what Spark would have granted had the entry stayed. Retrying is safe because Spark only removes the entry at a zero balance, so the failed call was granted nothing and the allocator's used and Spark's counters have not moved. After three attempts it throws the same SparkOutOfMemoryError it uses for a refused page, with the last exception as the cause, and the writers spill as they would for any other refusal. Any other exception is rethrown. A retry logs at info level and giving up logs a warning.

allocateArray needs its own override because the inherited MemoryConsumer.allocateArray asks Spark for the page directly rather than through allocate, and the sorters' pointer arrays come through it.

#6310 adds the same retry for Comet's native acquires in CometTaskMemoryManager, with its own matcher. Whichever of the two lands second will move both onto one shared helper.

The contributor guide's memory management and JVM shuffle pages describe the failure and the retry.

How are these changes tested?

A new CometUnifiedShuffleMemoryAllocatorSuite, registered in both PR build workflows, covers:

  • the race from the issue, for a page and for a pointer array: on a 100 byte off-heap UnifiedMemoryManager, the task holds 10 bytes through a CometTaskMemoryManager and another task holds 90. An allocation on a second thread parks in Spark below the task's minimum share, and the native consumer releases the task's last 10 bytes. On main the parked allocation throws key not found: 0. With this change it is granted in full once the other task frees its memory, with one retry logged and the allocator's accounting back at zero afterwards.
  • a NoSuchElementException with a different message, or without an ExecutionMemoryPool frame, is rethrown unchanged after a single call.
  • when every attempt throws the matching exception, the allocator gives up after the cap with SparkOutOfMemoryError (UNABLE_TO_ACQUIRE_MEMORY, the requested bytes, zero received, and the last exception as the cause), with the accounting unchanged.
  • an allocation that loses its entry on the first attempts is granted on the last one.

The suite passes on Spark 3.4, 3.5, 4.0 and 4.1, and the race tests fail without the retry.

… the task's entry

Spark's ExecutionMemoryPool registers a task's memoryForTask entry once, before
its wait loop, and reads it on every pass, while releaseMemory removes the
entry when the task's balance reaches zero. When the JVM shuffle allocator's
page or pointer array request was parked below the task's minimum share and
Comet's native memory consumer released the task's last bytes, the request
woke to NoSuchElementException "key not found: <taskAttemptId>" and failed the
task, because the shuffle writers only expect SparkOutOfMemoryError.

CometUnifiedShuffleMemoryAllocator now retries allocate and allocateArray when
Spark throws that exception from ExecutionMemoryPool. The retry registers the
task again and waits as the first attempt would have. Nothing is counted or
allocated before the exception, so retrying cannot leak. After three attempts
it refuses the request with the same SparkOutOfMemoryError it uses for a
refused page, with the last exception as its cause. The Spark issue is
SPARK-59827.
Only the shuffle allocator's callers are guarded; Spark's own operators in the
task can still hit the removed entry until Spark re-registers a waiting task
(apache/spark#59103).
@github-actions github-actions Bot added bug Something isn't working area:shuffle Shuffle (JVM and native) area:memory Memory pools, reservations, OOM handling labels Sep 29, 2026
@dwsmith1983
dwsmith1983 force-pushed the fix/6304-allocator-parked-acquire branch from ad2225d to 0de3286 Compare September 29, 2026 15:21
@dwsmith1983

Copy link
Copy Markdown
Contributor Author

@andygrove could you approve a CI run on 5f1e1da59 and take a look when you get a chance?

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

  • Prior state and problem: A waiting JVM shuffle allocation could fail with NoSuchElementException when another consumer released the task’s last reserved bytes.
  • Design approach: Retry the specific missing-entry failure for both shuffle pages and pointer arrays through one private helper.
  • Correctness / compatibility analysis: Checked Spark’s memory-pool, allocation and exception implementations in 3.4.3, 3.5.9, 4.0.4, 4.1.3 and 4.2.0. The relevant failure occurs before a grant or page allocation, so retrying preserves accounting and Spark’s fair-share checks.
  • Key design decisions: Message-and-stack matching limits retries to the intended failure. The helper keeps the abstraction local. Successful allocations still make one Spark call, with exception inspection and logging confined to the failure path.
  • Implementation sketch: Wrap allocatePage and inherited allocateArray, cap attempts at three, preserve the final exception as the cause, register the new suite in both CI workflows, and update the memory and shuffle documentation.
  • Behavioral changes worth calling out: Exhausted retries become SparkOutOfMemoryError, reaching existing shuffle refusal/spill handling. Unrelated exceptions propagate unchanged. Spark-owned callers remain outside this workaround.
  • Suggested improvements: None at P1/P2. No introduced P1/P2 issues found within this review.

Reviewed the entire six-file diff from 9f68a4144fdefc70ec26f3d80b0e4eafe72a8df4 to 20e1327dd353b32aec9be2833077766573c6057b. The PR is not a draft. Read the existing discussion and verified there are no reviews, inline comments or review threads containing unresolved concerns. Routed skills: review-comet-pr, review-comet-memory-pr, review-comet-shuffle-pr.

Exact-head CI: Comet CI and CodeQL report action_required, awaiting approval. Only the labeling check passed. There is no completed build/test verdict.

Validation: Independently compiled the exact-head allocator, supporting classes and new suite against cached Spark 4.1.3 dependencies using JDK 17. All five tests passed. Both race tests failed against the base allocator with NoSuchElementException: key not found: 0, confirming regression coverage. Full Maven/native builds, end-to-end shuffle tests and performance benchmarks were not run. Other supported Spark versions were source-checked only.

@dwsmith1983
dwsmith1983 requested a review from sunchao October 4, 2026 15:04

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

  • Prior state and problem: A waiting JVM shuffle allocation could throw NoSuchElementException when another consumer released the task’s last reserved bytes.
  • Design approach: Retry that specific Spark failure for both pages and pointer arrays through one private helper.
  • Correctness / compatibility analysis: Checked the relevant Spark sources in 3.4.3, 3.5.9, 4.0.4, 4.1.3 and 4.2.0. The missing-entry lookup fails before a grant or page allocation. Retrying re-registers the task and preserves Spark’s fair-share checks and allocation accounting.
  • Key design decisions: Message-and-stack matching narrows the retry to the intended failure. The shared helper keeps both allocation paths consistent. Successful requests still make one Spark allocation call through a supplier wrapper. Stack inspection and logging occur only on failure. Performance was not benchmarked.
  • Implementation sketch: Wrap allocatePage and inherited allocateArray, cap attempts at three, preserve the final exception as the cause, register the new suite in both CI workflows, and update the memory and shuffle documentation.
  • Behavioral changes worth calling out: Exhausted retries become SparkOutOfMemoryError, reaching existing allocation-refusal handling. Unrelated exceptions propagate unchanged. Compared with branch-1.1 at 992c806a7e38c2e88bd018aa5774164b0850e1fa, whose allocator matches the supplied base, this is an intended recovery improvement. Spark-owned callers remain exposed to the upstream race.
  • Suggested improvements: None at P1/P2. No introduced P1/P2 issues found within this review.

Reviewed the entire six-file diff from fef94f6cd78b18151dff57b7a936798385356de5 to 48458d0840317700ce4febb184e847243c8aee05. The PR is not a draft. Read the existing review and issue comment. There are no inline comments, review threads or substantiated unresolved P1/P2 concerns. Routed skills: review-comet-pr, review-comet-memory-pr, review-comet-shuffle-pr.

Exact-head CI: Comet CI and CodeQL report action_required, awaiting approval. The labeling check passed. There is no completed build/test CI verdict.

Validation: Independently compiled the exact-head allocator, supporting classes and suite against cached Spark 4.1.3 dependencies using JDK 21 with Java 17 source/target settings. All five tests passed. Both race tests failed against the supplied base allocator with NoSuchElementException: key not found: 0, confirming regression coverage. Full Maven/native builds, end-to-end shuffle tests and benchmarks were not run. Other supported Spark versions were source-checked only.

@dwsmith1983

Copy link
Copy Markdown
Contributor Author

@andygrove could you take a look at this one when you have a moment? It is merged up with main, and the CI run on the latest commit is waiting for workflow approval.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:memory Memory pools, reservations, OOM handling area:shuffle Shuffle (JVM and native) bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A JVM consumer's parked page allocation fails when another consumer of the task empties its balance

2 participants