Repository navigation
docs: add a Celeborn integration guide - #6373
Conversation
de3d569 to
f6db86c
Compare
ajsquared
left a comment
There was a problem hiding this comment.
Reviewed the seven documentation changes against the relevant Comet, Celeborn, and Spark source. No actionable findings. Tests were not run; CI and production systems were not inspected.
sunchao
left a comment
There was a problem hiding this comment.
Summary
Reviewed all seven documentation changes at dba82066dc34b4f0afd723336afbec3654208ea0 against merge base 8805e3f066440752ac98d5b48e197b43e4fb9ebd, with integration checks against current target 1628c520096e12ba3fb13988d6ee235070e77e07. No introduced or materially worsened P1/P2 findings. The 204 target-only file changes are separate from this PR.
- Prior state and problem: Setup, compatibility and native-only tuning were mixed together. Users needed a clear path for combining Comet execution with released Celeborn clients.
- Design approach: Add a dedicated Celeborn integration guide and connect it from configuration, tuning, metrics and navigation. Keep native remote-shuffle tuning in its own page.
- Correctness / compatibility analysis: Upstream Comet selection and delegation match the guide. Celeborn 0.6.0/0.7.0 source supports the native-client limitation, shaded-client naming, Kryo requirement and fallback behavior. Spark 3.5.0/3.5.9 source supports the startup-classpath and stage/RDD/shuffle-ID instructions. This is source inference, not executed reproduction.
- Key design decisions: Use ordinary Spark/Celeborn shuffle for released clients, retain acceleration of supported Comet operators, require dependencies on driver and executor startup classpaths, and distinguish plan selection from actual storage destination.
- Implementation sketch: Seven documentation files change. The new guide provides setup, verification and troubleshooting; the remaining edits update links, metrics interpretation and native admission guidance. Runtime implementation and configuration defaults are unchanged.
- Behavioral changes worth calling out: The example uses
auto; selectingnativedoes not enable native shuffle with the documented clients. Remote-read counters alone do not prove Celeborn storage. The native frame/admission settings do not tune Celeborn's existing Spark shuffle. No new runtime performance behavior is introduced. - Suggested improvements: None at the requested P1/P2 threshold.
Validation is source-only. Current head CI status metadata reports seven successful and sixteen skipped checks, including skipped Linux/macOS builds and Spark SQL suites. No tests, builds, rendered pages, link probes, packages, artifacts, CI logs or deployments were inspected or executed. Actual release artifact availability and deployed behavior remain unverified.
Which issue does this PR close?
Closes #6118.
Rationale for this change
Users need a clear setup guide for running Comet with Celeborn. With the currently released Celeborn 0.6.x and 0.7.x clients, Comet accelerates supported query operators while shuffle uses Celeborn's existing Spark integration. Setting
spark.comet.shuffle.mode=nativedoes not enable Comet native shuffle with those clients.What changes are included in this PR?
The guide was checked against #5352, its linked upstream PRs, subsequent compatibility fixes, and the current implementation and tests.
How are these changes tested?
git diff --checkpassed.Documentation only; no Spark/Celeborn cluster tests or full release-site generation were run.