Skip to content

docs: add a Celeborn integration guide - #6373

Merged
sunchao merged 2 commits into
apache:mainfrom
pingzh:pingzh-celeborn-integration-docs
Oct 1, 2026
Merged

sunchao merged 2 commits into
apache:mainfrom
pingzh:pingzh-celeborn-integration-docs

Conversation

@pingzh

@pingzh pingzh commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Closes #6118.

Rationale for this change

Users need a clear setup guide for running Comet with Celeborn. With the currently released Celeborn 0.6.x and 0.7.x clients, Comet accelerates supported query operators while shuffle uses Celeborn's existing Spark integration. Setting spark.comet.shuffle.mode=native does not enable Comet native shuffle with those clients.

What changes are included in this PR?

  • Add a Celeborn guide under Integrations with support status, client artifact coordinates, driver/executor startup classpaths, configuration, plan verification, and troubleshooting.
  • Keep the setup focused on the shuffle path available with released clients and omit instructions for enabling unavailable native shuffle.
  • Explain how to check shuffle activity and why Spark's remote-read byte counters alone do not establish that Celeborn stored the data.
  • Update the metrics and configuration references, and clarify that the existing native remote shuffle tuning settings do not tune Celeborn's Spark shuffle implementation.

The guide was checked against #5352, its linked upstream PRs, subsequent compatibility fixes, and the current implementation and tests.

How are these changes tested?

  • Prettier check on all changed Markdown files and git diff --check passed.
  • Sphinx HTML build passed with existing warnings for missing generated pages, existing cross-references, and unavailable Mermaid tooling. No warnings reference the edited pages.
  • Checked all 45 rendered local content links and anchors across the four revised pages; no broken links or references to removed sections remain.
  • Reviewed setup, compatibility, fallback, and metrics statements against source and tests.

Documentation only; no Spark/Celeborn cluster tests or full release-site generation were run.

image

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 29, 2026
@pingzh
pingzh force-pushed the pingzh-celeborn-integration-docs branch from de3d569 to f6db86c Compare September 29, 2026 15:05
@pingzh
pingzh marked this pull request as ready for review September 30, 2026 16:08

@ajsquared ajsquared left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the seven documentation changes against the relevant Comet, Celeborn, and Spark source. No actionable findings. Tests were not run; CI and production systems were not inspected.

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

Reviewed all seven documentation changes at dba82066dc34b4f0afd723336afbec3654208ea0 against merge base 8805e3f066440752ac98d5b48e197b43e4fb9ebd, with integration checks against current target 1628c520096e12ba3fb13988d6ee235070e77e07. No introduced or materially worsened P1/P2 findings. The 204 target-only file changes are separate from this PR.

  • Prior state and problem: Setup, compatibility and native-only tuning were mixed together. Users needed a clear path for combining Comet execution with released Celeborn clients.
  • Design approach: Add a dedicated Celeborn integration guide and connect it from configuration, tuning, metrics and navigation. Keep native remote-shuffle tuning in its own page.
  • Correctness / compatibility analysis: Upstream Comet selection and delegation match the guide. Celeborn 0.6.0/0.7.0 source supports the native-client limitation, shaded-client naming, Kryo requirement and fallback behavior. Spark 3.5.0/3.5.9 source supports the startup-classpath and stage/RDD/shuffle-ID instructions. This is source inference, not executed reproduction.
  • Key design decisions: Use ordinary Spark/Celeborn shuffle for released clients, retain acceleration of supported Comet operators, require dependencies on driver and executor startup classpaths, and distinguish plan selection from actual storage destination.
  • Implementation sketch: Seven documentation files change. The new guide provides setup, verification and troubleshooting; the remaining edits update links, metrics interpretation and native admission guidance. Runtime implementation and configuration defaults are unchanged.
  • Behavioral changes worth calling out: The example uses auto; selecting native does not enable native shuffle with the documented clients. Remote-read counters alone do not prove Celeborn storage. The native frame/admission settings do not tune Celeborn's existing Spark shuffle. No new runtime performance behavior is introduced.
  • Suggested improvements: None at the requested P1/P2 threshold.

Validation is source-only. Current head CI status metadata reports seven successful and sixteen skipped checks, including skipped Linux/macOS builds and Spark SQL suites. No tests, builds, rendered pages, link probes, packages, artifacts, CI logs or deployments were inspected or executed. Actual release artifact availability and deployed behavior remain unverified.

@sunchao
sunchao added this pull request to the merge queue Oct 1, 2026
Merged via the queue into apache:main with commit 9c42391 Oct 1, 2026
23 checks passed
@pingzh
pingzh deleted the pingzh-celeborn-integration-docs branch October 1, 2026 14:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

docs: add a user-facing Celeborn integration guide

3 participants