Skip to content

RPC: use RDMA completion channel to not spin - #29440

Merged
am17an merged 2 commits into
masterfrom
aman/rpc_rdma_idle
Sep 27, 2026
Merged

am17an merged 2 commits into
masterfrom
aman/rpc_rdma_idle

Conversation

@am17an

@am17an am17an commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Currently the RPC server spins 100% on rdma_poll, we can arm a completion channel to wake up the thread when there is an event or a peer disconnect. I saw no performance drop with RPC -sm layer or -sm tensor with this change, the only change was the CPU went from 100% to 0.1% usage.

Additional information

Requirements

@am17an
am17an requested a review from a team as a code owner September 25, 2026 18:42
@github-actions github-actions Bot added the ggml changes relating to the ggml tensor library for machine learning label Sep 25, 2026
Comment thread ggml/src/ggml-rpc/transport.cpp
@am17an am17an changed the title RPC: use RDMA completion queue to not spin RPC: use RDMA completion channel to not spin Sep 25, 2026
@ggerganov ggerganov self-assigned this Sep 25, 2026
@ggerganov

Copy link
Copy Markdown
Member

I'll run a few tests tomorrow. cc @rgerganov

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In transport-apple.cpp, add a TODO comment to implement similar mechanism. Add a reference to this PR.

static constexpr size_t RDMA_CHUNK = 256 * 1024; // 256 KiB per send/recv (fits default 8 MiB memlock)
static constexpr int RDMA_RX_DEPTH = 24; // pre-posted recv ring: 24 × 256 KiB = 6 MiB
// keep polling the CQ for this long after the last activity, then sleep until the next completion
static constexpr auto RDMA_SPIN_TIME = std::chrono::milliseconds(100);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't quite understand why we need this but I guess this is how it works

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My guess is that if we sleep unconditionally without a spin/polling phase, there will be a performance hit from the large amount of sleep + wake cycles.

@am17an
am17an merged commit d7fb90e into master Sep 27, 2026
21 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
* RPC: use RDMA completion queue to not spin

* add TODO for apple RDMA
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants