Skip to content

RPC: add -sm tensor - #26610

Open
am17an wants to merge 5 commits into
masterfrom
rpc_tensor
Open

am17an wants to merge 5 commits into
masterfrom
rpc_tensor

Conversation

@am17an

@am17an am17an commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Add RPC -sm tensor. This is on 2x Sparks connected via RDMA.

For RPC following changes are required:

  1. async graph_compute
  2. custom all_reduce
  3. graph uid cache like CUDA
  4. set_tensor_2d, get_tensor_2d

Looking for feedback @ggerganov @rgerganov

model size params backend ngl n_ubatch sm lm test t/s
deepseek4 ?B MXFP4 MoE 145.63 GiB 284.33 B RPC -1 2048 tensor dio pp2048 619.36 ± 15.42
deepseek4 ?B MXFP4 MoE 145.63 GiB 284.33 B RPC -1 2048 tensor dio tg128 19.75 ± 0.39

Additional information

  sequenceDiagram
      participant C as Client (RPC backend)
      box Server A (rank 0)
          participant A as A: rpc port
          participant Ac as A: comm port
      end
      box Server B (rank 1)
          participant Bc as B: comm port
          participant B as B: rpc port
      end

      Note over C,B: initialization
      C->>A: COMM_INIT (rank 0)
      C->>B: COMM_INIT (rank 1)
      Note over Ac: listen on comm port
      Bc-->>Ac: connect + caps negotiation (transport upgrade, e.g. RDMA)
      A->>C: response (ok)
      B->>C: response (ok)

      Note over C,B: for each subgraph (split at reduction boundaries)
      C->>A: GRAPH_COMPUTE (uid) [GRAPH_RECOMPUTE on reuse]
      C->>B: GRAPH_COMPUTE (uid)
      Note over A: compute subgraph async
      Note over B: compute subgraph async
  
      C->>A: COMM_ALLREDUCE (partial tensor) [fire and forget, no response]
      C->>B: COMM_ALLREDUCE (partial tensor)

      Note over A: sync pending graph<br/>cast F32→BF16 if ne ≥ 32768<br/>copy partial to send buffer
      Note over B: sync pending graph<br/>cast F32→BF16 if ne ≥ 32768<br/>copy partial to send buffer

      Ac-->>Bc: partial A (rank 0 sends first)
      Bc-->>Ac: partial B (rank 1 receives first)

      Note over A: upload partial B<br/>dst = dst + partial B (async ADD)
      Note over B: upload partial A<br/>dst = dst + partial A (async ADD)

      Note over C,B: read back the output
      C->>A: GET_TENSOR (output)
      Note over A: sync all backends
      A->>C: data
Loading

Requirements


Stack created with GitHub Stacks CLI • Give Feedback 💬

@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning labels Aug 5, 2026
@ryan5rdx

ryan5rdx commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

confirmed working on metal(RDMA 2 x M3 Ultra with #26421), testing same model (ds4 MXFP4):
tg2048: 9.61 t/s
pp2048: 166.05 t/s

it does however break with dspark applied because some ops it depends on appear to not be supported with TP (add across the split), so this is without any mtp.

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

@ryan5rdx yeah I know it doesn't work with dspark, but what do you get when for just -sm layer as tg2048? It seems kinda low based on the numbers you posted in the other PR

@ggerganov

Copy link
Copy Markdown
Member

This was supposed on top of the #25860 but that didn't happen.

Shouldn't you stack it on top of #26490, instead of #25860?

@am17an

am17an commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Sorry it is on top of the that
image

@ryan5rdx

ryan5rdx commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

@ryan5rdx yeah I know it doesn't work with dspark, but what do you get when for just -sm layer as tg2048? It seems kinda low based on the numbers you posted in the other PR

command for reference, let me know happy to test an alternate config:

./bin/llama-server -m ~/Downloads/DeepSeek-V4-Flash-0731-MXFP4.gguf -c 1048576 --reasoning on --rpc 192.168.0.13:50052   -np 1 --reasoning-preserve -ub 2048 -b 4096  --no-mmap -ngl 999 -fa on  --fit off -ts 1,1 --host 0.0.0.0 -kvu -sm tensor

and yup just confirmed - with -sm layer numbers align with what I have in #26421:
tg2048: 23.05 t/s
pp2048: 270.3 t/s

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

@ryan5rdx try using two rpc servers, one on each machine and connect via

llama-server --rpc <ip_1, ip_2> --device RPC0, RPC1

@Kononnable

Copy link
Copy Markdown
Contributor

It might be worth to add -sm layer results to bench results table - just to have performance baseline in a single place.

I don't know if this would be observable with RDNA, but sometimes manually moving tensors instead of standard -sm layer can increase performance when tcp/ip is used (just a sidenote for 'base' performance).
-sm layer -ts 0,1 -ot 'blk\.[0-1][0-9]?\.ffn_(up|down|gate|gate_up)_(ch|)exps=RPC0[127.0.0.1:50052]'

@ggerganov

Copy link
Copy Markdown
Member

@am17an The shared commits in the 2 branches differ:

Likely you've made changes to the dsv4-sm-tensor branch after you created the rpc_tensor branch. That's why the PR stack does not work.

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Yeah I messed up, I think #26490 should be okay to merge though

@ggerganov

Copy link
Copy Markdown
Member

Yeah I messed up, I think #26490 should be okay to merge though

Don't we want to fix the DSpark support first?

@am17an

am17an commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor Author

The support is broken over RPC I think(i.e. this PR), not in general. But I can check

@ryan5rdx

ryan5rdx commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

@ryan5rdx try using two rpc servers, one on each machine and connect via

llama-server --rpc <ip_1, ip_2> --device RPC0, RPC1

Sorry - to clarify - is this RPC servers on two RDMA linked nodes, and then a llama-server instance on one of them(if so I suppose this will just connect to the localhost RPC server)? Sorry I've never run a llama-cli/server instance where it's not also doing compute ha

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

@ryan5rdx yes, the llama-server is just a client of these two. Basically we're trying to activate the RPC<>RPC all-reduce path rather than the one you probably got (Metal<>RPC)

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

It might be worth to add -sm layer results to bench results table - just to have performance baseline in a single place.

Currently on master -sm layer is crashing for some reason. /home/aman/experiments/llama.cpp/ggml/src/ggml-rpc/ggml-rpc.cpp:519: Remote RPC server crashed or returned malformed response, but last time I checked it was around ~400 PP and ~15 TG

@ryan5rdx

ryan5rdx commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

@ryan5rdx yes, the llama-server is just a client of these two. Basically we're trying to activate the RPC<>RPC all-reduce path rather than the one you probably got (Metal<>RPC)

-sm tensor hangs after tensors are sharded to RPC nodes with any topology >1 RPC backend.

Tested with Qwen327B, 0.6B and ds4. -sm layer works on both qwen models with 2 RPC backends, ds4 breaks even on layer, but seems unrelated.

no obvious error on rpc servers or llama-server, I see the tensors copy on the RPC nodes then nothing.

final client logs, llama server web UI never comes up(hangs here):

0.00.236.455 I print_info: EOT token             = 151645 '<|im_end|>'
0.00.236.455 I print_info: PAD token             = 151654 '<|vision_pad|>'
0.00.236.455 I print_info: LF token              = 198 'Ċ'
0.00.236.456 I print_info: FIM PRE token         = 151659 '<|fim_prefix|>'
0.00.236.456 I print_info: FIM SUF token         = 151661 '<|fim_suffix|>'
0.00.236.456 I print_info: FIM MID token         = 151660 '<|fim_middle|>'
0.00.236.456 I print_info: FIM PAD token         = 151662 '<|fim_pad|>'
0.00.236.456 I print_info: FIM REP token         = 151663 '<|repo_name|>'
0.00.236.457 I print_info: FIM SEP token         = 151664 '<|file_sep|>'
0.00.236.457 I print_info: EOG token             = 128247 '</s>'
0.00.236.457 I print_info: EOG token             = 151643 '<|endoftext|>'
0.00.236.458 I print_info: EOG token             = 151645 '<|im_end|>'
0.00.236.458 I print_info: EOG token             = 151662 '<|fim_pad|>'
0.00.236.458 I print_info: EOG token             = 151663 '<|repo_name|>'
0.00.236.458 I print_info: EOG token             = 151664 '<|file_sep|>'
0.00.236.459 I print_info: max token length      = 256
0.00.236.459 I load_tensors: loading model tensors, this can take a while... (load_mode = none)
0.00.262.459 I load_tensors: offloading output layer to GPU
0.00.262.460 I load_tensors: offloading 27 repeating layers to GPU
0.00.262.461 I load_tensors: offloaded 29/29 layers to GPU
0.00.262.462 I load_tensors:          CPU model buffer size =   121.71 MiB
0.00.262.462 I load_tensors:       Meta() model buffer size =   251.44 MiB
0.00.734.830 I cmn  common_init_: added </s> logit bias = -inf
0.00.734.991 I cmn  common_init_: added <|endoftext|> logit bias = -inf
0.00.734.993 I cmn  common_init_: added <|im_end|> logit bias = -inf
0.00.734.993 I cmn  common_init_: added <|fim_pad|> logit bias = -inf
0.00.734.994 I cmn  common_init_: added <|repo_name|> logit bias = -inf
0.00.734.994 I cmn  common_init_: added <|file_sep|> logit bias = -inf
0.00.735.025 I llama_context: constructing llama_context
0.00.735.026 I llama_context: n_seq_max     = 1
0.00.735.027 I llama_context: n_ctx         = 20224
0.00.735.027 I llama_context: n_ctx_seq     = 20224
0.00.735.027 I llama_context: n_batch       = 4096
0.00.735.027 I llama_context: n_ubatch      = 2048
0.00.735.028 I llama_context: causal_attn   = 1
0.00.735.028 I llama_context: flash_attn    = enabled
0.00.735.028 I llama_context: kv_unified    = true
0.00.735.029 I llama_context: freq_base     = 1000000.0
0.00.735.029 I llama_context: freq_scale    = 1
0.00.735.030 I llama_context: n_rs_seq      = 0
0.00.735.030 I llama_context: n_outputs_max = 1
0.00.735.030 I llama_context: n_ctx_seq (20224) < n_ctx_train (40960) -- the full capacity of the model will not be utilized

command:

  ./bin/llama-server -m ~/Downloads/Qwen3-0.6B-UD-Q4_K_XL.gguf -c 20000 --reasoning on --rpc 192.168.0.13:50052,192.168.0.9:50052 --device RPC0,RPC1   -np 1 --reasoning-preserve -ub 2048 -b 4096  --no-mmap -ngl 999 -fa on  --fit off -ts 1,1 --host 0.0.0.0 -kvu -sm tensor -lv 4

@ryan5rdx

ryan5rdx commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

@ryan5rdx yes, the llama-server is just a client of these two. Basically we're trying to activate the RPC<>RPC all-reduce path rather than the one you probably got (Metal<>RPC)

for the all-reduce to work - don't RPC nodes need direct(RDMA) connections to each other in addition to (at least)TCP to the client?

with an A - B(just client) - C topology where A<>B and B<>C are RDMA links, but there is no A <> C link, can this work? (in MLX they achieve this with a mesh + rdma, but here RPC servers don't know about peers yet)

@rgerganov rgerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please create a mermaid sequence diagram (similar to this one ) which describe how peers communicate when -sm tensor is being used, I am still trying to understand the new flows being added and that would be very helpful. In fact, I think this should be part of our dev documentation (feel free to create an .md file) so we can maintain this in the long term.

Comment thread ggml/src/ggml-rpc/ggml-rpc.cpp
Comment thread ggml/src/ggml-rpc/ggml-rpc.cpp Outdated
return true;
}

// minor protocol version of each connected server, used to gate newer commands (comm collectives)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

no need to do this, we don't care about backward compatibility and we prefer to keep the code simple; just bump the version to 6.0.0 and expect all peers to be running this version

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

@ryan5rdx I run this using the directly the RDMA interfaces, for example in my case it is --rpc 10.10.20.1, 10.10.20.2 and you should see the RDMA being negotiated between the two peers. There is no mesh required if they can talk to each other

@ryan5rdx

ryan5rdx commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

@ryan5rdx I run this using the directly the RDMA interfaces, for example in my case it is --rpc 10.10.20.1, 10.10.20.2 and you should see the RDMA being negotiated between the two peers. There is no mesh required if they can talk to each other

in my case the RPC nodes cannot talk to each other via the --rpc passed IPs on the client (the node at 10.10.20.1 cannot reach the other rpc node via 10.10.20.2, it can however reach it via a different ip, if I were to connect them together via a TB cable, or I suppose via a completely different ip over Ethernet via tcp).

Possible I'm misunderstanding here - or maybe a mac difference because it only supports p2p (TB)rdma without routing vs the spark?

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

@ryan5rdx not sure, you can probably ask an LLM to debug?

@ryan5rdx

ryan5rdx commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

@ryan5rdx not sure, you can probably ask an LLM to debug?

Ok yup confirmed - so if I use GGML_RPC_NO_COMM=1 disable to worker <> worker comm, it works(even with RDMA, since it relays traffic over the client)

if I instead I use the LAN IPs - it still works becuase each worker can discover the other using the ip from the client, just super slow as expected (>>10s/token).

However - if we want this to work on metal, with worker<>worker RDMA we need a mesh and to allow for passing the TB IP of other workers to each rpc server(because peer addresses are not derivable from the addr we get from the client in a TB p2p network).

I'm not familiar with sparks but apparently it works because "all three nodes sit on one routable RDMA fabric"

@am17an

am17an commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

@rgerganov I added a mermaid diagram. Note that the Meta backend creates a lot of subgraphs so that's why graph caching is vital.

@rgerganov rgerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

you can improve the diagram by making a clear separation between messages sent to the standard RPC port and messages sent to the new "comm" port

return nullptr;
}
ggml_backend_rpc_context * rpc_ctx = (ggml_backend_rpc_context *) backends[i]->context;
// one rank per endpoint: a server processes its socket sequentially, so a second

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"one rank per endpoint" -- why having this limitation? i can have an endpoint with two devices which communicate very fast (because they are on the same physical host)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In that case you can just create two rpc servers, one for each device?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On local MoE models with -cmoe weights are copied to GPU during PP for faster compute. Utilizing separate processes per device would break this optimization.
The example is using local, but same optimization should be possible on rpc with multiple backends.

Utilizing same process can also be useful in general for AllReduce - 2 servers 2 GPU each could merge local results lowering the amount of network hops needed.

if (!parse_endpoint(ranks[0].endpoint, host0, port0)) {
return nullptr;
}
const uint32_t comm_port = (uint32_t) port0 + 1000;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

expect this port to be specified by the user via command line flag

Comment thread ggml/src/ggml-rpc/ggml-rpc.cpp
Comment thread ggml/src/ggml-rpc/ggml-rpc.cpp
}

static void * ggml_backend_rpc_comm_init(ggml_backend_t * backends, size_t n_backends) {
if (n_backends != 2 || std::getenv("GGML_RPC_NO_COMM") != nullptr) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this going to work with more than 2 backends or it will require major redesign?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's going to fall-back to meta backend's all-reduce, which is slow but works

@Geramy

Geramy commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

@am17an have you looked at implementing https://github.com/mk1-project/quickreduce/ QuickReduce seems pretty cool and it could probably be of huge benefit for RDMA performance.

@github-actions github-actions Bot added model Model specific CUDA Related to the CUDA backend labels Aug 14, 2026
@am17an
am17an requested a review from a team as a code owner August 30, 2026 08:27
@am17an

am17an commented Aug 30, 2026

Copy link
Copy Markdown
Contributor Author

it's rebased on latest master. The async events seem to have added a lot of latency in sync due to worker threads wake up. When using -sm tensor we're currently busy waiting on the worker thread. This will utilize the CPU at 100%, we can come with a better solution if required

@am17an
am17an force-pushed the rpc_tensor branch 2 times, most recently from a269ab9 to 28bd87b Compare August 30, 2026 16:03
HelloKS added a commit to HelloKS/llama.cpp that referenced this pull request Sep 2, 2026
@am17an

am17an commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Here are my MTP bench results on 2x sparks after #28387, #28390

  code_python        pred= 192 draft= 177 acc= 145 rate=0.819 tok/s=41.5
  code_cpp           pred= 192 draft= 218 acc= 136 rate=0.624 tok/s=36.1
  explain_concept    pred= 192 draft= 265 acc= 124 rate=0.468 tok/s=27.4
  summarize          pred=  39 draft=  48 acc=  27 rate=0.562 tok/s=27.9
  qa_factual         pred= 192 draft= 232 acc= 132 rate=0.569 tok/s=31.6
  translation        pred=  20 draft=  20 acc=  15 rate=0.750 tok/s=33.2
  creative_short     pred=  45 draft= 108 acc=  18 rate=0.167 tok/s=15.9
  stepwise_math      pred= 192 draft= 212 acc= 136 rate=0.641 tok/s=31.5
  long_code_review   pred= 192 draft= 291 acc= 117 rate=0.402 tok/s=24.3

Aggregate: {
  "n_requests": 9,
  "total_predicted": 1256,
  "total_draft": 1571,
  "total_draft_accepted": 850,
  "aggregate_accept_rate": 0.5411,
  "wall_s_total": 45.17
}

This is without using NCCL and any extra dsv4 specific optimizations (which in my tests bring the wall time < 40s)

@kh0pper

This comment was marked as low quality.

Comment on lines +2292 to +2298
// Aux graph contents are rewritten on every compute but are identical across calls while the subgraphs are reused,
// so they can get stable uids on rebuild. Only safe without a comm backend, where the fallback usage is deterministic.
if (backend_ctx->comm_ctx == nullptr) {
for (ggml_cgraph * cgraph_aux : backend_ctx->cgraphs_aux) {
cgraph_aux->uid = ggml_graph_next_uid();
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not safe to do in general because with the API a user can pass arbitrary graphs. The correct logic would I think not be easy to implement and introduce non-negligible complexity. What is the specific motivation for adding this change?

@am17an am17an Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry I just saw this. These changes are not required for this to work.

@ggerganov
ggerganov force-pushed the rpc_tensor branch 3 times, most recently from fc8b53b to 718ecbc Compare September 25, 2026 05:14
@am17an
am17an added this pull request to stack #29414 September 25, 2026 05:19
@Alexander2995

Copy link
Copy Markdown

Hi, I have a bit different graph cache implementation made as separate patch.

Could it be beneficial for this sm tensor feature branch?

Discussion #29253

Kononnable added a commit to Kononnable/llama.cpp that referenced this pull request Sep 25, 2026
Partial cherry-pick of the graph caching and the RPC_CMD_SET_TENSOR_2D /
RPC_CMD_GET_TENSOR_2D parts of PR ggml-org#26610. The -sm tensor comm support is
not included.

- server caches computed graphs per uid with bounded eviction
- client tracks cached graph uids and reuses them
- add RPC_CMD_SET_TENSOR_2D / RPC_CMD_GET_TENSOR_2D
- bump the RPC protocol to 7

Assisted-by: DeepSeek V4 Flash
@am17an

am17an commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

@rgerganov do you have time to review this PR? I'm pretty confident of the changes and I volunteer to maintain them

@ggerganov
ggerganov force-pushed the rpc_tensor branch 2 times, most recently from cf00d3b to d09d1c7 Compare September 29, 2026 05:28
Kononnable added a commit to Kononnable/llama.cpp that referenced this pull request Sep 30, 2026
Partial cherry-pick of the graph caching and the RPC_CMD_SET_TENSOR_2D /
RPC_CMD_GET_TENSOR_2D parts of PR ggml-org#26610. The -sm tensor comm support is
not included.

- server caches computed graphs per uid with bounded eviction
- client tracks cached graph uids and reuses them
- add RPC_CMD_SET_TENSOR_2D / RPC_CMD_GET_TENSOR_2D
- bump the RPC protocol to 7

Assisted-by: DeepSeek V4 Flash

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

AMD ZenDNN Issues related to the AMD ZenDNN backend Apple Metal https://en.wikipedia.org/wiki/Metal_(API) Ascend NPU issues specific to Ascend NPUs build Compilation issues conversion CUDA Related to the CUDA backend devops improvements to build systems and github actions documentation Improvements or additions to documentation examples ggml changes relating to the ggml tensor library for machine learning Hexagon IBM zDNN issues specific to IBM zDNN Accelerator model Model specific mtmd Related to multimodal functionality (video/image/audio) OpenCL Issues specific to the OpenCL backend OpenVINO server/ui server SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language testing Everything test related vendor Vulkan Issues specific to the Vulkan backend WebGPU

Projects

None yet

Development

Successfully merging this pull request may close these issues.