Skip to content

rpc : reject invalid top-level graph nodes in graph_compute - #25670

Open
liminfei-amd wants to merge 1 commit into
ggml-org:masterfrom
liminfei-amd:fix/25299-rpc-null-node
Open

liminfei-amd wants to merge 1 commit into
ggml-org:masterfrom
liminfei-amd:fix/25299-rpc-null-node

Conversation

@liminfei-amd

@liminfei-amd liminfei-amd commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Overview

rpc_server::graph_compute() deserializes each top-level graph node id and
resolves it through create_node(). The old check only rejected a nullptr
result when the wire id was non-zero, treating a nullptr result for id 0
as "expected" and letting it into graph->nodes. id 0 is the wire sentinel
for "no tensor" and is valid for optional tensor sources (src[]/view_src),
but it is not a valid top-level graph node. The backend later dereferences
that node during graph planning, crashing the RPC server on a single crafted
GRAPH_COMPUTE message (device=0, n_nodes=1, node_id=0, n_tensors=0).

Fixes #25299.

Reject every failed top-level node reconstruction before calling
ggml_backend_graph_compute(), regardless of id.

Additional information

Verified with a localhost-only RPC client sending the malformed message above
against a CPU-only build (no ROCm/HIP): before the fix the server exits with
SIGSEGV; after the fix the malformed connection is closed and a subsequent
client can still connect and complete HELLO normally.

A regression test (test-rpc-invalid-graph-node) is included; it starts the
RPC server, sends the crafted message, confirms the connection is dropped, and
verifies the server stays alive for new connections.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - an AI assistant helped locate the root cause and drafted this fix; I reviewed the surrounding create_node/graph_compute code, reproduced the crash and the fix's effect with a local RPC client, and can explain every line of the change.

@github-actions github-actions Bot added examples ggml changes relating to the ggml tensor library for machine learning labels Jul 14, 2026
@ruixiang63 ruixiang63 closed this Jul 15, 2026
@ruixiang63 ruixiang63 reopened this Jul 15, 2026
@liminfei-amd
liminfei-amd marked this pull request as ready for review July 17, 2026 07:04
@liminfei-amd
liminfei-amd requested a review from a team as a code owner July 17, 2026 07:04
@liminfei-amd
liminfei-amd force-pushed the fix/25299-rpc-null-node branch from 91b62a7 to 3900550 Compare August 31, 2026 06:33
id 0 is a valid "no tensor" sentinel for optional sources, not for a
top-level graph node. Reject any failed node reconstruction regardless
of id, before running the graph.

Fixes ggml-org#25299
Reported-by: professor-moody

Assisted-by: GitHub Copilot
@liminfei-amd
liminfei-amd force-pushed the fix/25299-rpc-null-node branch from 3900550 to 181b199 Compare August 31, 2026 06:55
@liminfei-amd

Copy link
Copy Markdown
Contributor Author

@ruixiang63 you reopened this PR back in July — would you be able to take a look,
or point it at whoever should review it?

Rebased onto current master. The branch had gone stale and was conflicting. The
only conflict was in tools/rpc/CMakeLists.txt, which gained the
test-rpc-multi-server block in #26500. Both tests are kept. Mine is now gated
on LLAMA_BUILD_TESTS — it was on the generic BUILD_TESTING before, which this
project does not use, so it would not have registered in a normal test build.

I also removed a branch that my own change had made dead: once the guard returns
on nullptr, the following if (graph->nodes[i] != nullptr) is always true.
Removing it is safe because create_node() returns non-null only after
tensor_ptrs.find(id) succeeds, so the tensor_ptrs.at(id) below it cannot
throw.

The crash still reproduces on current master. Same test harness on both sides,
with only the guard change differing:

master this PR
test-rpc-invalid-graph-node FAILED passed
server after a malformed GRAPH_COMPUTE dies, SIGSEGV logs the rejection and keeps serving

test-rpc-multi-server also passes on this branch, which covers the normal graph
path through the code I touched. Both builds are CPU-only.

One thing I cannot do from here: the workflow runs on this head are all
action_required, so CI has not started. If you are able to approve them, that
would help.

@Jasmine-tim

Copy link
Copy Markdown

Independent reproduction — still reproducible at d1d3c33

Confirming this NULL-pointer dereference is still live on the current tree. I reproduced it independently on commit
d1d3c33 (build 0.4.1-dev, build 10987), both in-process and against the real
ggml-rpc-server over TCP.

Root cause

rpc_server::deserialize_tensor (ggml/src/ggml-rpc/ggml-rpc.cpp, ~line 1370) copies the attacker-controlled op,
op_params, and src[] verbatim into a real ggml_tensor with no check that an op which requires input sources actually
has them. graph_compute (ggml-rpc.cpp:1768) then hands the graph to the CPU backend, and
ggml_backend_cpu_device_supports_op (ggml-cpu.cpp:424+) dereferences src0->type / src1->type with no null check →
SIGSEGV.

A single 316-byte unauthenticated GRAPH_COMPUTE message (one tensor, op=GGML_OP_ADD, no src tensors) crashes the whole
server process.

Reproduction

Build (RPC enabled):

git checkout d1d3c33
cmake -B build-rpc -DCMAKE_BUILD_TYPE=Release -DGGML_RPC=ON
cmake --build build-rpc -j "$(nproc)"

A. In-process (ground truth) — drives the raw blob through the public rpc_server::graph_compute():

[08b] blob=s07_op_nosrc.bin size=316
[08b] calling graph_compute (expect SEGV if null-deref still live)...
Segmentation fault (core dumped)
EXIT=139

B. Remote / TCP — against the real ggml-rpc-server:

LD_LIBRARY_PATH=$PWD/build-rpc/bin ./build-rpc/bin/ggml-rpc-server --host 127.0.0.1 --port 19040

The client completes a HELLO handshake (cmd=14, 24-byte conn_caps), then sends GRAPH_COMPUTE (cmd=10, payload = the
316-byte blob):

HELLO rsp size=28
recv: b'\x00\x00\x00\x00\x00\x00\x00\x00' # 8-byte header flushed just before the crash

Server result — process dies with SIGSEGV, port freed:

[1]+ Segmentation fault (core dumped) ...ggml-rpc-server --host 127.0.0.1 --port 19040
PORT FREE (crashed)

Server log stops at Accepted client connection — the crash happens inside graph_compute before any response is
written.

Trigger payload

s07_op_nosrc.bin — sha256 = 086dd889364fe7b26c1a35709aaecf8ccc1f5109f2f8ddfbc4ae3b87dcc285e8
(sibling s11_badparams.bin, op=ADD + op_params=0x7FFFFFFF, SIGSEGVs identically.)

Environment

  • Linux x86_64, Ubuntu kernel 7.0.0-34-generic
  • g++ (Ubuntu 15.2.0-16ubuntu1) 15.2.0
  • commit d1d3c33

Impact

Remote unauthenticated denial-of-service: one crafted 316-byte message to the RPC server's TCP port (default 1234)
crashes the process, taking down every connected client. This maps to CVE-2026-78148 (sibling: CVE-2026-78147, issue
#25289); the fix in this PR (#25670) is not yet merged, so the tree remains exposed.

Note (scope): RPC is an opt-in feature — it requires a GGML_RPC=ON build and an operator explicitly starting
ggml-rpc-server; it is not exposed by llama-server by default.

This is an independent confirmation, not a new report — the sink is already tracked as CVE-2026-78148 / issue #25299.
Posting here to help the fix land.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

examples ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Remote unauthenticated NULL-pointer dereference in ggml-rpc graph_compute() via a node id of 0 (ggml_graph_plan/ggml_is_empty)

3 participants