Skip to content

ggml-hrx: publish deferred host writebacks before a host-staging upload (engine #286) - #82

Merged
bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-host-writeback-flush
Oct 5, 2026
Merged

bong-water-water-bong merged 1 commit into
1bit/hrx-vulkan-patchedfrom
1bit/hrx-host-writeback-flush

Conversation

@bong-water-water-bong

Copy link
Copy Markdown

Root cause for 1bit-MONSTER/engine#286. PR #73 was the placement guard, which
hides the failing layout; this is the actual bug it was working around.

Root cause

An HRX command program whose output binding lives in host memory stages a
download buffer, enqueues the device copy, and then only queues the memcpy to
host memory:

  • download_prepared_host_staging() → add_host_writeback() records
    {host_destination, mapped_source, size} (runtime/command-program-executor.cpp);
  • the memcpy itself happens in GraphReplayStreamState::mark_stream_synchronized();
  • that is reached only from backend_synchronize() (ggml-hrx.cpp:509), cache
    trimming (runtime/graph-program-cache-limit.cpp:56) and the
    GGML_HRX_DIAGNOSTIC_GRAPH_SYNC path (command-program-executor.cpp:1909).

The split scheduler does not synchronize between splits. ggml-backend.cpp
computes every split with no sync:

if (!sched->callback_eval) {
    enum ggml_status ec = ggml_backend_graph_compute_async(split_backend, &split->graph);
    ...
}

and synchronizes a backend only when it inserts a cross-backend input copy,
i.e. only inside for (input_id = 0; input_id < split->n_inputs; input_id++).

So when an HRX split writes a tensor that lives in a host buffer and a later HRX
split reads it back, upload_prepared_host_staging() uploads the previous
contents of that memory.

Why this is engine ggml-org#286

With the expert MUL_MAT_IDs on the CPU, the gate ADD_ID output lands in a host
buffer, SWIGLU_OAI (a later HRX split) reads it across the CPU split, and the
host memcpy has not been applied yet. That is the missing-bias/garbage-value
signature (KLD 1.65).

It also explains every observation the issue already lists:

observation why
llama-eval-callback hides it that path takes the else branch and calls ggml_backend_synchronize(split_backend) after every node, which flushes the writeback
GGML_HRX_DEBUG_SERIAL_EXECUTION=1 still wrong DebugSerialExecutionTrace::sync() calls hrx_stream_synchronize() only; it never calls mark_stream_synchronized(), so the memcpy stays pending — syncing the stream is not enough
SWIGLU_OAI on the CPU is correct the tensor is then consumed by the CPU backend, the scheduler copies it and synchronizes the producing backend
default MUL_MAT_ID-on-HRX is correct the tensor stays device resident, no host staging

Change

Flush pending writebacks (stream synchronize, then publish) before an upload
reads host memory. The check is cheap: pending writebacks are normally empty, and
the flush clears them, so at most one extra synchronize happens per program that
reads host memory produced by an earlier replay.

This is deliberately the narrow fix — it makes the HRX backend stop assuming its
own earlier host writes are already published. The reason a tensor written by an
HRX node ends up in a host buffer at all (ggml-alloc in-place reuse ignoring
buffer_id, which let the gate ADD_ID output share the CPU MUL_MAT_ID buffer) is
a separate question and is not touched here.

Verification

  • Full semantic compile of the translation unit with the real HRX public
    headers: g++ -fsyntax-only -std=c++17 -I ggml/src -I ggml/include -I ggml/src/ggml-hrx -I <hrx>/libhrx/include ggml/src/ggml-hrx/runtime/command-program-executor.cpp → clean. The same command against a deliberately broken copy fails, so the check is meaningful.
  • Not verified on hardware: no Strix Halo box was available here, so the
    gpt-oss-20b KLD and greedy repro has not been re-run. That is the one thing
    this needs before merge:
llama-perplexity -c 512 -b 512 --chunks 8 -fa on   (KLD vs CPU)
GGML_HRX_DISABLE_DISPATCH=mul_mat_id.f32_f32_wmma

Expected: KLD ~0.03 and correct greedy text with the placement guard still in
place, and GGML_HRX_DEBUG_SERIAL_EXECUTION=1 no longer changing the result.

@github-actions github-actions Bot added the ggml label Oct 3, 2026
Root cause for 1bit-MONSTER/engine#286.

An HRX command program whose output binding lives in host memory stages a
download buffer, enqueues the device copy and then only *queues* the memcpy to
host memory: download_prepared_host_staging() calls add_host_writeback()
(runtime/command-program-executor.cpp), which records
{host_destination, mapped_source, size}. The memcpy itself happens in
GraphReplayStreamState::mark_stream_synchronized(), and that is reached only
from backend_synchronize() (ggml-hrx.cpp), cache trimming and the
GGML_HRX_DIAGNOSTIC_GRAPH_SYNC path.

The split scheduler does not synchronize between splits: the normal path
computes every split with ggml_backend_graph_compute_async() and no sync
(ggml/src/ggml-backend.cpp, "if (!sched->callback_eval)"). It synchronizes a
backend only when it inserts a cross-backend input copy, i.e. only inside
"if (split->n_inputs > 0)". So when an HRX split writes a tensor that lives in
a host buffer and a later HRX split reads it back, upload_prepared_host_staging()
uploads the previous contents of that memory.

That is the gpt-oss-20b failure in engine ggml-org#286: with the expert MUL_MAT_IDs on
the CPU, the gate ADD_ID output lands in a host buffer, SWIGLU_OAI (a later HRX
split) reads it across the CPU split, and the host memcpy has not been applied
yet.

It also explains the observations the issue lists:

* llama-eval-callback hides it, because that path takes the
  "else" branch and calls ggml_backend_synchronize(split_backend) after every
  node, which does flush the writeback.
* GGML_HRX_DEBUG_SERIAL_EXECUTION=1 does not help, because
  DebugSerialExecutionTrace::sync() calls hrx_stream_synchronize() only; it
  never calls mark_stream_synchronized(), so the pending memcpy stays pending.
* SWIGLU_OAI on the CPU is correct: the tensor is then consumed by the CPU
  backend, the scheduler copies it and synchronizes the producing backend.
* The default MUL_MAT_ID-on-HRX layout is correct because the tensor stays
  device resident and no host staging is involved.

Flush the pending writebacks (stream synchronize, then publish) before an
upload reads host memory. The check is cheap: pending writebacks are normally
empty and are cleared by the same flush.

Verified by a full semantic compile of the translation unit
(g++ -fsyntax-only, C++17, ggml + libhrx headers) with the same command run
against a deliberately broken copy to confirm the check is meaningful. Not
verified on hardware: no Strix Halo box was available, so the gpt-oss-20b
KLD/greedy repro from engine ggml-org#286 has not been re-run.
@bong-water-water-bong
bong-water-water-bong force-pushed the 1bit/hrx-host-writeback-flush branch from 09d922a to e0f637b Compare October 3, 2026 01:57
@bong-water-water-bong

Copy link
Copy Markdown
Author

Review (orchestrator, PR-review duty while coder-llm is down)

Read against f5b7f4a. The diagnosis holds:

  • add_host_writeback only queues the memcpy.
  • mark_stream_synchronized() is reached only from backend_synchronize, cache trim and the diagnostic sync.
  • The scheduler's compute_async loop doesn't synchronize between same-backend splits.

So an HRX split that uploads host memory written by an earlier HRX split's download reads stale bytes. That covers both upload paths: source_host_buffer (device copy from the host buffer) and upload_async. stream_synchronize_sleeping exists in runtime/hrx-sleeping-wait.h, and the per-loop check is cheap because the flush empties the list.

Other readers of host memory are already covered:

  • the CPU backend consuming the tensor: the scheduler syncs the input backend before the copy;
  • logits via tensor_get_async + synchronize.

So the narrow fix is enough for ggml-org#286.

Verdict: approve, pending the hardware run in the description (gpt-oss-20b KLD vs CPU with GGML_HRX_DISABLE_DISPATCH=mul_mat_id.f32_f32_wmma, and GGML_HRX_DEBUG_SERIAL_EXECUTION=1 no longer changing the result). strixhalo is reserved from Sat 10:00 to Sun 10:00 ADT for the release, so this is queued for the post-release merge window together with #75-#80.

@bong-water-water-bong

Copy link
Copy Markdown
Author

Verified on strixhalo (gfx1151, balanced power mode 85 W). Builds: base f5b7f4a and f5b7f4a + e0f637b, same tree and flags.

The base already has the placement guard (ba0be9a), which hides the failing layout. To reach the bug, I used a test-only build of both binaries in which an env var skips the moe_tail_claimable hook. This patch was not committed.

gpt-oss-20b MXFP4, llama-perplexity -c 512 -b 512 --chunks 8 -fa on, KLD vs CPU logits, GGML_HRX_DISABLE_DISPATCH=mul_mat_id.f32_f32_wmma:

config base this PR
guard bypassed, run 1 1.651073 (same top 11.86%) 0.029298 (same top 88.48%)
guard bypassed, run 2 1.651073 0.029298
guard bypassed + GGML_HRX_DEBUG_SERIAL_EXECUTION=1 1.651073 0.029298
guard on 0.029289 0.029289
default (MUL_MAT_ID on HRX) 0.028448 0.028448

Greedy /completion ("The history of the Roman Empire began", 32 tokens, 3 identical requests):

  • guard bypassed, base: " with the fall of the Roman Empire, and the fall of the fall, as the fall in the end? ..."
  • guard bypassed, this PR: " with the founding of the city of Rome in 753 BCE. According to legend, ..." This is the same text the default path gives.

Each config gave 3/3 identical top-5 logprobs, with no NaN.

Default path, llama-bench gpt-oss-20b, 3 interleaved rounds, medians:

pp512 tg128
base 1005.72 28.58
this PR 1008.32 28.27

Both are inside the run-to-run spread: tg samples 28.27 to 28.63 on both binaries.

@bong-water-water-bong
bong-water-water-bong merged commit a679f78 into 1bit/hrx-vulkan-patched Oct 5, 2026
10 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant