Skip to content

1bit serve --prefill-device hrx: prompt prefill on HRX, decode on Vulkan, one shared KV cache (zero copy) - #30

Merged
bong-water-water-bong merged 1 commit into
mainfrom
hrx-prefill-split
Sep 24, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
hrx-prefill-split

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

HRX prefills faster than Vulkan, Vulkan decodes faster, same GPU. 1bit serve --device vulkan --prefill-device hrx uses both, with no copies: HRX owns the KV cache in its device memory and exports it as dma-bufs (HSA); Vulkan maps the same memory. Only the KV cell metadata (positions, sequences; a few KB) moves between the two contexts.

llama.cpp (1bit/hrx-vulkan-patched 267d864, on top of #29's commits; our code in our own files under the Apache notice, hooks only in upstream files):

  • ggml-hrx ggml-hrx-dmabuf.cpp (export), ggml-vulkan ggml-vulkan-dmabuf.inc (import, VK_EXT_external_memory_dma_buf)
  • llama llama-kv-share.cpp: llama_kv_share_next / llama_kv_share_from / llama_kv_cells_copy; fixed 4 KiB layout in <= 1 GiB chunks with a layout hash per chunk
  • llama-server server-prefill.cpp: ONEBIT_PREFILL_DEVICE, prompt prefix in whole ubatches (HRX prefill halves on a partial one)

Engine: --prefill-device hrx and --prefill-min-tokens N in 1bit serve; ctest serve_e2e_vulkan_prefill_hrx; docs.

Measured on Strix Halo, llama-server on Vulkan0:

Model Prompt Prefill, Vulkan alone -> split Request
Qwen2.5-7B Q4_K_M 8192 7347 -> 4646 ms 10376 -> 7668 ms (-26%)
Qwen2.5-7B Q4_K_M 2048 1541 -> 1078 ms 4314 -> 3890 ms (-10%)
Qwen3-0.6B Q4_K_M 8192 1289 -> 1089 ms 2267 -> 2032 ms (-10%)
  • Decode on the shared KV: within 3% of Vulkan with its own KV (a host-memory design cost 2x and was dropped).
  • Teacher-forced KL against Vulkan alone: mean 0.0002-0.0006 nats, max 0.007, top token 64-65/65 (Q4_K_M vs BF16 is 0.063).
  • tests/serve_e2e.sh passes for vulkan, hrx and vulkan --prefill-device hrx, at the model's full 40960-token context and at 16384. (The full-context case found a bug: one shared buffer past ~4 GiB read garbage in Vulkan, hence the 1 GiB chunks.)

🤖 Generated with Claude Code

…kan, one shared KV cache

- third_party/llama.cpp -> 1bit/hrx-vulkan-patched 267d864: adds zero-copy KV
  sharing (HRX exports its KV memory as dma-bufs, Vulkan maps them; only the KV
  cell metadata moves) and ONEBIT_PREFILL_DEVICE in llama-server, on top of the
  IQ3_XXS / honest-claims commits.
- app/serve.cpp: --prefill-device hrx (with --device vulkan) runs the HRX build's
  llama-server on Vulkan0 with -fa on and ONEBIT_PREFILL_DEVICE=HRX0;
  --prefill-min-tokens overrides the 1024-token threshold.
- ctest serve_e2e_vulkan_prefill_hrx; docs/hrx.md "Prefill on HRX, decode on
  Vulkan", docs/serve.md.

Strix Halo, Qwen2.5-7B Q4_K_M: 8192-token prompt, request 10376 -> 7668 ms
(-26%); 2048, 4314 -> 3890 ms. Teacher-forced KL against Vulkan alone: mean
0.0003-0.0006, top token 64-65/65.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit c55cc77 into main Sep 24, 2026
1 check passed
@bong-water-water-bong
bong-water-water-bong deleted the hrx-prefill-split branch September 24, 2026 08:03
bong-water-water-bong pushed a commit that referenced this pull request Sep 27, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 27, 2026
…orised (engine#124) (#180)

* Pin llama.cpp a34a6b7: HRX decode-split multipass output pass is vectorised (engine#124)

Brings 1bit-MONSTER/llama.cpp#30. The multipass reducer's output pass gave each
workitem one channel and walked every KV block serially, recomputing
expf(partial_max - maximum) per (block, element); only half the workgroup is live
at value_head_size=128, which was the whole residual >2048 decode gap.

The per-block scale is now computed once in the lane-strided sum pass into a
per-row LDS stage, and one vectorised all-rows pass (4 channels per workitem) does
the output. Per-channel block order is unchanged, so the reduce is bit-identical
(1.24 GB of decode-path logits compare equal to the previous pin).

d2100 +15.4%, d3000 +19.8%, d4800 +22.0%; d2100/d2000 0.881 -> 0.976, so the
boundary cliff is gone. Nothing changes at or below capacity 2048. Details and
repro in benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md.

* Pin llama.cpp fcd83eb (llama.cpp #30 merged; same tree as a34a6b7)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: agent-872bf5 <agent-872bf5@localhost>
Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant