Repository navigation
1bit serve --prefill-device hrx: prompt prefill on HRX, decode on Vulkan, one shared KV cache (zero copy) - #30
Merged
Conversation
…kan, one shared KV cache - third_party/llama.cpp -> 1bit/hrx-vulkan-patched 267d864: adds zero-copy KV sharing (HRX exports its KV memory as dma-bufs, Vulkan maps them; only the KV cell metadata moves) and ONEBIT_PREFILL_DEVICE in llama-server, on top of the IQ3_XXS / honest-claims commits. - app/serve.cpp: --prefill-device hrx (with --device vulkan) runs the HRX build's llama-server on Vulkan0 with -fa on and ONEBIT_PREFILL_DEVICE=HRX0; --prefill-min-tokens overrides the 1024-token threshold. - ctest serve_e2e_vulkan_prefill_hrx; docs/hrx.md "Prefill on HRX, decode on Vulkan", docs/serve.md. Strix Halo, Qwen2.5-7B Q4_K_M: 8192-token prompt, request 10376 -> 7668 ms (-26%); 2048, 4314 -> 3890 ms. Teacher-forced KL against Vulkan alone: mean 0.0003-0.0006, top token 64-65/65. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 24, 2026
bong-water-water-bong
pushed a commit
that referenced
this pull request
Sep 27, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
added a commit
that referenced
this pull request
Sep 27, 2026
…orised (engine#124) (#180) * Pin llama.cpp a34a6b7: HRX decode-split multipass output pass is vectorised (engine#124) Brings 1bit-MONSTER/llama.cpp#30. The multipass reducer's output pass gave each workitem one channel and walked every KV block serially, recomputing expf(partial_max - maximum) per (block, element); only half the workgroup is live at value_head_size=128, which was the whole residual >2048 decode gap. The per-block scale is now computed once in the lane-strided sum pass into a per-row LDS stage, and one vectorised all-rows pass (4 channels per workitem) does the output. Per-channel block order is unchanged, so the reduce is bit-identical (1.24 GB of decode-path logits compare equal to the previous pin). d2100 +15.4%, d3000 +19.8%, d4800 +22.0%; d2100/d2000 0.881 -> 0.976, so the boundary cliff is gone. Nothing changes at or below capacity 2048. Details and repro in benchmarks/NOTE-hrx-124-multipass-output-2026-09-27.md. * Pin llama.cpp fcd83eb (llama.cpp #30 merged; same tree as a34a6b7) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: agent-872bf5 <agent-872bf5@localhost> Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
HRX prefills faster than Vulkan, Vulkan decodes faster, same GPU.
1bit serve --device vulkan --prefill-device hrxuses both, with no copies: HRX owns the KV cache in its device memory and exports it as dma-bufs (HSA); Vulkan maps the same memory. Only the KV cell metadata (positions, sequences; a few KB) moves between the two contexts.llama.cpp (
1bit/hrx-vulkan-patched267d864, on top of #29's commits; our code in our own files under the Apache notice, hooks only in upstream files):ggml-hrx-dmabuf.cpp(export), ggml-vulkanggml-vulkan-dmabuf.inc(import,VK_EXT_external_memory_dma_buf)llama-kv-share.cpp:llama_kv_share_next/llama_kv_share_from/llama_kv_cells_copy; fixed 4 KiB layout in <= 1 GiB chunks with a layout hash per chunkserver-prefill.cpp:ONEBIT_PREFILL_DEVICE, prompt prefix in whole ubatches (HRX prefill halves on a partial one)Engine:
--prefill-device hrxand--prefill-min-tokens Nin1bit serve; ctestserve_e2e_vulkan_prefill_hrx; docs.Measured on Strix Halo, llama-server on Vulkan0:
tests/serve_e2e.shpasses for vulkan, hrx andvulkan --prefill-device hrx, at the model's full 40960-token context and at 16384. (The full-context case found a bug: one shared buffer past ~4 GiB read garbage in Vulkan, hence the 1 GiB chunks.)🤖 Generated with Claude Code