Repository navigation
Pin ROCmFPX fc664ba: Hadamard rotation per tensor, any drafter type beside a rotated file - #262
Conversation
…eside a rotated file ROCmFPX#7 makes the Hadamard activation rotation per tensor: the loader flags the Q4_0 weights of a file stamped onebit.hadamard_q4_0 = 32 and only those get rotated activations. A plain Q4_0 DFlash2 drafter beside Qwen3.8-27B-H32 now accepts 216/263 drafts on code at 45.6 tok/s (Q8_0: 217/260, 44.4); with the old process-wide GGML_Q4_0_HADAMARD it accepted 0/337. serve no longer sets GGML_Q4_0_HADAMARD (the backend reads the stamp; a value in the environment still forces the old process-wide rotation), and the #259 guard that refused a Q4_0 drafter goes, with gguf_tensor_type_count. hadamard_route.sh: no process-wide variable, Q4_0 and Q8_0 drafters both accepted; long_route.sh follows. docs/lean.md, serve.md and tools/hadamard_q4_0.py describe the per-tensor rotation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
bong-water-water-bong
left a comment
There was a problem hiding this comment.
Review (PR-Agent duty). Looks good.
- The pin fc664ba is the tip of ROCmFPX
1bit/vulkan-rocmi4(checked with ls-remote) and is the merge of ROCmFPX#7. - Not setting
GGML_Q4_0_HADAMARDis the right default now that the backend reads the stamp, and keeping a value set in the environment as an override is sensible. - Dropping the #259 guard and
gguf_tensor_type_countis consistent with that: the per-tensor flag makes a plain Q4_0 drafter correct, and the PR has the numbers (216/263 accepted, 45.6 tok/s). - The tests now cover both drafter types, and checks pass.
One non-blocking note: serve and the ROCm llama-server have to come from the same build. A ROCm llama-server built before ROCmFPX#7 and paired with this serve gets no rotation at all, and its output is silently wrong rather than refused. Since the engine build produces both from the pinned submodule, that only happens with a mixed install. A version stamp check (or setting the env when the backend predates the flag) would close it someday.
Merging.
ROCmFPX#7 makes the Hadamard activation rotation per tensor: the loader flags
the Q4_0 weights of a file stamped onebit.hadamard_q4_0 = 32 and only those get
rotated activations. A plain Q4_0 DFlash2 drafter beside Qwen3.8-27B-H32 now
accepts 216/263 drafts on code at 45.6 tok/s (Q8_0: 217/260, 44.4); with the old
process-wide GGML_Q4_0_HADAMARD it accepted 0/337.
serve no longer sets GGML_Q4_0_HADAMARD (the backend reads the stamp; a value in
the environment still forces the old process-wide rotation), and the #259
guard that refused a Q4_0 drafter goes, with gguf_tensor_type_count.
hadamard_route.sh: no process-wide variable, Q4_0 and Q8_0 drafters both
accepted; long_route.sh follows. docs/lean.md, serve.md and
tools/hadamard_q4_0.py describe the per-tensor rotation.
🤖 Generated with Claude Code