Repository navigation
NPU: the build's own lax kernels are the default; Qwen3.6-35B-A3B through Lemonade - #72
Merged
Merged
Conversation
…PU_LAX) A Qwen3.6-35B-A3B directory needs no --npu-kernels: -DONEBIT_NPU_LAX builds the open kernels with scripts/build-lax.sh into the build tree (once), or -DONEBIT_NPU_LAX_KERNELS names a build. Lemonade hands 1bit serve only the model directory; through the fork's onebit recipe the 35B loaded in 30 s on full ELFs and answered at 16.0 tok/s. Kernels at the pin match the goldens. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
bong-water-water-bong
pushed a commit
that referenced
this pull request
Oct 2, 2026
…512 tokens, deterministic flash attention llama.cpp fork since cebcd70 (balanced mode, figures from each PR): - #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX; #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000, tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029. - #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156. - #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959). - #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files. - #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B. Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
added a commit
that referenced
this pull request
Oct 2, 2026
…tic FA, attention sinks) (#295) * Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention llama.cpp fork since cebcd70 (balanced mode, figures from each PR): - #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX; #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000, tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029. - #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156. - #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959). - #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files. - #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B. Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss) #68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output. Docs: docs/hrx.md "Our patches". Registry pin updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode) #74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now matches the CPU. Found by the release format matrix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
This was referenced Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Checklist item 6 (the NPU from Lemonade), engine half.
Change:
1bit servenow finds the lax kernels without a flag. The lookup order is--npu-kernels,<dir>/npu/lax,$ONEBIT_NPU_LAX_KERNELS, then the build's own.-DONEBIT_NPU_LAX=ONbuilds them once into the build tree withscripts/build-lax.sh, from the open kernels inthird_party/OpenFlowLM-Next.-DONEBIT_NPU_LAX_KERNELS=<dir>points at an existing build instead.Measured on strixhalo:
build-lax.shat the pin:lax_ainsts.elfmd5426b8dd…andinsts.binmd54f3c749…, both matchingtests/golden/npu_lax.1bit serve -m <35B dir>with no kernel flag: ready in about 36 s, 16.9 tok/s, full ELFs.onebit-npu, catalog entryQwen3.6-35B-A3B-NPU-1bit→FastFlowLM/Qwen3.6-35B-A3B-NPU2): Lemonade hash-verified the checkpoint repo, loaded it in 30 s, and answered "Paris" at 16.0 tok/s. The log shows "full ELFs from …/lax/kernels"; no xclbin is read.Docs:
docs/npu-lax.md, and the checklist indocs/lemonade.md. Still open there: the small NPU models,reasoning_contenton the NPU route, and our own HF repos.🤖 Generated with Claude Code