Skip to content

NPU: the build's own lax kernels are the default; Qwen3.6-35B-A3B through Lemonade - #72

Merged
bong-water-water-bong merged 1 commit into
mainfrom
npu-lax-default
Sep 25, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
npu-lax-default

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Checklist item 6 (the NPU from Lemonade), engine half.

Change: 1bit serve now finds the lax kernels without a flag. The lookup order is --npu-kernels, <dir>/npu/lax, $ONEBIT_NPU_LAX_KERNELS, then the build's own. -DONEBIT_NPU_LAX=ON builds them once into the build tree with scripts/build-lax.sh, from the open kernels in third_party/OpenFlowLM-Next. -DONEBIT_NPU_LAX_KERNELS=<dir> points at an existing build instead.

Measured on strixhalo:

  • build-lax.sh at the pin: lax_a insts.elf md5 426b8dd… and insts.bin md5 4f3c749…, both matching tests/golden/npu_lax.
  • 1bit serve -m <35B dir> with no kernel flag: ready in about 36 s, 16.9 tok/s, full ELFs.
  • Through Lemonade (fork onebit-npu, catalog entry Qwen3.6-35B-A3B-NPU-1bit → FastFlowLM/Qwen3.6-35B-A3B-NPU2): Lemonade hash-verified the checkpoint repo, loaded it in 30 s, and answered "Paris" at 16.0 tok/s. The log shows "full ELFs from …/lax/kernels"; no xclbin is read.

Docs: docs/npu-lax.md, and the checklist in docs/lemonade.md. Still open there: the small NPU models, reasoning_content on the NPU route, and our own HF repos.

🤖 Generated with Claude Code

…PU_LAX)

A Qwen3.6-35B-A3B directory needs no --npu-kernels: -DONEBIT_NPU_LAX builds the open kernels
with scripts/build-lax.sh into the build tree (once), or -DONEBIT_NPU_LAX_KERNELS names a build.
Lemonade hands 1bit serve only the model directory; through the fork's onebit recipe the 35B
loaded in 30 s on full ELFs and answered at 16.0 tok/s. Kernels at the pin match the goldens.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 93a8b02

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Incorrect Error Message

The error message in run_serve function has been updated to include a new suggestion about the build's own lax kernels, but it still refers to the old path <dir>/npu/lax instead of the new default search order. This could mislead users about where to place the kernels.

"Qwen3.6-35B-A3B needs the lax kernels: --npu-kernels, <dir>/npu/lax, "
"$ONEBIT_NPU_LAX_KERNELS or a -DONEBIT_NPU_LAX build)");
Default Kernel Path Not Fully Documented

The help message in run_unified function mentions the new default search order, but it doesn't clearly indicate that -DONEBIT_NPU_LAX=ON builds the kernels into the build tree. This could lead to confusion about how the default path is determined.

"              else <model dir>/npu/lax, else $ONEBIT_NPU_LAX_KERNELS, else this build's\n"
"              (-DONEBIT_NPU_LAX, docs/npu-lax.md)\n"

@bong-water-water-bong
bong-water-water-bong merged commit 2b2d7eb into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the npu-lax-default branch September 25, 2026 11:06
bong-water-water-bong pushed a commit that referenced this pull request Oct 2, 2026
…512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
…tic FA, attention sinks) (#295)

* Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss)

#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode)

#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant