Repository navigation
NPU: Qwen3.6-35B-A3B on full ELFs, one runlist submit per token, served by 1bit serve - #52
Merged
Merged
Conversation
…mental) The whole 40-layer MoE token (30 DeltaNet + 10 full-attention layers, routed experts retargeted and enqueued on the device) as ONE XRT runlist submit, from the open kernels pinned in third_party/OpenFlowLM-Next (MIT; branch 1bit/lax-35b @ 2490fa6, our fixes and optimisations of the lax design). - scripts/build-lax.sh: the lax kernels (both control texts, one xclbin), the final norm and lm_head, and the XRT harness, from the pinned source. From a clean checkout the insts.bin are byte-identical to the verified builds. - tests/npu_lax_parity.sh: 3 positions vs the fp64 reference (make_decode.py --requant), then one chat turn. On Strix Halo: corr 0.999998 / 0.999993 / 0.999998, argmax 846 / 198 / 3710 = reference; "The capital of France is Paris."; 29 s. - docs/npu-lax.md: build, check, chat; 16.1 tok/s (model.q4nx sha256 688f1e15...3cf8de), the steps from 11.1 tok/s, the three faults that made it correct, the GPU comparison, and what is left (full ELFs + a C++ driver so 1bit serve can route here; batched prefill). - NOTICE: the open kernels (Cyrus Attoun / phlegm, AMD OpenFlowLM; MIT) and vegah/LLMNpuTest (Apache-2.0), from their own license files. Experimental and outside 1bit serve: it runs on the classic xclbin path with a Python driver, not yet the full-ELF C++ path rule 4 asks of served kernels. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
… Python
The host side of the 35B MoE decode (docs/npu-lax.md), ported from the pinned
open kernels' Python (third_party/OpenFlowLM-Next @ 2490fa6, MIT) into the engine.
The layouts are re-implemented from their specification, not copied.
- npu/lax_pack.{h,cpp}: the packer. It builds the 40 layer pools (512 MiB) and
consts, the q8 lm head pool, the final norm and the ptab from model.q4nx, with
every projection at q4_1 (--requant). q8 is re-quantized as the reference's
NumPy does it: no FMA, and a tied min or max goes to the later element.
- npu/lax_stream.{h,cpp}: the full-attention position patches (KV window
length and offset, row drain, ptab record) and the cfg words (pool address
+ 0x80000000, the 8 column MM2S queues).
- npu/lax.{h,cpp}, npu/lax_kernels.h: the XRT decoder. It packs in parallel
straight into the BOs, then runs 40 layers as one runlist, then ln and the lm
head. Kernel loading sits behind KernelSet, so a full-ELF set can replace the
classic xclbin + insts.bin one.
- app/npu_lax.{h,cpp}: `1bit npu-lax`, greedy chat (Qwen3.6 template, thinking
off, one session across turns) and --parity (xres<t>.bin at position t,
compare_decode.py's metric).
Tests, on Strix Halo (model.q4nx sha256 688f1e15...3cf8de, kernels from
scripts/build-lax.sh at the pin, lax_a insts.bin md5 4f3c749d):
- npu_lax_host (CI): requant, signed-nibble transcode, every permutation,
ptab, patches and cfg words, against hashes of the Python output.
- npu_lax_pack_model: all 83 buffers byte-identical to the Python packer
(tests/golden/npu_lax/sha256.tsv), 17 s.
- npu_lax_patches: lax_a patches = the reference harness's table.
- npu_lax_e2e (tests/npu_lax_cpp.sh): parity corr 0.999998 / 0.999993 /
0.999998, argmax 846 / 198 / 3710 (same as the Python driver), then the chat
answers "The capital of France is Paris."; 12.6 s.
Speed, back to back with lax_chat.py on the same prompt (128 tokens, same
text): ready 3.8 s vs 11.7 s, prompt 20.7 vs 18.7 tok/s, decode 16.4 vs
16.2 tok/s (bound by the device).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ve --device npu`
The lax kernels now load as full ELFs (PDI + control code, assembled in memory
by npu/full_elf; no xclbin opened), the default kernel set behind
npu/lax_kernels.h; the classic xclbin path stays selectable (--transport
classic) for A/B.
- npu/lax_elf.{h,cpp}: the seven ELFs (laxinit, lxf, axf<pos>, ln, lm) from the
build's insts.elf + main.pdi; unfold() clears mlir-aie's +0x80000000 fold on
args >= 5 (XRT's ELF patcher adds it itself) and refreshes the UID note.
Byte-identical to the ELFs the reference harness ran (lax-elf 65ad44d).
- npu/lax_elf_kernels.cpp: layer context from laxinit, lxf and one PDI-less
axf<pos> config per position reached; laxinit heads every token's runlist
(ln / lm_head reconfigure the array in their own contexts).
- npu/lax_stream: elf_position_sites derives the window length word and, for
the drain, record and window offset, the DDR_PATCH word plus the addend of
the one relocation on its BD (3222; 2762/68; 2984/74; 3240/78 on the pin).
- npu/lax.cpp: two runlist slots, the next position prepared while the device
runs the current one; Decoder::reset; per-run timings; lax::Session
(greedy generation, cache continued only on an exact token prefix).
- `1bit serve -m <35B dir> --device npu` routes to it through unified
(--npu-kernels / <dir>/npu/lax / $ONEBIT_NPU_LAX_KERNELS); usage reports
prompt_tokens_details.cached_tokens.
- scripts/build-lax.sh exports insts.elf + main.pdi for all four kinds.
- Tests: synthetic-ELF site tests (CI), npu_lax_elf (sites + golden ELFs),
npu_lax_e2e (ELF) / npu_lax_e2e_classic (parity + reset), npu_lax_serve.
Strix Halo, model.q4nx 688f1e15...3cf8de, build/1bit from this branch:
parity corr 0.999998/0.999993/0.999998, argmax 846/198/3710, logits
bit-identical to classic; bench 59.7-60.8 vs 61.2-62.1 ms/token; chat decode
16.3/16.5 vs 16.0/16.1 tok/s; serve answers "The capital of France is Paris."
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… re-running the history
Qwen3.6's chat template re-renders earlier answers without their think block, so a
chat client's follow-up never extended the tokens on the device and started over at
position 0. The KV rows can be overwritten from any position (a position reads rows
0..pos-1 only); the DeltaNet state cannot rewind, so lax::Session now keeps snapshots of
it (30 x 2342912 B = 70 MB, Decoder::save_state / load_state) in host memory, taken
before each prompt's last token and before its last <|im_start|>, LRU-capped. A request
starts from the live cache, the longest snapshot that is a proper prefix of its prompt,
or position 0, whichever feeds the fewest tokens.
- npu/lax_turns.{h,cpp}: the XRT-free bookkeeping (store, plan, snapshot points), with
CI host tests in npu_lax_host including a three-turn chat on a stand-in decoder.
- `1bit npu-lax --snapshot-check` (ctest npu_lax_snapshot): a snapshot restored after a
different continuation gives logits bit-identical to the prefix fed from scratch (28/28
positions, cmp-identical dumps), and a Session follow-up the same tokens.
- `1bit serve --npu-snapshots N` (default 4); cached_tokens counts snapshot reuse,
timings.cache names the source; streams honour stream_options.include_usage.
- tests/npu_lax_turns.py: the three-turn chat measurement. Turn 2/3 time to first token
12.83 s -> 3.78 s and 15.48 s -> 3.09 s (267 -> 77 and 322 -> 62 tokens fed).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
marked this pull request as ready for review
September 25, 2026 02:54
bong-water-water-bong
enabled auto-merge (squash)
September 25, 2026 02:54
This was referenced Sep 25, 2026
bong-water-water-bong
added a commit
that referenced
this pull request
Oct 1, 2026
…ma.cpp #52) (#257) * Pin llama.cpp cde002d: GDN snapshot decay ratios in log space, fixes the MTP-rollback NaN (llama.cpp #52) Qwen3.8-27B with --mtp on HRX0 aborted with NaN logits on some prompts (Q8_0, UD-Q5_K_XL, UD-Q6_K files). The same bug caused an AMDGPU memory fault on a repeated request without MTP. Both are fixed by the snapshot publish path forming c_t/c_s as exp(log c_t - log c_s). Registry regenerated (pin line); docs/hrx.md notes the fix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs/hrx.md: label the MTP-rollback fix speeds as balanced power mode Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The whole 40-layer Qwen3.6-35B-A3B MoE token on the NPU as one XRT runlist submit. That is 30 DeltaNet layers and 10 full-attention layers, with the routed experts retargeted and enqueued on the device. It runs on the open
laxkernels, pinned inthird_party/OpenFlowLM-Next(MIT). Experimental and outside1bit serve; seedocs/npu-lax.md.The submodule pins
2490fa6on branch1bit/lax-35bofbong-water-water-bong/OpenFlowLM-Next, and a fresh clone fetches it.What is here
scripts/build-lax.shbuilds everything from the pinned source:laxcontrol texts (one xclbin);From a clean checkout this took 3 min 21 s, and the
insts.binfiles came out byte-identical to the builds the parity runs used.tests/npu_lax_parity.shchecks 3 positions against the fp64 reference, then runs one chat turn.docs/npu-lax.mdcovers build, check, chat, the numbers, the fixes, the GPU comparison, and what is left.NOTICE entries for the open kernels (Cyrus Attoun / phlegm and AMD OpenFlowLM, MIT) and for vegah/LLMNpuTest (Apache-2.0). The copyright lines are taken from their own license files.
README status and a PORTING row (3x).
Results on Strix Halo
The model is
model.q4nx, sha256688f1e153d10…3cf8de.tests/npu_lax_parity.shHow it got to 16.1 tok/s
C++ driver (aae6a58):
1bit npu-lax, no Pythonnpu/lax_pack, written from the recipe with no OpenFlowLM code copied. All 83 buffers are byte-identical to the Python packer, verified against the sha256 table intests/golden/npu_lax/sha256.tsv.npu/lax, with a swappable kernel set behindnpu/lax_kernels.h, so a full-ELF set can drop in later.1bit npu-laxfor chat, with a--paritymode.Tests:
npu_lax_hostruns in CI, with no model.npu_lax_pack_model,npu_lax_patchesandnpu_lax_e2e(tests/npu_lax_cpp.sh, 12.6 s).Full ELFs and
1bit serve(3c3e797)1bit serve -m <35B dir> --device npuanswers "The capital of France is Paris." on full ELFs; strace shows 0 xclbin opens.npu/lax_elfunfolds it and assembles all seven ELFs. They are byte-identical to the harness's working set.npu/lax_stream::elf_position_sites.--transport classic/--npu-transport classic.--bench 256)Tests: 10 ctests pass on Strix Halo.
npu_lax_hostruns in CI with 15 new ELF-site checks. On hardware:npu_lax_elf(seven ELFs byte-for-byte),npu_lax_e2eon both transports (logits byte-identical between them), andnpu_lax_serve(Paris; a prompt extending a cached turn answers Berlin; streaming; 0 xclbin). The fast-laneserve_e2estill passes.Chat follow-ups restore a state snapshot (70011b8)
The DeltaNet state of the 30 linear layers (70 MB) is snapshotted at two points: before the prompt's last token, and before its last
<|im_start|>. A follow-up restores the longest snapshot whose tokens are a prefix of the new prompt, then feeds only the remaining tokens. KV rows past that point are overwritten before anything reads them.Test: logits after a restore are bit-identical to feeding the same prefix from scratch (
npu_lax_snapshot).3-turn chat through
1bit serve:--npu-snapshots N;usage.prompt_tokens_details.cached_tokenscounts reuse from snapshots.What is left
🤖 Generated with Claude Code