Skip to content

NPU: Qwen3.6-35B-A3B on full ELFs, one runlist submit per token, served by 1bit serve - #52

Merged
bong-water-water-bong merged 7 commits into
mainfrom
npu/lax-35b
Sep 25, 2026
Merged

bong-water-water-bong merged 7 commits into
mainfrom
npu/lax-35b

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

The whole 40-layer Qwen3.6-35B-A3B MoE token on the NPU as one XRT runlist submit. That is 30 DeltaNet layers and 10 full-attention layers, with the routed experts retargeted and enqueued on the device. It runs on the open lax kernels, pinned in third_party/OpenFlowLM-Next (MIT). Experimental and outside 1bit serve; see docs/npu-lax.md.

The submodule pins 2490fa6 on branch 1bit/lax-35b of bong-water-water-bong/OpenFlowLM-Next, and a fresh clone fetches it.

What is here

  • scripts/build-lax.sh builds everything from the pinned source:

    • the two lax control texts (one xclbin);
    • the final norm and the lm_head;
    • the XRT harness.

    From a clean checkout this took 3 min 21 s, and the insts.bin files came out byte-identical to the builds the parity runs used.

  • tests/npu_lax_parity.sh checks 3 positions against the fp64 reference, then runs one chat turn.

  • docs/npu-lax.md covers build, check, chat, the numbers, the fixes, the GPU comparison, and what is left.

  • NOTICE entries for the open kernels (Cyrus Attoun / phlegm and AMD OpenFlowLM, MIT) and for vegah/LLMNpuTest (Apache-2.0). The copyright lines are taken from their own license files.

  • README status and a PORTING row (3x).

Results on Strix Halo

The model is model.q4nx, sha256 688f1e153d10…3cf8de.

check result
tests/npu_lax_parity.sh PASS; logits corr 0.999998 / 0.999993 / 0.999998; argmax 846 / 198 / 3710, equal to the reference; "The capital of France is Paris."
chat, 128 tokens 16.1 tok/s end to end; load 12.1 s (weights packed at startup, no pre-packed files)

How it got to 16.1 tok/s

step 40 layers tok/s
first correct decode 75.8 ms 11.1
full-attention trim 67.5 ms 12.3
DeltaNet schedule 53.5 ms 14.7
router 47.7–48.1 ms 16.1

C++ driver (aae6a58): 1bit npu-lax, no Python

  • Packer: re-implemented in C++ in npu/lax_pack, written from the recipe with no OpenFlowLM code copied. All 83 buffers are byte-identical to the Python packer, verified against the sha256 table in tests/golden/npu_lax/sha256.tsv.
  • Decoder: in npu/lax, with a swappable kernel set behind npu/lax_kernels.h, so a full-ELF set can drop in later.
  • Command: 1bit npu-lax for chat, with a --parity mode.
  • Results: parity PASS (corr 0.999998 / 0.999993 / 0.999998, argmax 846 / 198 / 3710), and the chat answers "The capital of France is Paris."
C++ Python
load, process start to ready 3.8 s 11.7 s
prompt 20.7 tok/s 18.7 tok/s
decode 16.4 tok/s 16.2 tok/s

Tests:

  • npu_lax_host runs in CI, with no model.
  • On Strix Halo: npu_lax_pack_model, npu_lax_patches and npu_lax_e2e (tests/npu_lax_cpp.sh, 12.6 s).

Full ELFs and 1bit serve (3c3e797)

1bit serve -m <35B dir> --device npu answers "The capital of France is Paris." on full ELFs; strace shows 0 xclbin opens.

  • Why the ELF path used to fault: mlir-aie folds +0x80000000 into buffer args 5 and up, aiebu moves that into the relocation addends, and XRT adds it again. npu/lax_elf unfolds it and assembles all seven ELFs. They are byte-identical to the harness's working set.
  • Per-token init: the init ELF heads every token's runlist, because a full-ELF context loses its array configuration after the ln or lm_head context runs.
  • Per-position configs: one per position, carrying no PDI, patched at the relocation sites found by npu/lax_stream::elf_position_sites.
  • Overlap: the next position is prepared while the device runs the current one.
  • Classic path: still available as --transport classic / --npu-transport classic.
full ELF classic
per token (--bench 256) 59.7–60.8 ms 61.2–62.1 ms
chat decode 16.3–16.5 tok/s 16.0–16.1 tok/s
load 3.6–3.8 s 3.6 s

Tests: 10 ctests pass on Strix Halo. npu_lax_host runs in CI with 15 new ELF-site checks. On hardware: npu_lax_elf (seven ELFs byte-for-byte), npu_lax_e2e on both transports (logits byte-identical between them), and npu_lax_serve (Paris; a prompt extending a cached turn answers Berlin; streaming; 0 xclbin). The fast-lane serve_e2e still passes.

Chat follow-ups restore a state snapshot (70011b8)

The DeltaNet state of the 30 linear layers (70 MB) is snapshotted at two points: before the prompt's last token, and before its last <|im_start|>. A follow-up restores the longest snapshot whose tokens are a prefix of the new prompt, then feeds only the remaining tokens. KV rows past that point are overwritten before anything reads them.

Test: logits after a restore are bit-identical to feeding the same prefix from scratch (npu_lax_snapshot).

3-turn chat through 1bit serve:

turn prompt tokens fed after TTFT before TTFT after
2 267 77 12.83 s 3.78 s
3 322 62 15.48 s 3.09 s
  • Cost: save 4.7 ms, restore 5 ms; 4 snapshots take 281 MB.
  • Flags: --npu-snapshots N; usage.prompt_tokens_details.cached_tokens counts reuse from snapshots.

What is left

  • Chat cache: the previous answer is still fed again, because the template re-renders it without its think block. Snapshots are in memory only.
  • Prefill: batched prompt processing.
  • Placement: the GPU (Vulkan) is faster for this model (Q8_0: 53 tok/s decode). The NPU's role is a concurrent second stream.

🤖 Generated with Claude Code

…mental)

The whole 40-layer MoE token (30 DeltaNet + 10 full-attention layers, routed
experts retargeted and enqueued on the device) as ONE XRT runlist submit, from
the open kernels pinned in third_party/OpenFlowLM-Next (MIT; branch
1bit/lax-35b @ 2490fa6, our fixes and optimisations of the lax design).

- scripts/build-lax.sh: the lax kernels (both control texts, one xclbin), the
  final norm and lm_head, and the XRT harness, from the pinned source. From a
  clean checkout the insts.bin are byte-identical to the verified builds.
- tests/npu_lax_parity.sh: 3 positions vs the fp64 reference (make_decode.py
  --requant), then one chat turn. On Strix Halo: corr 0.999998 / 0.999993 /
  0.999998, argmax 846 / 198 / 3710 = reference; "The capital of France is
  Paris."; 29 s.
- docs/npu-lax.md: build, check, chat; 16.1 tok/s (model.q4nx sha256
  688f1e15...3cf8de), the steps from 11.1 tok/s, the three faults that made it
  correct, the GPU comparison, and what is left (full ELFs + a C++ driver so
  1bit serve can route here; batched prefill).
- NOTICE: the open kernels (Cyrus Attoun / phlegm, AMD OpenFlowLM; MIT) and
  vegah/LLMNpuTest (Apache-2.0), from their own license files.

Experimental and outside 1bit serve: it runs on the classic xclbin path with a
Python driver, not yet the full-ELF C++ path rule 4 asks of served kernels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 3b9fbfe

… Python

The host side of the 35B MoE decode (docs/npu-lax.md), ported from the pinned
open kernels' Python (third_party/OpenFlowLM-Next @ 2490fa6, MIT) into the engine.
The layouts are re-implemented from their specification, not copied.

- npu/lax_pack.{h,cpp}: the packer. It builds the 40 layer pools (512 MiB) and
  consts, the q8 lm head pool, the final norm and the ptab from model.q4nx, with
  every projection at q4_1 (--requant). q8 is re-quantized as the reference's
  NumPy does it: no FMA, and a tied min or max goes to the later element.
- npu/lax_stream.{h,cpp}: the full-attention position patches (KV window
  length and offset, row drain, ptab record) and the cfg words (pool address
  + 0x80000000, the 8 column MM2S queues).
- npu/lax.{h,cpp}, npu/lax_kernels.h: the XRT decoder. It packs in parallel
  straight into the BOs, then runs 40 layers as one runlist, then ln and the lm
  head. Kernel loading sits behind KernelSet, so a full-ELF set can replace the
  classic xclbin + insts.bin one.
- app/npu_lax.{h,cpp}: `1bit npu-lax`, greedy chat (Qwen3.6 template, thinking
  off, one session across turns) and --parity (xres<t>.bin at position t,
  compare_decode.py's metric).

Tests, on Strix Halo (model.q4nx sha256 688f1e15...3cf8de, kernels from
scripts/build-lax.sh at the pin, lax_a insts.bin md5 4f3c749d):
- npu_lax_host (CI): requant, signed-nibble transcode, every permutation,
  ptab, patches and cfg words, against hashes of the Python output.
- npu_lax_pack_model: all 83 buffers byte-identical to the Python packer
  (tests/golden/npu_lax/sha256.tsv), 17 s.
- npu_lax_patches: lax_a patches = the reference harness's table.
- npu_lax_e2e (tests/npu_lax_cpp.sh): parity corr 0.999998 / 0.999993 /
  0.999998, argmax 846 / 198 / 3710 (same as the Python driver), then the chat
  answers "The capital of France is Paris."; 12.6 s.

Speed, back to back with lax_chat.py on the same prompt (128 tokens, same
text): ready 3.8 s vs 11.7 s, prompt 20.7 vs 18.7 tok/s, decode 16.4 vs
16.2 tok/s (bound by the device).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ve --device npu`

The lax kernels now load as full ELFs (PDI + control code, assembled in memory
by npu/full_elf; no xclbin opened), the default kernel set behind
npu/lax_kernels.h; the classic xclbin path stays selectable (--transport
classic) for A/B.

- npu/lax_elf.{h,cpp}: the seven ELFs (laxinit, lxf, axf<pos>, ln, lm) from the
  build's insts.elf + main.pdi; unfold() clears mlir-aie's +0x80000000 fold on
  args >= 5 (XRT's ELF patcher adds it itself) and refreshes the UID note.
  Byte-identical to the ELFs the reference harness ran (lax-elf 65ad44d).
- npu/lax_elf_kernels.cpp: layer context from laxinit, lxf and one PDI-less
  axf<pos> config per position reached; laxinit heads every token's runlist
  (ln / lm_head reconfigure the array in their own contexts).
- npu/lax_stream: elf_position_sites derives the window length word and, for
  the drain, record and window offset, the DDR_PATCH word plus the addend of
  the one relocation on its BD (3222; 2762/68; 2984/74; 3240/78 on the pin).
- npu/lax.cpp: two runlist slots, the next position prepared while the device
  runs the current one; Decoder::reset; per-run timings; lax::Session
  (greedy generation, cache continued only on an exact token prefix).
- `1bit serve -m <35B dir> --device npu` routes to it through unified
  (--npu-kernels / <dir>/npu/lax / $ONEBIT_NPU_LAX_KERNELS); usage reports
  prompt_tokens_details.cached_tokens.
- scripts/build-lax.sh exports insts.elf + main.pdi for all four kinds.
- Tests: synthetic-ELF site tests (CI), npu_lax_elf (sites + golden ELFs),
  npu_lax_e2e (ELF) / npu_lax_e2e_classic (parity + reset), npu_lax_serve.

Strix Halo, model.q4nx 688f1e15...3cf8de, build/1bit from this branch:
parity corr 0.999998/0.999993/0.999998, argmax 846/198/3710, logits
bit-identical to classic; bench 59.7-60.8 vs 61.2-62.1 ms/token; chat decode
16.3/16.5 vs 16.0/16.1 tok/s; serve answers "The capital of France is Paris."

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong bong-water-water-bong changed the title NPU: Qwen3.6-35B-A3B lax decode, one runlist submit per token (experimental) NPU: Qwen3.6-35B-A3B on full ELFs, one runlist submit per token, served by 1bit serve Sep 25, 2026
bong-water-water-bong and others added 4 commits September 24, 2026 21:42
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… re-running the history

Qwen3.6's chat template re-renders earlier answers without their think block, so a
chat client's follow-up never extended the tokens on the device and started over at
position 0. The KV rows can be overwritten from any position (a position reads rows
0..pos-1 only); the DeltaNet state cannot rewind, so lax::Session now keeps snapshots of
it (30 x 2342912 B = 70 MB, Decoder::save_state / load_state) in host memory, taken
before each prompt's last token and before its last <|im_start|>, LRU-capped. A request
starts from the live cache, the longest snapshot that is a proper prefix of its prompt,
or position 0, whichever feeds the fewest tokens.

- npu/lax_turns.{h,cpp}: the XRT-free bookkeeping (store, plan, snapshot points), with
  CI host tests in npu_lax_host including a three-turn chat on a stand-in decoder.
- `1bit npu-lax --snapshot-check` (ctest npu_lax_snapshot): a snapshot restored after a
  different continuation gives logits bit-identical to the prefix fed from scratch (28/28
  positions, cmp-identical dumps), and a Session follow-up the same tokens.
- `1bit serve --npu-snapshots N` (default 4); cached_tokens counts snapshot reuse,
  timings.cache names the source; streams honour stream_options.include_usage.
- tests/npu_lax_turns.py: the three-turn chat measurement. Turn 2/3 time to first token
  12.83 s -> 3.78 s and 15.48 s -> 3.09 s (267 -> 77 and 322 -> 62 tokens fed).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong marked this pull request as ready for review September 25, 2026 02:54
@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 25, 2026 02:54
@bong-water-water-bong
bong-water-water-bong merged commit 940b4a8 into main Sep 25, 2026
4 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the npu/lax-35b branch September 25, 2026 02:56
bong-water-water-bong added a commit that referenced this pull request Oct 1, 2026
…ma.cpp #52) (#257)

* Pin llama.cpp cde002d: GDN snapshot decay ratios in log space, fixes the MTP-rollback NaN (llama.cpp #52)

Qwen3.8-27B with --mtp on HRX0 aborted with NaN logits on some prompts
(Q8_0, UD-Q5_K_XL, UD-Q6_K files). The same bug caused an AMDGPU memory fault
on a repeated request without MTP. Both are fixed by the snapshot publish path
forming c_t/c_s as exp(log c_t - log c_s). Registry regenerated (pin line);
docs/hrx.md notes the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs/hrx.md: label the MTP-rollback fix speeds as balanced power mode

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant