Repository navigation
npu: GGUF-on-NPU serving route (Qwen2.5-7B, MiniCPM4-8B, MiniCPM5-1B answer ' Paris' on the NPU) - #179
Conversation
1bit serve -m <model.gguf> --device npu now repacks the GGUF into a Q4NX model directory (scripts/repack_gguf.py: FLM_Q4NX_Converter + a config.json derived from the GGUF metadata) and serves it in process through the model-generic forward (npu/forward/q4nx_forward.cpp, ported from 1bit-MONSTER's npu-infer), behind the OpenAI endpoints (app/forward_serve.cpp). The existing Q4NX-model-directory route (unified.cpp) is unchanged. The forward decodes any arch whose full-ELF geometry is in the kernel set; the per-shape ELFs live in the model-generic geometry set. Verified on Strix Halo: Qwen2.5-7B answers 'The capital of France is' -> ' Paris.'.
The forward defaults tie_embeddings=1 when config.json omits tie_word_embeddings, so an untied model (separate output.weight) got the embedding table as its LM head and produced garbage. Derive it from the GGUF's tensor list (output.weight present = untied), and carry the tokenizer's bos/eos ids + the arch's activation through.
- repack_gguf.py now derives vocab_size from the Q4NX's own embed_tokens shape (the converter pads vocab rows to the 32-row tile: 73448 -> 73472) instead of the unpadded tokenizer vocab, and carries LongRoPE rope_factors_short.weight through config.json as rope_scaling.short_factor. - the model-generic forward reads that short_factor and multiplies the RoPE inv_freq by it (inv *= short_factor[i]), which MiniCPM4's LongRoPE requires. Without the padded vocab the lm_head dequant failed (73448 % 32 != 0); without the short_factor MiniCPM4's RoPE was off.
The <think>/</think> idiom is a Qwen3-lane detail; the model-generic route serves arbitrary archs whose tokenizers do not know those tokens, so the enable_thinking=false block corrupted the prompt (Qwen2.5-7B chat returned garbage instead of 'Paris').
The .gguf/npu route could not run at all: - repack_gguf.py defaulted the Q4NX converter to a dead path (~/1bit-MONSTER-iso-build/...), so every repack failed with 'converter not found'. Resolve $ONEBIT_Q4NX_CONVERTER, then a sibling checkout, then the 1bit-MONSTER checkout, and say how to set it on failure. - serve.cpp defaulted the repack python to a per-user venv path, and rejected a bare command name because fs::exists() was applied to it. Default to python3 and only stat real paths. - serve.cpp passed an empty --kernels unless ONEBIT_NPU_KERNELS was set, so the forward looked for /full_i8_*.elf. Resolve $ONEBIT_NPU_KERNELS -> <model dir>/npu -> the configured ONEBIT_NPU_KERNELS_DIR, and fail with an actionable message. CMakeLists gains that cache path and bakes it in. Verified: 1bit serve -m MiniCPM5-1B-Q4_K_M.gguf --device npu repacks, resolves the kernel directory with no env var, loads the four full ELFs and reports ready. The chat completion still hangs (the llama 'empty reply' item, task-3).
…ad tiles
Three fixes found by running the route end to end:
- llama-class models were decoding to special-token soup because the forward's
hardcoded ChatML prompt never carries the BOS that their chat template opens
with (MiniCPM5-1B: '{{- bos_token }}'). repack_gguf.py now derives
add_bos_token from tokenizer.chat_template and the forward prepends
bos_token_id for chat requests. MiniCPM5-1B goes from
'|end><|fim_middle|>' to coherent text.
- the repack script path was hardcoded to another checkout; bake
ONEBIT_NPU_REPACK_SCRIPT at configure time and use it as the default.
- run_lm_head re-transposed the whole NV x H head (a T-float strided gather,
~150M cache misses) on every token; build the [tile][H][T] layout once.
Measured: Qwen3-0.6B answers coherently through the forward (layer path is
sound) at ~19 s/token, and Qwen2.5-7B loads with all five full-ELF designs, so
the stale 'crashes at load' row no longer applies.
MiniCPM4-8B carries minicpm.embedding_scale (12.0), minicpm.residual_scale (scale_depth/sqrt(NL) = 0.2475 for 32 layers) and minicpm.logit_scale (16.0). The model-generic forward must apply the first two (the logit scale is a constant on the output and does not move the argmax). Without them the MiniCPM4-8B hidden state runs at the wrong magnitude and the logits explode to punctuation soup instead of ' Paris'. write_config now carries any of the three present in the GGUF under the generic names the forward reads.
…' Paris') Ports the three verified MiniCPM4-8B fixes from np-model-generic (edd6739d8) into the engine route's forward. Agent -9beb31 verified MiniCPM4-8B is SentencePiece and -b35e4d noted the serve criterion is additionally tokenizer-gated; this is the forward half. 1. LongRoPE sign: ggml applies the rope factor as rope_yarn(theta / ff, ...) (ggml-cpu/ops.cpp ggml_rope_cache_init), so the angle is DIVIDED by freq_factors. 599ac77 multiplied (inv *= short_factor), inflating the high dims up to 31x. Divide now; NPU_INFER_ROPE_FACTOR_MUL=1 keeps the A/B. 2. embedding_scale / residual_scale from config.json (1.0 elsewhere). MiniCPM4 scales the embedding by 12.0 and each residual branch by scale_depth/sqrt(NL) = 0.2475; without them the hidden state runs at the wrong magnitude and the logits explode (absmax 136, punctuation soup). Verified in np-model-generic: host step-0 argmax 11225 '▁Paris' for ids 1,1507,8107,1379,8360,1410, matching the llama.cpp oracle.
Ports 2928002a9 from np-model-generic. Reads minicpm.logit_scale (16.0) from config.json and divides the final logits by it; 1.0 elsewhere. The forward's logits were ~16x the llama.cpp oracle's (oracle top-1 prob 0.929 vs effectively one-hot here). Refitting a temperature against the oracle with the fix in place collapses it 0.0586 -> 0.9451, i.e. exactly the missing logit_scale, with own top-1 1.0000 -> 0.9449 and the slice-KL unchanged. The argmax does not move (constant positive divisor), which is why no argmax-level check could see this. Kept config-driven rather than baked into the converter so the Q4NX stays faithful to the GGUF.
MiniCPM4-8B decoded garbage for hours because the repack's config.json dropped embedding_scale / residual_scale / logit_scale, which the GGUF declares and llama.cpp applies -- and no argmax-level check could see it (the first two move the answer, the third does not move an argmax at all). This checks the class directly: every scale-like scalar the GGUF declares must survive into the repacked config.json. MiniCPM4-8B -> 3 declared, all present PASS MiniCPM5-1B -> 0 declared PASS (so no hidden config scale) Qwen2.5-7B -> 0 declared PASS negative control (scales stripped) -> FAIL (3 problem(s)), exit 1 Idea from @agent-b35e4d: cmp the repacked config.json against a reference dir's / the GGUF's own scale keys, instant and oracle-free, as a better guard than a per-model oracle probe.
@agent-b35e4d's point: the failure mode is silent, so a check nobody runs is weaker than the bug it targets. repack_gguf.py now runs the same declared-vs- carried scan inline after write_config and returns 1 if the GGUF declares a scale-like key (embedding_scale / residual_scale / logit_scale / scale_emb / scale_depth / dim_model_base / softcaps) that config.json drops. Verified: MiniCPM4-8B repack prints '[OK] every scale-like key the GGUF declares is carried in config.json' then 'repack complete'; the standalone scripts/check_repack_config.py agrees (PASS), and its negative control (those keys stripped from a copy) still FAILs with exit 1.
@agent-b35e4d's shape, adopted: FAIL on the known-critical list (the factors we have proven the forward must apply -- MiniCPM4-8B's three cost hours), WARN but do not fail on any other scale-ish key the GGUF declares and the config drops. The next architecture's factor will have a name nobody listed, so a silent skip would repeat this bug; but some scale-ish keys are legitimately not carried -- deepseek2.expert_weights_scale (1.8) is a routing scale, not a logit scale -- so a blanket rule would false-fail correct repacks. Verified on the warn tier: GLM-4.7-Flash warns on expert_weights_scale and still PASSes. check_repack_config.py is refactored into a reusable check(gguf, dir) used by both the CLI and the inline repack call, so the two cannot drift. docs/npu.md 'From a GGUF' now states the behaviour change (rc=1 on a dropped critical key), the known-critical list, why an argmax-level check cannot see this class, the warn tier, the negative control, and the provenance stamp still owed -- so an unexplained rc=1 is documented where the repack is documented rather than left to a resume.
@agent-b35e4d's point, and it was the guard's real hole: repack_gguf() returned a cached dir on 'model.q4nx exists' alone, so the repack -- and therefore the declared-scale guard -- was bypassed on every subsequent serve. A dir built before the MiniCPM4 scale fix would have been served silently forever, and the guard would only ever have protected freshly built artifacts. A cache hit now runs the same metadata-only scan (GGUF header + config.json; no reconversion, no NPU, milliseconds) and discards + re-repacks on failure instead of serving it. The checker resolves as a sibling of ONEBIT_NPU_REPACK_SCRIPT, or via ONEBIT_Q4NX_CHECK; if neither is present it warns loudly and reuses, so a missing checker cannot brick serving. Verified: builds and links (ninja 1bit, 5/5); and the exact command the resolver constructs exits 1 on a stale dir (the three MiniCPM4 scales stripped, which is what a pre-3d717e7 dir looks like) and 0 on the good dir -- so a stale artifact is discarded and rebuilt rather than served. Folds in with the provenance stamp: the stamp comparison (converter path + commit, q/k-reorder flag) belongs in this same gate, at which point reuse requires BOTH halves. Documented in docs/npu.md.
@agent-b35e4d, two closing points: 1. When the checker cannot be resolved, warn-and-reuse applies to every dir for the whole process, i.e. the guard is silently absent -- the same weakness relocated, and a repeated line is easily lost in serve output. That warning is now printed once per process, framed and unmissable, and docs/npu.md states it is the only state where the guard does not apply (so a baked path resolving to a moved checkout is obvious rather than invisible). 2. docs/npu.md now scopes the guard explicitly: it covers newly repacked directories, while the DIRECTORY route (1bit serve -m <Q4NX dir>) consumes a dir with no repack and is trusted as published -- correct for foreign dirs (FLM's Qwen2.5-7B-NPU2 / MiniCPM5-1B-NPU2, the 0.6B fast lane), but stated so nobody reads the guard as covering every NPU serve.
|
Docs7 for 1bit-monster/engine
Commit |
PR Reviewer Guide 🔍(Review updated until commit f62762b)Here are some key observations to aid the review process:
|
18 files (+273 lines) were missing the full notice, which fails CI on PR #179. Applied with 'python3 tools/copyright.py --fix'; --check now passes. Comment-only change plus the previously-uncommitted MLA config fields in repack_gguf.py (host glue: it writes the deepseek2 MLA geometry -- q_lora_rank, kv_lora_rank, rope.dimension_count, key_length_mla/value_length_mla, expert_*, leading_dense_* -- from the GGUF into config.json, which the forward needs to derive its designs). No kernel detail is added by this commit.
The declared-vs-carried scan cannot see which converter produced a dir, and the model-dir cache reuses one on `model.q4nx` existence alone, so the other half of the guard was missing: a stale dir from an older converter was reusable forever. - repack_gguf.py writes `repack-stamp.txt` (converter path + git commit, arch/model_type, the repack script) and gains `--check-stamp <dir>`: 0 current, 1 stale, 2 cannot tell. - app/serve.cpp checks that stamp before the declared-scale scan on a cache hit; a missing or mismatched stamp discards and repacks. 'Cannot tell' (converter unresolvable) fails open with a once-per-process warning, so a missing converter cannot brick serving. Verified end to end with ONEBIT_Q4NX_CONVERTER pointed at the reorder-fixed checkout: a pre-stamp cache dir is discarded and rebuilt (READY 44s), and the next serve reuses it with no repack (READY 5s); the three check-stamp exit codes and a no-stamp dir were each exercised. docs/npu.md now documents the pair instead of listing the stamp as owed, and records what is still converter-side (the q/k reorder flag).
|
Persistent review updated to latest commit b4ddc7e |
…dump The repack, stamp check and declared-scale check were built as shell strings from ONEBIT_Q4NX_* environment variables and run with std::system (CodeQL: command injection, 3 critical). Run them from an argument list with fork/exec (_spawnvp on Windows) instead. This also fixes the stamp check's "cannot tell" branch: std::system returned the raw wait status, so exit code 2 arrived as 512 and never matched; run_program returns the exit code. NPU_INFER_DUMP_HIDDEN wrote /tmp/q4nx_hidden_<pos>.bin with fopen (0666 before umask, and it would follow a planted link); open it 0600 with O_NOFOLLOW. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Persistent review updated to latest commit f62762b |
…and PORTING status (#182) The decode-split q8 pack read global output behind an LDS-only barrier; fixed in llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised (#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…UF on the NPU, 1BP; blog post (#214) README gets a Direction note (HRX with Loom kernels plus the NPU, as geramyL proposed; Vulkan stays the default until the RFC #213 gates are met) and current status lines: six GGUF architectures answer on the NPU (0.006-0.17 tok/s), DwarfStar reads 1BP, and Laya classifies conversations with its scorer on HRX at 15-16 ms. PORTING's NPU and Laya rows gain #179/#181/#209 and #187/#206/#208. The blog post covers the same. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
What this is
The GGUF-on-NPU serving route: a GGUF is repacked to a Q4NX model directory and served
in-process through the model-generic forward on the XDNA 2 NPU.
1bit serve -m <model.gguf> --device npurepacks (converter →model.q4nx+config.json+tokenizer.json), resolves the geometry-keyedfull_i8_*.elfkerneldirectory, and serves behind the OpenAI-compatible API. No generated NPU artifacts are
included in this repo (kernels live outside it, and none appear anywhere in this diff).
Verified on the real NPU (full-ELF backend, i8 GEMMs on device)
12095112258181Each equals the independent llama.cpp oracle for the same ids.
Notable fixes in this branch
embedding_scale(12.0) andresidual_scale(0.2475 =scale_depth/sqrt(n_layer)) applied to the embedding andevery residual branch;
logit_scale(16.0) as a divisor on the final logits; andthe LongRoPE
short_factorapplied astheta / ff(ggml's convention) rather thana multiply. Copying
llama.py's q/k reorder into theminicpmarch needed the headdim from
attention.key_length, because MiniCPM4's GGUF has norope.dimension_count.repack_gguf.pynow fails (rc=1) if theGGUF declares a known-critical scale-like scalar that
config.jsondrops, and warns(never fails) on any unlisted scale-ish key — so
deepseek2.expert_weights_scale, arouting scale, does not false-fail. This class cost hours: the three MiniCPM4 scales
were declared by the GGUF, applied by llama.cpp, and dropped by the repack, and no
argmax-level check can see it.
model.q4nxexistence alone, bypassing the repack and therefore the guard. A hit nowruns the same metadata-only scan and discards + rebuilds on failure.
logit_scalefidelity. Confirmed by refitting a temperature against the oracle:with the divisor in place the best-fit T collapses 0.0586 → 0.9451 and the top-1
probability moves 1.0000 → 0.9449 against the oracle's 0.9291, while the argmax does
not move — which is exactly why every argmax-level check was blind to it.
Known gaps (stated, not hidden)
there is no MLA path (
q_a/q_b/kv_a/k_b/v_b) and no expert path. This row isnot claimed as supported.
(
npu/tokenizer.cpp), so1bit serveon that row falls back to bytes; the explicit-iddecode above is unaffected. The other three are BPE.
(
1bit serve -m <dir>) is trusted as published, anddocs/npu.mdsays so.see
docs/npu.md"From a GGUF".Review notes
docs/npu.mddocuments the repack, the guard (its two tiers, the negative control, andthe single state in which it is inert), and the cache-hit behaviour.
scripts/check_repack_config.pyruns the scan standalone against a GGUF + a produced directory and has a negative control.