Skip to content

npu: GGUF-on-NPU serving route (Qwen2.5-7B, MiniCPM4-8B, MiniCPM5-1B answer ' Paris' on the NPU) - #179

Merged
bong-water-water-bong merged 18 commits into
mainfrom
feat/gguf-npu
Sep 27, 2026
Merged

bong-water-water-bong merged 18 commits into
mainfrom
feat/gguf-npu

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

What this is

The GGUF-on-NPU serving route: a GGUF is repacked to a Q4NX model directory and served
in-process through the model-generic forward on the XDNA 2 NPU.

1bit serve -m <model.gguf> --device npu repacks (converter → model.q4nx +
config.json + tokenizer.json), resolves the geometry-keyed full_i8_*.elf kernel
directory, and serves behind the OpenAI-compatible API. No generated NPU artifacts are
included in this repo (kernels live outside it, and none appear anywhere in this diff).

Verified on the real NPU (full-ELF backend, i8 GEMMs on device)

model anchor "The capital of France is" token
Qwen2.5-7B-Instruct 12095 ' Paris'
MiniCPM4-8B 11225 '▁Paris'
MiniCPM5-1B 8181 ' Paris'

Each equals the independent llama.cpp oracle for the same ids.

Notable fixes in this branch

  • MiniCPM4-8B needed three things together: the GGUF's embedding_scale (12.0) and
    residual_scale (0.2475 = scale_depth/sqrt(n_layer)) applied to the embedding and
    every residual branch; logit_scale (16.0) as a divisor on the final logits; and
    the LongRoPE short_factor applied as theta / ff (ggml's convention) rather than
    a multiply. Copying llama.py's q/k reorder into the minicpm arch needed the head
    dim from attention.key_length, because MiniCPM4's GGUF has no
    rope.dimension_count.
  • Repack guard (enforced, not advisory). repack_gguf.py now fails (rc=1) if the
    GGUF declares a known-critical scale-like scalar that config.json drops, and warns
    (never fails) on any unlisted scale-ish key — so deepseek2.expert_weights_scale, a
    routing scale, does not false-fail. This class cost hours: the three MiniCPM4 scales
    were declared by the GGUF, applied by llama.cpp, and dropped by the repack, and no
    argmax-level check can see it
    .
  • Cache-hit validation. The serve route used to reuse a cached Q4NX directory on
    model.q4nx existence alone, bypassing the repack and therefore the guard. A hit now
    runs the same metadata-only scan and discards + rebuilds on failure.
  • logit_scale fidelity. Confirmed by refitting a temperature against the oracle:
    with the divisor in place the best-fit T collapses 0.0586 → 0.9451 and the top-1
    probability moves 1.0000 → 0.9449 against the oracle's 0.9291, while the argmax does
    not move — which is exactly why every argmax-level check was blind to it.

Known gaps (stated, not hidden)

  • GLM-4.7-Flash (deepseek2, MLA + MoE) has a Q4NX and its kernels, but no forward:
    there is no MLA path (q_a/q_b/kv_a/k_b/v_b) and no expert path. This row is
    not claimed as supported.
  • MiniCPM4-8B is SentencePiece/unigram, and this route's tokenizer is BPE-only
    (npu/tokenizer.cpp), so 1bit serve on that row falls back to bytes; the explicit-id
    decode above is unaffected. The other three are BPE.
  • The guard covers newly repacked directories; a foreign Q4NX directory
    (1bit serve -m <dir>) is trusted as published, and docs/npu.md says so.
  • A provenance stamp (converter path + commit, arch, q/k-reorder applied) is still owed;
    see docs/npu.md "From a GGUF".

Review notes

docs/npu.md documents the repack, the guard (its two tiers, the negative control, and
the single state in which it is inert), and the cache-hit behaviour. scripts/check_repack_config.py
runs the scan standalone against a GGUF + a produced directory and has a negative control.

agent and others added 14 commits September 26, 2026 22:14
1bit serve -m <model.gguf> --device npu now repacks the GGUF into a Q4NX
model directory (scripts/repack_gguf.py: FLM_Q4NX_Converter + a config.json
derived from the GGUF metadata) and serves it in process through the
model-generic forward (npu/forward/q4nx_forward.cpp, ported from
1bit-MONSTER's npu-infer), behind the OpenAI endpoints (app/forward_serve.cpp).
The existing Q4NX-model-directory route (unified.cpp) is unchanged.

The forward decodes any arch whose full-ELF geometry is in the kernel set;
the per-shape ELFs live in the model-generic geometry set. Verified on Strix
Halo: Qwen2.5-7B answers 'The capital of France is' -> ' Paris.'.
The forward defaults tie_embeddings=1 when config.json omits
tie_word_embeddings, so an untied model (separate output.weight) got the
embedding table as its LM head and produced garbage. Derive it from the
GGUF's tensor list (output.weight present = untied), and carry the
tokenizer's bos/eos ids + the arch's activation through.
- repack_gguf.py now derives vocab_size from the Q4NX's own embed_tokens
  shape (the converter pads vocab rows to the 32-row tile: 73448 -> 73472)
  instead of the unpadded tokenizer vocab, and carries LongRoPE
  rope_factors_short.weight through config.json as rope_scaling.short_factor.
- the model-generic forward reads that short_factor and multiplies the RoPE
  inv_freq by it (inv *= short_factor[i]), which MiniCPM4's LongRoPE requires.

Without the padded vocab the lm_head dequant failed (73448 % 32 != 0);
without the short_factor MiniCPM4's RoPE was off.
The <think>/</think> idiom is a Qwen3-lane detail; the model-generic route
serves arbitrary archs whose tokenizers do not know those tokens, so the
enable_thinking=false block corrupted the prompt (Qwen2.5-7B chat returned
garbage instead of 'Paris').
The .gguf/npu route could not run at all:
- repack_gguf.py defaulted the Q4NX converter to a dead path
  (~/1bit-MONSTER-iso-build/...), so every repack failed with
  'converter not found'. Resolve $ONEBIT_Q4NX_CONVERTER, then a sibling
  checkout, then the 1bit-MONSTER checkout, and say how to set it on failure.
- serve.cpp defaulted the repack python to a per-user venv path, and rejected a
  bare command name because fs::exists() was applied to it. Default to python3
  and only stat real paths.
- serve.cpp passed an empty --kernels unless ONEBIT_NPU_KERNELS was set, so the
  forward looked for /full_i8_*.elf. Resolve $ONEBIT_NPU_KERNELS ->
  <model dir>/npu -> the configured ONEBIT_NPU_KERNELS_DIR, and fail with an
  actionable message. CMakeLists gains that cache path and bakes it in.

Verified: 1bit serve -m MiniCPM5-1B-Q4_K_M.gguf --device npu repacks, resolves
the kernel directory with no env var, loads the four full ELFs and reports
ready. The chat completion still hangs (the llama 'empty reply' item, task-3).
…ad tiles

Three fixes found by running the route end to end:

- llama-class models were decoding to special-token soup because the forward's
  hardcoded ChatML prompt never carries the BOS that their chat template opens
  with (MiniCPM5-1B: '{{- bos_token }}'). repack_gguf.py now derives
  add_bos_token from tokenizer.chat_template and the forward prepends
  bos_token_id for chat requests. MiniCPM5-1B goes from
  '|end><|fim_middle|>' to coherent text.
- the repack script path was hardcoded to another checkout; bake
  ONEBIT_NPU_REPACK_SCRIPT at configure time and use it as the default.
- run_lm_head re-transposed the whole NV x H head (a T-float strided gather,
  ~150M cache misses) on every token; build the [tile][H][T] layout once.

Measured: Qwen3-0.6B answers coherently through the forward (layer path is
sound) at ~19 s/token, and Qwen2.5-7B loads with all five full-ELF designs, so
the stale 'crashes at load' row no longer applies.
MiniCPM4-8B carries minicpm.embedding_scale (12.0), minicpm.residual_scale
(scale_depth/sqrt(NL) = 0.2475 for 32 layers) and minicpm.logit_scale (16.0).
The model-generic forward must apply the first two (the logit scale is a
constant on the output and does not move the argmax). Without them the
MiniCPM4-8B hidden state runs at the wrong magnitude and the logits explode to
punctuation soup instead of ' Paris'. write_config now carries any of the three
present in the GGUF under the generic names the forward reads.
…' Paris')

Ports the three verified MiniCPM4-8B fixes from np-model-generic (edd6739d8) into
the engine route's forward. Agent -9beb31 verified MiniCPM4-8B is SentencePiece
and -b35e4d noted the serve criterion is additionally tokenizer-gated; this is the
forward half.

1. LongRoPE sign: ggml applies the rope factor as rope_yarn(theta / ff, ...)
   (ggml-cpu/ops.cpp ggml_rope_cache_init), so the angle is DIVIDED by
   freq_factors.  599ac77 multiplied (inv *= short_factor), inflating the high
   dims up to 31x.  Divide now; NPU_INFER_ROPE_FACTOR_MUL=1 keeps the A/B.
2. embedding_scale / residual_scale from config.json (1.0 elsewhere).  MiniCPM4
   scales the embedding by 12.0 and each residual branch by
   scale_depth/sqrt(NL) = 0.2475; without them the hidden state runs at the
   wrong magnitude and the logits explode (absmax 136, punctuation soup).

Verified in np-model-generic: host step-0 argmax 11225 '▁Paris' for ids
1,1507,8107,1379,8360,1410, matching the llama.cpp oracle.
Ports 2928002a9 from np-model-generic.  Reads minicpm.logit_scale (16.0) from
config.json and divides the final logits by it; 1.0 elsewhere.

The forward's logits were ~16x the llama.cpp oracle's (oracle top-1 prob 0.929 vs
effectively one-hot here).  Refitting a temperature against the oracle with the
fix in place collapses it 0.0586 -> 0.9451, i.e. exactly the missing
logit_scale, with own top-1 1.0000 -> 0.9449 and the slice-KL unchanged.  The
argmax does not move (constant positive divisor), which is why no argmax-level
check could see this.  Kept config-driven rather than baked into the converter so
the Q4NX stays faithful to the GGUF.
MiniCPM4-8B decoded garbage for hours because the repack's config.json dropped
embedding_scale / residual_scale / logit_scale, which the GGUF declares and
llama.cpp applies -- and no argmax-level check could see it (the first two move
the answer, the third does not move an argmax at all).  This checks the class
directly: every scale-like scalar the GGUF declares must survive into the
repacked config.json.

  MiniCPM4-8B  -> 3 declared, all present  PASS
  MiniCPM5-1B  -> 0 declared               PASS   (so no hidden config scale)
  Qwen2.5-7B   -> 0 declared               PASS
  negative control (scales stripped) -> FAIL (3 problem(s)), exit 1

Idea from @agent-b35e4d: cmp the repacked config.json against a reference dir's /
the GGUF's own scale keys, instant and oracle-free, as a better guard than a
per-model oracle probe.
@agent-b35e4d's point: the failure mode is silent, so a check nobody runs is
weaker than the bug it targets.  repack_gguf.py now runs the same declared-vs-
carried scan inline after write_config and returns 1 if the GGUF declares a
scale-like key (embedding_scale / residual_scale / logit_scale / scale_emb /
scale_depth / dim_model_base / softcaps) that config.json drops.

Verified: MiniCPM4-8B repack prints '[OK] every scale-like key the GGUF declares
is carried in config.json' then 'repack complete'; the standalone
scripts/check_repack_config.py agrees (PASS), and its negative control (those keys
stripped from a copy) still FAILs with exit 1.
@agent-b35e4d's shape, adopted: FAIL on the known-critical list (the factors we
have proven the forward must apply -- MiniCPM4-8B's three cost hours), WARN but do
not fail on any other scale-ish key the GGUF declares and the config drops.  The
next architecture's factor will have a name nobody listed, so a silent skip would
repeat this bug; but some scale-ish keys are legitimately not carried --
deepseek2.expert_weights_scale (1.8) is a routing scale, not a logit scale -- so a
blanket rule would false-fail correct repacks.  Verified on the warn tier:
GLM-4.7-Flash warns on expert_weights_scale and still PASSes.

check_repack_config.py is refactored into a reusable check(gguf, dir) used by both
the CLI and the inline repack call, so the two cannot drift.

docs/npu.md 'From a GGUF' now states the behaviour change (rc=1 on a dropped
critical key), the known-critical list, why an argmax-level check cannot see this
class, the warn tier, the negative control, and the provenance stamp still owed --
so an unexplained rc=1 is documented where the repack is documented rather than
left to a resume.
@agent-b35e4d's point, and it was the guard's real hole: repack_gguf() returned a
cached dir on 'model.q4nx exists' alone, so the repack -- and therefore the
declared-scale guard -- was bypassed on every subsequent serve.  A dir built
before the MiniCPM4 scale fix would have been served silently forever, and the
guard would only ever have protected freshly built artifacts.

A cache hit now runs the same metadata-only scan (GGUF header + config.json; no
reconversion, no NPU, milliseconds) and discards + re-repacks on failure instead of
serving it.  The checker resolves as a sibling of ONEBIT_NPU_REPACK_SCRIPT, or via
ONEBIT_Q4NX_CHECK; if neither is present it warns loudly and reuses, so a missing
checker cannot brick serving.

Verified: builds and links (ninja 1bit, 5/5); and the exact command the resolver
constructs exits 1 on a stale dir (the three MiniCPM4 scales stripped, which is
what a pre-3d717e7 dir looks like) and 0 on the good dir -- so a stale artifact is
discarded and rebuilt rather than served.

Folds in with the provenance stamp: the stamp comparison (converter path + commit,
q/k-reorder flag) belongs in this same gate, at which point reuse requires BOTH
halves.  Documented in docs/npu.md.
@agent-b35e4d, two closing points:
1. When the checker cannot be resolved, warn-and-reuse applies to every dir for
   the whole process, i.e. the guard is silently absent -- the same weakness
   relocated, and a repeated line is easily lost in serve output.  That warning is
   now printed once per process, framed and unmissable, and docs/npu.md states it
   is the only state where the guard does not apply (so a baked path resolving to a
   moved checkout is obvious rather than invisible).
2. docs/npu.md now scopes the guard explicitly: it covers newly repacked
   directories, while the DIRECTORY route (1bit serve -m <Q4NX dir>) consumes a dir
   with no repack and is trusted as published -- correct for foreign dirs (FLM's
   Qwen2.5-7B-NPU2 / MiniCPM5-1B-NPU2, the 0.6B fast lane), but stated so nobody
   reads the guard as covering every NPU serve.
@context7

context7 Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit f62762b

Comment thread app/serve.cpp Fixed
Comment thread app/serve.cpp Fixed
Comment thread app/serve.cpp Fixed
Comment thread app/serve.cpp Fixed
Comment thread npu/forward/src/q4nx_forward.cpp Fixed
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit f62762b)

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 5 🔵🔵🔵🔵🔵
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Possible Issue

The read_rope_short_factor function parses the short_factor from config.json but does not validate that the parsed values are positive. If the short_factor contains negative or zero values, it could lead to incorrect RoPE computations, especially when these values are used as divisors in the RoPE angle calculation.

while (q < e) {
    char* end = nullptr;
    const float v = strtof(q, &end);
    if (end == q) break;
    out.push_back(v);
    q = end;
}
return out;
Possible Issue

The read_config_float function parses floating-point values from config.json but does not validate that the parsed value is finite. If the configuration contains invalid floating-point values (e.g., NaN or infinity), it could lead to incorrect behavior in downstream computations, especially in scaling factors like embedding_scale, residual_scale, and logit_scale.

    const char* p = strstr(s.c_str(), key);
    if (!p) return dflt;
    const char* c = strchr(p, ':');
    if (!c) return dflt;
    return (float)strtod(c + 1, nullptr);
}
Possible Issue

The load_ple_block function assumes that the tensor data is either in a plain int8 format or a quantized format. If the tensor data is in an unexpected format, the function may not correctly dequantize the data, leading to incorrect PLE computations. This could happen if the tensor's data size does not match the expected format.

const size_t plain_bytes =
    (size_t)sem_.vocab_size_per_layer_input * (size_t)cols;
if ((size_t)td->data_size == plain_bytes) {
    // PLAIN int8 + per-32-column scale (Gemma4): only the requested band of
    // 32 rows is needed, so it is dequantized in place rather than the whole
    // [vocab, NL*ple_dim] table.
    TensorDesc* sd =
        model_tensor_by_name(mw_, "model.per_layer_token_embd.weight.scale");
    const float* sc = sd ? (const float*)model_tensor_data(mw_, sd) : nullptr;
    const uint8_t* band = raw + (size_t)block * Q4NX_TILE_ROWS * (size_t)cols;
    const float* band_sc =
        sc ? sc + (size_t)block * Q4NX_TILE_ROWS * (size_t)(cols / 32) : nullptr;
    if (q4nx_dequant_plain_i8(band, band_sc, Q4NX_TILE_ROWS, cols,
                              ple_block_.data()) != 0) {
        err_ = "per_layer_token_embd dequant failed";
        return false;
    }
} else {
    const size_t band = q4nx_tensor_bytes(Q4NX_TILE_ROWS, cols);
    if (q4nx_dequant_tensor(raw + (size_t)block * band, band, Q4NX_TILE_ROWS,
                            cols, ple_block_.data()) != 0) {
        err_ = "per_layer_token_embd dequant failed";
        return false;
    }
}
ple_block_row_ = block;
return true;
Possible Issue

The gdn_layer function performs several operations including convolution, normalization, and gated recurrent updates. If any of these operations are not correctly implemented or if the input data is not properly validated, it could lead to incorrect outputs. Specifically, the convolution and normalization steps are critical for the correctness of the GDN layer.

bool Q4nxNpuForward::gdn_layer(int l, const std::vector<float>& xn,
                               std::vector<float>& out) {
    const int H = c_.H;
    const int KD = lin_kd_, VD = lin_vd_, CD = 2 * lin_kd_ + lin_vd_;
    const int NKH = lin_nkh_, NVH = lin_nvh_, KHD = lin_khd_, VHD = lin_vhd_;
    const int CK = lin_ck_;
    const int i_qkv = idx_of(Q4NX_DESIGN_QKVLIN);
    const int i_z = idx_of(Q4NX_DESIGN_Z);
    const int i_o = idx_of(Q4NX_DESIGN_O);
    if (i_qkv < 0 || i_z < 0 || i_o < 0) {
        err_ = "GDN designs missing (QKVLIN/Z/O)";
        return false;
    }
    std::vector<float> qkv, z;
    {
        const Q4nxDesignGeom* g = &designs_[i_qkv];
        Q4nxProjections p{};
        p.ple = wqkvlin_.data();
        if (q4nx_assemble_B(g, &p, B_.data()) != 0) { err_ = "assemble gdn qkv"; return false; }
        npu_gemm(*ctx_[i_qkv], xn.data(), g->K, g->N, B_.data(), qkv);
    }
    {
        const Q4nxDesignGeom* g = &designs_[i_z];
        Q4nxProjections p{};
        p.ple = wz_.data();
        if (q4nx_assemble_B(g, &p, B_.data()) != 0) { err_ = "assemble gdn z"; return false; }
        npu_gemm(*ctx_[i_z], xn.data(), g->K, g->N, B_.data(), z);
    }
    // depthwise causal conv1d (kernel CK) + SiLU over CD channels;
    // full = [state (CK-1) | current], new state = full[1:]
    std::vector<float> co(CD);
    float* st = conv_state_[l].data();
    for (int c = 0; c < CD; c++) {
        const float* cs = st + (size_t)c * (CK - 1);
        float s = 0;
        for (int j = 0; j < CK - 1; j++) s += cs[j] * lin_conv_w_[(size_t)j * CD + c];
        s += qkv[c] * lin_conv_w_[(size_t)(CK - 1) * CD + c];
        co[c] = s / (1.0f + expf(-s));
    }
    for (int c = 0; c < CD; c++) {
        float* cs = st + (size_t)c * (CK - 1);
        for (int j = 0; j + 1 < CK - 1; j++) cs[j] = cs[j + 1];
        if (CK > 1) cs[CK - 2] = qkv[c];
    }
    // alpha/beta on the host (N = NVH < 128, below the GEMM tile width)
    std::vector<float> ah(NVH), bh(NVH);
    for (int h = 0; h < NVH; h++) {
        const float* ra = wa_.data() + (size_t)h * H;
        const float* rb = wb_.data() + (size_t)h * H;
        float sa = 0, sb = 0;
        for (int i = 0; i < H; i++) { sa += xn[i] * ra[i]; sb += xn[i] * rb[i]; }
        ah[h] = sa;
        bh[h] = sb;
    }
    // q/k L2-normalised with NVH/NKH head repetition; v per v-head
    const int rep = NVH / NKH;
    std::vector<float> ql((size_t)NVH * KHD), kl((size_t)NVH * KHD), vv((size_t)NVH * VHD);
    for (int h = 0; h < NVH; h++) {
        const int src = h / rep;
        for (int d = 0; d < KHD; d++) ql[(size_t)h * KHD + d] = co[(size_t)src * KHD + d];
        for (int d = 0; d < KHD; d++) kl[(size_t)h * KHD + d] = co[KD + (size_t)src * KHD + d];
        for (int d = 0; d < VHD; d++) vv[(size_t)h * VHD + d] = co[2 * KD + (size_t)h * VHD + d];
    }
    for (int h = 0; h < NVH; h++) {
        float sq = 0, sk = 0;
        for (int d = 0; d < KHD; d++) { sq += ql[(size_t)h*KHD+d]*ql[(size_t)h*KHD+d]; sk += kl[(size_t)h*KHD+d]*kl[(size_t)h*KHD+d]; }
        const float iq = 1.0f / sqrtf(sq + 1e-6f), ik = 1.0f / sqrtf(sk + 1e-6f);
        for (int d = 0; d < KHD; d++) { ql[(size_t)h*KHD+d] *= iq; kl[(size_t)h*KHD+d] *= ik; }
    }
    // recurrent gated delta rule, then scale-free gated RMSNorm, then out_proj
    float* rs = rec_state_[l].data();
    const float scale = 1.0f / sqrtf((float)KHD);
    std::vector<float> core((size_t)NVH * VHD), kvm(VHD);
    for (int h = 0; h < NVH; h++) {
        const float eg = expf(lin_ssm_a_[h] * softplus_f(ah[h] + lin_dt_[h]));
        const float beta = 1.0f / (1.0f + expf(-bh[h]));
        float* sh = rs + (size_t)h * KHD * VHD;
        const float* kh = &kl[(size_t)h * KHD];
        const float* qh = &ql[(size_t)h * KHD];
        const float* vh = &vv[(size_t)h * VHD];
        for (int i = 0; i < KHD * VHD; i++) sh[i] *= eg;
        for (int vd = 0; vd < VHD; vd++) {
            float s = 0;
            for (int kd = 0; kd < KHD; kd++) s += sh[kd * VHD + vd] * kh[kd];
            kvm[vd] = s;
        }
        for (int vd = 0; vd < VHD; vd++) {
            const float delta = (vh[vd] - kvm[vd]) * beta;
            for (int kd = 0; kd < KHD; kd++) sh[kd * VHD + vd] += kh[kd] * delta;
        }
        float ss = 0;
        for (int vd = 0; vd < VHD; vd++) {
            float s = 0;
            for (int kd = 0; kd < KHD; kd++) s += sh[kd * VHD + vd] * qh[kd];
            core[(size_t)h * VHD + vd] = s * scale;
            ss += core[(size_t)h * VHD + vd] * core[(size_t)h * VHD + vd];
        }
        const float inv = 1.0f / sqrtf(ss / (float)VHD + sem_.eps);
        for (int vd = 0; vd < VHD; vd++) {
            const float zz = z[(size_t)h * VHD + vd];
            core[(size_t)h * VHD + vd] = core[(size_t)h * VHD + vd] * inv *
                lin_norm_w_[vd] * (zz / (1.0f + expf(-zz)));
        }
    }
    if (getenv("NPU_INFER_DUMP_LAYERS")) {
        auto dmp = [&](const char* nm, const float* v, int n) {
            double s = 0;
            float mx = 0;
            for (int i = 0; i < n; i++) { float a = fabsf(v[i]); s += a; if (a > mx) mx = a; }
            fprintf(stderr, "  [G] l=%d %-5s %.6f %.4f\n", l, nm, s / n, mx);
        };
        dmp("qkv", qkv.data(), CD);
        dmp("co", co.data(), CD);
        dmp("z", z.data(), VD);
        dmp("a", ah.data(), NVH);
        dmp("b", bh.data(), NVH);
        dmp("core", core.data(), (int)core.size());
    }
    {
        const Q4nxDesignGeom* g = &designs_[i_o];
        Q4nxProjections p{};
        p.o = woutlin_.data();
        if (q4nx_assemble_B(g, &p, B_.data()) != 0) { err_ = "assemble gdn out"; return false; }
        npu_gemm(*ctx_[i_o], core.data(), g->K, g->N, B_.data(), out);
    }
    return true;
}
Possible Issue

The attention function performs RoPE and attention computations. If the RoPE parameters are not correctly computed or if the attention scores are not properly normalized, it could lead to incorrect attention outputs. This is especially critical in models like MiniCPM4 where LongRoPE is used.

bool Q4nxNpuForward::attention(const std::vector<float>& qkv, int pos,
                               std::vector<float>& ctx, int head_dim,
                               int n_kv_heads, float theta, int rotary_dim,
                               bool proportional, int window,
                               int kv_layer, bool append_kv) {
    const int NH = c_.NH;
    const int HD = head_dim;      /* per layer type (Gemma4: 256 or 512) */
    const int NKV = n_kv_heads;
    const int gqa = NH / NKV;

    std::vector<float> q((size_t)NH * HD), k((size_t)NKV * HD), v((size_t)NKV * HD);
    for (int i = 0; i < NH * HD; i++) q[i] = qkv[i];
    for (int i = 0; i < NKV * HD; i++) k[i] = qkv[NH * HD + i];
    for (int i = 0; i < NKV * HD; i++) v[i] = qkv[NH * HD + NKV * HD + i];
    // Qwen2.5-style attention biases: the q/k/v projections carry a per-output
    // bias (config attention_bias=true). Added before qk-norm/RoPE.
    if (sem_.attention_bias) {
        for (int i = 0; i < NH * HD; i++) q[i] += q_bias_[cur_layer_][i];
        for (int i = 0; i < NKV * HD; i++) k[i] += k_bias_[cur_layer_][i];
        for (int i = 0; i < NKV * HD; i++) v[i] += v_bias_[cur_layer_][i];
    }

    // per-head RMSNorm on q and k, only for architectures that have it.  A KV-shared
    // layer (Gemma4) computed no k at all, so its k is never normalized or used.
    if (sem_.qk_norm) {
        std::vector<float> t(HD);
        for (int h = 0; h < NH; h++) {
            rmsnorm(&q[(size_t)h * HD], q_norm_[cur_layer_].data(), HD, t.data());
            memcpy(&q[(size_t)h * HD], t.data(), sizeof(float) * HD);
        }
        if (append_kv) {
            for (int h = 0; h < NKV; h++) {
                rmsnorm(&k[(size_t)h * HD], k_norm_[cur_layer_].data(), HD, t.data());
                memcpy(&k[(size_t)h * HD], t.data(), sizeof(float) * HD);
            }
        }
    }
    // Gemma4 normalises v with a SCALE-FREE RMSNorm (its v_norm has no weight).
    if (sem_.v_norm && append_kv) {
        std::vector<float> tv(HD);
        for (int h = 0; h < NKV; h++) {
            rmsnorm_noweight(&v[(size_t)h * HD], HD, tv.data());
            memcpy(&v[(size_t)h * HD], tv.data(), sizeof(float) * HD);
        }
    }

    // RoPE: half-split over the first `rot` channels (rot == HD means full).
    // Gemma4's full layers use "proportional" RoPE: the pairing spans the WHOLE
    // head_dim, but the frequencies past rot/2 are zero, so those pairs are the
    // identity (_compute_proportional_rope_parameters).
    const int rot = (rotary_dim > 0 && rotary_dim <= HD) ? rotary_dim : HD;
    {
        const int half = proportional ? HD / 2 : rot / 2;
        std::vector<float> cos_(half), sin_(half);
        for (int i = 0; i < half; i++) {
            double inv;
            if (proportional) {
                inv = 1.0 / std::pow((double)theta, 2.0 * (double)i / (double)HD);
                if (i >= rot / 2) inv = 0.0;
            } else {
                inv = 1.0 / std::pow((double)theta, (double)i / (double)half);
                // LongRoPE (MiniCPM4): ggml applies the factor as
                // rope_yarn(theta / ff, ...) (ggml-cpu/ops.cpp
                // ggml_rope_cache_init) -- the angle is DIVIDED by freq_factors.
                // Multiplying inflates the high dims by up to 31x and yields
                // garbage.  NPU_INFER_ROPE_FACTOR_MUL=1 keeps the old multiply.
                if (i < (int)rope_scale_.size()) {
                    const char* mul = getenv("NPU_INFER_ROPE_FACTOR_MUL");
                    if (mul && mul[0] == '1') inv *= (double)rope_scale_[i];
                    else                       inv /= (double)rope_scale_[i];
                }
            }
            const double ang = (double)pos * inv;
            cos_[i] = (float)std::cos(ang);
            sin_[i] = (float)std::sin(ang);
        }
        auto rot_half = [&](float* x) {
            // pairs (i, i + half) inside the rotated block; the rest passes through
            for (int i = 0; i < half; i++) {
                const float x1 = x[i], x2 = x[i + half];
                x[i] = x1 * cos_[i] - x2 * sin_[i];
                x[i + half] = x1 * sin_[i] + x2 * cos_[i];
            }
        };
        for (int h = 0; h < NH; h++) rot_half(&q[(size_t)h * HD]);
        for (int h = 0; h < NKV; h++) rot_half(&k[(size_t)h * HD]);
    }

    // append to cache (a KV-shared layer reuses its owner's, which the owner has
    // already appended earlier in this same step).
    if (append_kv) {
        kcache_[kv_layer].insert(kcache_[kv_layer].end(), k.begin(), k.end());
        vcache_[kv_layer].insert(vcache_[kv_layer].end(), v.begin(), v.end());
    }

    const int T = pos + 1;
    // Gemma4 sets attention scaling to 1.0 (q/k carry an RMSNorm); everyone else
    // uses the usual 1/sqrt(head_dim).
    const float scale = sem_.attn_scaling_one ? 1.0f : 1.0f / std::sqrt((float)HD);
    // Sliding-window attention: ignore positions older than the window.
    const int t0 = (window > 0 && pos + 1 > window)
                       ? pos + 1 - window : 0;
    ctx.assign((size_t)NH * HD, 0.0f);
    std::vector<float> score(T);
    for (int h = 0; h < NH; h++) {
        const int kh = h / gqa;
        const float* qh = &q[(size_t)h * HD];
        float mx = -1e30f;
        for (int t = 0; t < T; t++) {
            if (t < t0) { score[t] = -1e30f; continue; }
            const float* kt = &kcache_[kv_layer][((size_t)t * NKV + kh) * HD];
            float d = 0.0f;
            for (int i = 0; i < HD; i++) d += qh[i] * kt[i];
            score[t] = d * scale;
            if (score[t] > mx) mx = score[t];
        }
        float sum = 0.0f;
        for (int t = 0; t < T; t++) {
            if (t < t0) { score[t] = 0.0f; continue; }
            score[t] = std::exp(score[t] - mx);
            sum += score[t];
        }
        const float inv = sum > 0 ? 1.0f / sum : 0.0f;
        float* out = &ctx[(size_t)h * HD];
        for (int t = t0; t < T; t++) {
            const float p = score[t] * inv;
            const float* vt = &vcache_[kv_layer][((size_t)t * NKV + kh) * HD];
            for (int i = 0; i < HD; i++) out[i] += p * vt[i];
        }
    }
    return true;
}
Possible Issue

The run_lm_head function performs a tiled LM head computation. If the tile size or the way the tiles are processed is incorrect, it could lead to incorrect logits. This is especially important for models that have a large vocabulary size.

// tiled LM head: logits[K*t + n] = dot(embed[K*t + n], h); K = H, tile = QKV N
bool Q4nxNpuForward::run_lm_head(const std::vector<float>& h, std::vector<float>& logits) {
    const int H = c_.H, NV = c_.NV;
    if (host_gemm_) {
        // Offline mode: plain float matvec against the (possibly dequantized)
        // tied table, matching the Python reference.  No I8Ctx is touched.
        const float inv_es0 = (sem_.tie_embeddings && sem_.embed_scale > 0)
                                  ? 1.0f / sem_.embed_scale : 1.0f;
        logits.assign((size_t)NV, 0.0f);
        for (int n = 0; n < NV; n++) {
            float s = 0;
            for (int k = 0; k < H; k++) s += h[(size_t)k] * inv_es0 * lm_at(n, k);
            if (sem_.logit_softcap > 0)
                s = sem_.logit_softcap * std::tanh(s / sem_.logit_softcap);
            logits[(size_t)n] = s;
        }
        return true;
    }
    const int i_qkv = idx_of(Q4NX_DESIGN_QKV);
    const int T = designs_[i_qkv].N;                 // reuse the QKV geometry
    const int ntiles = (NV + T - 1) / T;
    logits.assign(NV, 0.0f);

    // The stored embedding is PRE-SCALED by the architecture's embedding scale
    // (Gemma4: sqrt(hidden_size)), so the LM head must divide that scale back out
    // to recover the raw tied weight.  embed_scale is 1 for every other model.
    const float inv_es = (sem_.tie_embeddings && sem_.embed_scale > 0)
                             ? 1.0f / sem_.embed_scale : 1.0f;
    std::vector<float> hs((size_t)H);
    for (int k = 0; k < H; k++) hs[(size_t)k] = h[(size_t)k] * inv_es;

    const size_t tile_elems = (size_t)H * (size_t)T;
    if (lm_head_tiles_.size() != (size_t)ntiles * tile_elems || lm_head_tiles_T_ != T) {
        // Once per model: [tile][H][T], so each token copies sequentially instead
        // of gathering the whole head with a T-float stride.
        lm_head_tiles_.assign((size_t)ntiles * tile_elems, 0.0f);
        for (int t = 0; t < ntiles; t++) {
            const int n0 = t * T, n1 = std::min(NV, n0 + T);
            float* dst = lm_head_tiles_.data() + (size_t)t * tile_elems;
            for (int n = n0; n < n1; n++)
                for (int k = 0; k < H; k++)
                    dst[(size_t)k * T + (n - n0)] = lm_at(n, k);
        }
        lm_head_tiles_T_ = T;
    }
    std::vector<float> Bt(tile_elems);
    for (int t = 0; t < ntiles; t++) {
        const int n0 = t * T, n1 = std::min(NV, n0 + T);
        std::memcpy(Bt.data(), lm_head_tiles_.data() + (size_t)t * tile_elems,
                    tile_elems * sizeof(float));
        float sout = 1.0f;
        ctx_[i_qkv]->packB(0, Bt.data(), H, T, sout);
        float amax = 0.0f;
        for (int k = 0; k < H; k++) { float a = std::fabs(hs[(size_t)k]); if (a > amax) amax = a; }
        const float as = amax > 0 ? amax / 127.0f : 1.0f;
        ctx_[i_qkv]->quantize_async(hs.data(), 1, H, as);
        xrt::run r = ctx_[i_qkv]->sync_and_launch(0);
        ctx_[i_qkv]->wait_kernel(r);
        ctx_[i_qkv]->readback();
        const std::vector<float>& gs = ctx_[i_qkv]->group_scales[0];
        for (int n = n0; n < n1; n++) {
            float v = (float)ctx_[i_qkv]->Cm[n - n0] * as * gs[(size_t)(n - n0)];
            if (sem_.logit_softcap > 0)
                v = sem_.logit_softcap * std::tanh(v / sem_.logit_softcap);
            logits[(size_t)n] = v;
        }
    }
    return true;
}

⚠️ Review coverage: The following files were not included in this review because of the token budget:

  • app/forward_serve.cpp
  • app/serve.cpp
  • scripts/repack_gguf.py
  • scripts/check_repack_config.py
  • npu/forward/src/model.c
  • npu/forward/engine-src/npu_engine_i8ctx_inc.h
  • npu/forward/engine-src/gu_i4_pack.h
  • npu/forward/src/q4nx_semantics.c
  • npu/forward/include/model_config.h
  • npu/forward/src/q4nx_pack.c
  • npu/forward/engine-src/q4nx_raw.h
  • npu/forward/include/q4nx_forward.h
  • npu/forward/include/model.h
  • npu/forward/include/q4nx_semantics.h
  • npu/forward/src/q4nx_dequant.c
  • npu/forward/include/q4nx_pack.h
  • npu/forward/include/common.h
  • npu/forward/include/q4nx_dequant.h
  • docs/npu.md
  • npu/forward/include/device.h
  • app/forward_serve.h

pi agent and others added 2 commits September 27, 2026 17:54
18 files (+273 lines) were missing the full notice, which fails CI on PR #179.
Applied with 'python3 tools/copyright.py --fix'; --check now passes.  Comment-only
change plus the previously-uncommitted MLA config fields in repack_gguf.py (host
glue: it writes the deepseek2 MLA geometry -- q_lora_rank, kv_lora_rank,
rope.dimension_count, key_length_mla/value_length_mla, expert_*, leading_dense_*
-- from the GGUF into config.json, which the forward needs to derive its designs).
No kernel detail is added by this commit.
The declared-vs-carried scan cannot see which converter produced a dir, and the
model-dir cache reuses one on `model.q4nx` existence alone, so the other half of
the guard was missing: a stale dir from an older converter was reusable forever.

- repack_gguf.py writes `repack-stamp.txt` (converter path + git commit,
  arch/model_type, the repack script) and gains `--check-stamp <dir>`: 0 current,
  1 stale, 2 cannot tell.
- app/serve.cpp checks that stamp before the declared-scale scan on a cache hit;
  a missing or mismatched stamp discards and repacks.  'Cannot tell' (converter
  unresolvable) fails open with a once-per-process warning, so a missing converter
  cannot brick serving.

Verified end to end with ONEBIT_Q4NX_CONVERTER pointed at the reorder-fixed
checkout: a pre-stamp cache dir is discarded and rebuilt (READY 44s), and the next
serve reuses it with no repack (READY 5s); the three check-stamp exit codes and a
no-stamp dir were each exercised.  docs/npu.md now documents the pair instead of
listing the stamp as owed, and records what is still converter-side (the q/k
reorder flag).
Comment thread app/serve.cpp Fixed
Comment thread app/serve.cpp Fixed
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit b4ddc7e

bong-water-water-bong and others added 2 commits September 27, 2026 19:01
…dump

The repack, stamp check and declared-scale check were built as shell strings from
ONEBIT_Q4NX_* environment variables and run with std::system (CodeQL: command
injection, 3 critical). Run them from an argument list with fork/exec
(_spawnvp on Windows) instead. This also fixes the stamp check's "cannot tell"
branch: std::system returned the raw wait status, so exit code 2 arrived as 512
and never matched; run_program returns the exit code.

NPU_INFER_DUMP_HIDDEN wrote /tmp/q4nx_hidden_<pos>.bin with fopen (0666 before
umask, and it would follow a planted link); open it 0600 with O_NOFOLLOW.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread app/serve.cpp Dismissed
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit f62762b

@bong-water-water-bong
bong-water-water-bong merged commit 41dd3c0 into main Sep 27, 2026
10 of 11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the feat/gguf-npu branch September 27, 2026 22:07
bong-water-water-bong added a commit that referenced this pull request Sep 28, 2026
…and PORTING status (#182)

The decode-split q8 pack read global output behind an LDS-only barrier; fixed in
llama.cpp 00adc2b (#176), kernel on by default again (#178), multipass vectorised
(#180). Recap: architecture gaps closed (323 HF architectures mapped), GGUF on the
NPU (#179). Every number from docs/hrx.md, docs/registry.md, docs/npu.md.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 29, 2026
…UF on the NPU, 1BP; blog post (#214)

README gets a Direction note (HRX with Loom kernels plus the NPU, as geramyL proposed;
Vulkan stays the default until the RFC #213 gates are met) and current status lines:
six GGUF architectures answer on the NPU (0.006-0.17 tok/s), DwarfStar reads 1BP, and
Laya classifies conversations with its scorer on HRX at 15-16 ms. PORTING's NPU and
Laya rows gain #179/#181/#209 and #187/#206/#208. The blog post covers the same.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants