What happens
Qwen3.8-27B UD-Q4_K_XL runs on HRX0 with --spec-type draft-mtp and the MTP head (-md mtp-Qwen3.8-27B-Q4_0.gguf -ngld 99 -devd HRX0 --spec-draft-n-max 3 --spec-draft-p-min 0.75). The test sends 6 identical /completion requests (greedy, n_probs: 5, cache_prompt: false, prompt "The B-tree insert algorithm works as follows:", 64 tokens). The top-5 logprobs alternate every other request:
- Requests 1, 3 and 5 are bit-identical, and requests 2, 4 and 6 are bit-identical.
- Between the two groups, the first difference is at generated step 15, with |dlogprob| up to 0.0124. The top-5 ids change at step 53.
This uses llama.cpp fork 388290b plus the FA masked-V fix (1bit/hrx-fa-masked-v). That fix removes the separate kernel-side dependence on earlier requests.
Localization
| setting |
result |
| defaults (checkpoints on, max 32) |
alternates every other request |
--ctx-checkpoints 0 |
6/6 identical |
--cache-ram 0 (prompt cache off) |
still alternates |
draft on CPU (-devd none -ngld 0) |
still varies, with a 4-request cycle |
The behaviour does not depend on the HRX kernels, because it persists with the draft on CPU.
The -lv 5 trace shows the following, in order:
- Every request creates one checkpoint at prompt pos 4 (pos_min = pos_max = 4, n_tokens = 5, 149.6 MiB), so the 10-token prompt is processed as 5 + 5.
- The checkpoint is erased at the next request (pos_next = 0) and is never restored.
- The MTP draft head's first prediction already differs between requests: draft candidate 0 at pos 0 has p = 0.862 vs 0.847.
- That changes how many drafts are accepted, so the target's verify batch shapes differ, and the target's numerics differ with them.
Places to look:
common/speculative.cpp: the draft-mtp state (pending_g_last, pending_pos_last, verify_g, and the cross-ubatch bridge in process()), and get_state/set_state around checkpoint creation.
tools/server/server-context.cpp: create_checkpoint.
- Recurrent
state_write with n_rs_seq > 0 (src/llama-memory-recurrent.cpp).
Effect
With f32 activations the alternation stays below the greedy threshold: 8/8 identical chat texts. With the Q8_1-activation MTP kernels (1bit/hrx-kquant-tokens-q8share, b9fcba5b4 + FA fix), chat x8 gives 2 alternating texts.
With --ctx-checkpoints 0 the same build gives 8/8 identical text, equal to the f32 build. It is also faster: median 18.3 tok/s vs 16.9 with checkpoints (balanced 85 W).
Repro
Script ~/wt/loom-probdet-scratch/t8.sh on strixhalo, about 75 s per config.
- It takes
box.lock and starts llama-server under systemd-run MemoryMax=70G.
- It sends N identical requests and compares the top-5 logprobs.
- Environment:
SPECS="label:bindir:env:extra,args" and NREQ=6. Example: SPECS="nockpt:bin-fix3::--ctx-checkpoints,0".
- Logs and responses go to
d27/mp-<label>/.
Workaround until this is fixed: run MTP with --ctx-checkpoints 0.
What happens
Qwen3.8-27B UD-Q4_K_XL runs on HRX0 with
--spec-type draft-mtpand the MTP head (-md mtp-Qwen3.8-27B-Q4_0.gguf -ngld 99 -devd HRX0 --spec-draft-n-max 3 --spec-draft-p-min 0.75). The test sends 6 identical/completionrequests (greedy,n_probs: 5,cache_prompt: false, prompt "The B-tree insert algorithm works as follows:", 64 tokens). The top-5 logprobs alternate every other request:This uses llama.cpp fork 388290b plus the FA masked-V fix (
1bit/hrx-fa-masked-v). That fix removes the separate kernel-side dependence on earlier requests.Localization
--ctx-checkpoints 0--cache-ram 0(prompt cache off)-devd none -ngld 0)The behaviour does not depend on the HRX kernels, because it persists with the draft on CPU.
The
-lv 5trace shows the following, in order:Places to look:
common/speculative.cpp: the draft-mtp state (pending_g_last,pending_pos_last,verify_g, and the cross-ubatch bridge inprocess()), andget_state/set_statearound checkpoint creation.tools/server/server-context.cpp:create_checkpoint.state_writewithn_rs_seq > 0(src/llama-memory-recurrent.cpp).Effect
With f32 activations the alternation stays below the greedy threshold: 8/8 identical chat texts. With the Q8_1-activation MTP kernels (
1bit/hrx-kquant-tokens-q8share, b9fcba5b4 + FA fix), chat x8 gives 2 alternating texts.With
--ctx-checkpoints 0the same build gives 8/8 identical text, equal to the f32 build. It is also faster: median 18.3 tok/s vs 16.9 with checkpoints (balanced 85 W).Repro
Script
~/wt/loom-probdet-scratch/t8.shon strixhalo, about 75 s per config.box.lockand startsllama-serverundersystemd-run MemoryMax=70G.SPECS="label:bindir:env:extra,args"andNREQ=6. Example:SPECS="nockpt:bin-fix3::--ctx-checkpoints,0".d27/mp-<label>/.Workaround until this is fixed: run MTP with
--ctx-checkpoints 0.