Skip to content

MTP: context checkpoints make repeated identical requests give results that alternate every other request #290

Description

@bong-water-water-bong

What happens

Qwen3.8-27B UD-Q4_K_XL runs on HRX0 with --spec-type draft-mtp and the MTP head (-md mtp-Qwen3.8-27B-Q4_0.gguf -ngld 99 -devd HRX0 --spec-draft-n-max 3 --spec-draft-p-min 0.75). The test sends 6 identical /completion requests (greedy, n_probs: 5, cache_prompt: false, prompt "The B-tree insert algorithm works as follows:", 64 tokens). The top-5 logprobs alternate every other request:

  • Requests 1, 3 and 5 are bit-identical, and requests 2, 4 and 6 are bit-identical.
  • Between the two groups, the first difference is at generated step 15, with |dlogprob| up to 0.0124. The top-5 ids change at step 53.

This uses llama.cpp fork 388290b plus the FA masked-V fix (1bit/hrx-fa-masked-v). That fix removes the separate kernel-side dependence on earlier requests.

Localization

setting result
defaults (checkpoints on, max 32) alternates every other request
--ctx-checkpoints 0 6/6 identical
--cache-ram 0 (prompt cache off) still alternates
draft on CPU (-devd none -ngld 0) still varies, with a 4-request cycle

The behaviour does not depend on the HRX kernels, because it persists with the draft on CPU.

The -lv 5 trace shows the following, in order:

  1. Every request creates one checkpoint at prompt pos 4 (pos_min = pos_max = 4, n_tokens = 5, 149.6 MiB), so the 10-token prompt is processed as 5 + 5.
  2. The checkpoint is erased at the next request (pos_next = 0) and is never restored.
  3. The MTP draft head's first prediction already differs between requests: draft candidate 0 at pos 0 has p = 0.862 vs 0.847.
  4. That changes how many drafts are accepted, so the target's verify batch shapes differ, and the target's numerics differ with them.

Places to look:

  • common/speculative.cpp: the draft-mtp state (pending_g_last, pending_pos_last, verify_g, and the cross-ubatch bridge in process()), and get_state/set_state around checkpoint creation.
  • tools/server/server-context.cpp: create_checkpoint.
  • Recurrent state_write with n_rs_seq > 0 (src/llama-memory-recurrent.cpp).

Effect

With f32 activations the alternation stays below the greedy threshold: 8/8 identical chat texts. With the Q8_1-activation MTP kernels (1bit/hrx-kquant-tokens-q8share, b9fcba5b4 + FA fix), chat x8 gives 2 alternating texts.

With --ctx-checkpoints 0 the same build gives 8/8 identical text, equal to the f32 build. It is also faster: median 18.3 tok/s vs 16.9 with checkpoints (balanced 85 W).

Repro

Script ~/wt/loom-probdet-scratch/t8.sh on strixhalo, about 75 s per config.

  • It takes box.lock and starts llama-server under systemd-run MemoryMax=70G.
  • It sends N identical requests and compares the top-5 logprobs.
  • Environment: SPECS="label:bindir:env:extra,args" and NREQ=6. Example: SPECS="nockpt:bin-fix3::--ctx-checkpoints,0".
  • Logs and responses go to d27/mp-<label>/.

Workaround until this is fixed: run MTP with --ctx-checkpoints 0.

Activity

  1. bong-water-water-bong commented on Oct 5, 2026

    @bong-water-water-bong
    CollaboratorAuthor

    Status 2026-10-04: fix in fork #83 (MTP no longer bridges the previous h-row into a fresh sequence). Verified on 3 Oct: the 27B with an MTP head gave 6/6 identical results with context checkpoints. It waits for the release, which is on hold.

  2. bong-water-water-bong commented on Oct 5, 2026

    @bong-water-water-bong
    CollaboratorAuthor

    Fork #83 is merged into 1bit/hrx-vulkan-patched (86eee8985). This issue closes when an engine pin includes it.

  3. bong-water-water-bong commented on Oct 5, 2026

    @bong-water-water-bong
    CollaboratorAuthor

    Related: fork #87 (merged, f3c771f) fixes the decode slowdown after a prompt, caused by the sleeping wait. Separately, a Flash-Next fork PR is coming with an MTP fix: hidden-state rows copied at n_embd instead of n_embd_out gave all-NaN draft logits.

  4. bong-water-water-bong commented on Oct 5, 2026

    @bong-water-water-bong
    CollaboratorAuthor

    Fixed: MTP no longer bridges the previous h-row into a fresh sequence (fork #83, 34e545f; 6/6 identical results with context checkpoints on the 27B), in the engine since llama.cpp pin e57beb97 (#311, main 4aa4ac2). The sleeping-wait decode slowdown (fork #87) and the Flash-Next n_embd_out MTP fix (fork #88) are in the same pin.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions