Skip to content

memory : copy Hadamard matrix to k_rot tensor only if it has buffer assigned - #27967

Merged
fairydreaming merged 1 commit into
ggml-org:masterfrom
fairydreaming:context-shift-null-k-rot-buffer
Aug 30, 2026
Merged

fairydreaming merged 1 commit into
ggml-org:masterfrom
fairydreaming:context-shift-null-k-rot-buffer

Conversation

@fairydreaming

Copy link
Copy Markdown
Contributor

Overview

This PR prevents crashes during context shift of unquantized K cache in models that use Lightning Indexer.

Additional information

For Lightning Indexer k_rot Hadamard rotation tensors are always created regardless of the K cache type, however context shift graph uses them only if K cache is quantized. When the cache is not quantized k_rot buffer is null during context shift, which results in crashes during set_input_k_rot() call.

Requirements

…ssigned to prevent crashes during context shift of unquantized K cache
@fairydreaming fairydreaming added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Aug 29, 2026
@fairydreaming
fairydreaming merged commit bdf3955 into ggml-org:master Aug 30, 2026
23 of 27 checks passed
TheTom pushed a commit to TheTom/llama-cpp-turboquant that referenced this pull request Sep 3, 2026
Ports every upstream qwen4exp (Qwen3.8-Flash-Next) commit from the past
10 days that this fork's manual PR port had not received:

- reduce graph splits by hoisting the PLE embedding gather out of the
  per-layer loop (ggml-org#27880)
- sum indexer heads via strided adds instead of transpose+sum_rows (ggml-org#28023)
- support recurrent state rollback for MTP speculative decoding (ggml-org#28123)
- rewrite QSA sparse-attention block/bias selection: fixes NaN-producing
  bias rows for short sequences, fixes cross-sequence block pooling in a
  unified KV cache, adds mrope duplicate-position ranking, and fixes a
  CUDA rms_norm gridDim.y overflow (ggml-org#27941)
- indexer cache seq_cp staleness fix, ext.x/ext.y state-restore fix,
  PLE-must-be-linear-attention validation, correct -sm tensor
  disablement (ggml-org#27941)
- Hadamard k_rot context-shift crash fix, shared with other archs (ggml-org#27967)

Also replaces raw GGML_ASSERT aborts in hparams loading with proper
error messages, and adds test coverage: a PLE fixture in
test-llama-archs (which required porting the per_layer_token_embd
row-count-from-metadata fix to make it loadable) and a state
round-trip test in test-save-load-state.

Verified against the real Qwen3.8-Flash-Next model: correct generation
at short and long (~66k token) context, and test-llama-archs passes
qwen4exp on both CUDA and CPU.
OllyJohnston pushed a commit to OllyJohnston/llama.cpp that referenced this pull request Sep 6, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967)

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
(cherry picked from commit bdf3955)
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967)

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967)

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967)

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967)

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants