memory : copy Hadamard matrix to k_rot tensor only if it has buffer assigned - #27967
Merged
fairydreaming merged 1 commit intoAug 30, 2026
Merged
Conversation
…ssigned to prevent crashes during context shift of unquantized K cache
ggerganov
approved these changes
Aug 29, 2026
CISC
approved these changes
Aug 29, 2026
TheTom
pushed a commit
to TheTom/llama-cpp-turboquant
that referenced
this pull request
Sep 3, 2026
Ports every upstream qwen4exp (Qwen3.8-Flash-Next) commit from the past 10 days that this fork's manual PR port had not received: - reduce graph splits by hoisting the PLE embedding gather out of the per-layer loop (ggml-org#27880) - sum indexer heads via strided adds instead of transpose+sum_rows (ggml-org#28023) - support recurrent state rollback for MTP speculative decoding (ggml-org#28123) - rewrite QSA sparse-attention block/bias selection: fixes NaN-producing bias rows for short sequences, fixes cross-sequence block pooling in a unified KV cache, adds mrope duplicate-position ranking, and fixes a CUDA rms_norm gridDim.y overflow (ggml-org#27941) - indexer cache seq_cp staleness fix, ext.x/ext.y state-restore fix, PLE-must-be-linear-attention validation, correct -sm tensor disablement (ggml-org#27941) - Hadamard k_rot context-shift crash fix, shared with other archs (ggml-org#27967) Also replaces raw GGML_ASSERT aborts in hparams loading with proper error messages, and adds test coverage: a PLE fixture in test-llama-archs (which required porting the per_layer_token_embd row-count-from-metadata fix to make it loadable) and a state round-trip test in test-save-load-state. Verified against the real Qwen3.8-Flash-Next model: correct generation at short and long (~66k token) context, and test-llama-archs passes qwen4exp on both CUDA and CPU.
OllyJohnston
pushed a commit
to OllyJohnston/llama.cpp
that referenced
this pull request
Sep 6, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com> (cherry picked from commit bdf3955)
thecodacus
pushed a commit
to thecodacus/llama.cpp
that referenced
this pull request
Sep 7, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
zbrad
pushed a commit
to zbrad/llama.cpp
that referenced
this pull request
Sep 10, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
pl752
pushed a commit
to pl752/llama.cpp
that referenced
this pull request
Sep 15, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
zsogitbe
pushed a commit
to zsogitbe/llama.cpp
that referenced
this pull request
Sep 17, 2026
…ssigned to prevent crashes during context shift of unquantized K cache (ggml-org#27967) Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: AesSedai <7980540+AesSedai@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR prevents crashes during context shift of unquantized K cache in models that use Lightning Indexer.
Additional information
For Lightning Indexer
k_rotHadamard rotation tensors are always created regardless of the K cache type, however context shift graph uses them only if K cache is quantized. When the cache is not quantized k_rot buffer is null during context shift, which results in crashes duringset_input_k_rot()call.Requirements