Repository navigation
Sync upstream master (2026-10-08, ggml-org a11f57ba9) - #18
Merged
Merged
Conversation
* hexagon: ssm-conv double-buffered DMA for prefill and decode restructuring * hex-ssm-conv: remove divs from loops and fix trace events * hex-dma: improved SSM_CONV dma pipeline and streamlined dma_queue --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* llama: re-pool each shared k-pool rep once With shared cells every pool is re-pooled, and since the pooled keys are always scattered, the pools a seq_cp shares between sequences wrote the same rep row from several scatter entries, a data race on the CPU backend. Mark each rep once: the sharing sequences read the same row through pool_cells. * llama: assert whole-sequence seq_cp in the hybrid idx memory The recurrent state is always copied whole whatever the range, and a k-pool cell shared by a partial copy could carry two pool groupings with a single pooled row. Every caller copies whole sequences, so reject partial ranges instead of supporting them. * llama: drop the k-pool cache_safe mode With whole-sequence seq_cp, sequences sharing cells share their pools too, so the pooled row of a shared rep is valid for all of them. Mark each rep once in every ubatch instead of re-pooling everything while cells are shared, which removes the sharing scan and the stale-all workarounds in seq_rm, state_read and state_drop. seq_cp now only stales the destination.
* ggml-openvino: skip unselected graph branches and support DUP Upstream ggml-org#29622 adds a mixed token/embd branch to every input embedding graph through ggml_build_forward_select(). Its nodes are not flagged for compute, but the backend translated them anyway, and the DUP in that branch was unsupported, so the scheduler split the graph and passed the embeddings across the split with a fixed token count. The first single-token decode then failed (test-thread-safety on CPU and GPU). Build the OV model from the compute nodes only, and translate a same-type contiguous DUP like CONT so the graph stays on one backend. * ggml-openvino: make inp_scale_rows token dim dynamic ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n, gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim and pad it per chunk on the static (NPU) path. * ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU plugin fails to compile that form for some row counts with "clFinish, error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model weights use a u4 zero point and are not affected. Report these cases as unsupported on GPU until the plugin is fixed. Op tests check support before allocating, so the check matches unbound weights only; model loading probes with a dummy buffer and keeps its weights on the GPU. * ggml-openvino: create FILL in the output type translate_fill always built an f32 constant, so an f16 FILL produced f32 data and the copy back overran the f16 output buffer. Use the output type for the constant. * ggml-openvino: reject CONCAT with a quantized type Quantized inputs are dequantized when translated, so the backend cannot write a quantized CONCAT output. Report it as unsupported, as for CPY to a quantized type. * ggml-openvino: handle the single recurrent state gather of build_rs ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS on the s_copy leaf and take the ubatch and extra states as views of it. The stateful path matched only the previous form, a GET_ROWS per view of s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU ("is_axis_valid(axis, r)" in a Concat). For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the active-state gather, keep the rank-4 layout of reshapes that read a view of it, and map the copy of the empty extra-state view to the single-slot remainder writeback. Do not warn about the dynamic dim of empty views. * openvino: align eltwise operand ranks to work around a GPU-plugin defect * openvino: match the MoE fusion on the rank-3 stateful graph * ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract whose operand ranks differ. In gemma-3 the lower-rank operand of the post-attention residual add is the norm output, and unsqueezing it makes the GPU plugin compute the layer wrongly: gemma-3 returns empty answers on GPU with stateful execution. Skip the rewrite when the lower-rank operand is an RMS norm output. * docs : update OpenVINO validated models --------- Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>
…-org#30034) With GGML_BACKEND_DL=ON, backends must be loaded explicitly before creating models. Assisted-by: Codex
* ggml: refactor selective expert copying to user code * tests: enroll two models into selective expert copy test * tests: use deepseek2 as test model * improve comment in ggml-backend.h Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * cont: fix whitespace * cont : better comments Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…gml-org#30020) The speculative MTP init enables NextN extraction on the target and draft contexts after both were created and their schedulers reserved. With unmasked extraction the trunk graph keeps every token through the last layer instead of cropping to the output rows, so the first decode reallocates to that batch's shape and the next, wider batch trips GGML_SCHED_DEBUG_REALLOC. Invalidate the reserve when the flags change so the next compute re-reserves with the new graph shape. Assisted-by: Claude
…0040) * add text_config as fallback * remove redundant llm_config check Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…#30017) * mimo2 : always emit h_nextn the other nextn-capable models set it unconditionally Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD * models : consolidate nextn row cropping into shared helpers - replace the duplicated crop conditions and the per-model flags (narrow_early, crop_before_ffn, crop_last_layer, emit_h_nextn) with two helpers on llm_graph_context: crop_before_nextn() / crop_after_nextn() - models that only tested embeddings_nextn_masked now share the same condition, so they crop the last layer before the nextn capture whenever extraction is off - t_h_nextn is now set unconditionally in mimo2, qwen4exp and deepseek4 (as in the other nextn-capable models); host-side reads stay gated by cparams.embeddings_nextn Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
A 1.0 loader (e.g. Android 8.1) has no vkEnumerateInstanceVersion, so backend init called a null pointer. Treat it like any loader under 1.2. Fixes ggml-org#29871. Assisted-by: Claude Opus 5.5
* ggml-cuda: chunk large BF16/FP16 to F32 conversions * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update ggml/src/ggml-cuda/ggml-cuda.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * ggml-cuda: respect dst stride in chunked cuBLAS matmul --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* ggml: fix CLAMP on non-contiguous views (CPU, CUDA) CUDA clamped ggml_nelements values flat and ignored the view strides. CPU addressed row j as j*nb01 and ignored nb02/nb03. Both now follow the strides of dims 1..3; CUDA supports_op requires contiguous rows, like Metal. test_clamp gains a non-contiguous view case. * cuda: clamp kernel uses fastdiv for the view strides
* rpc: allow -sm tensor * fix flush for apple rdma * move graph_uids to rpc_dispatcher * cont : fix conflict * cont: stop spinning dispatcher thread * remove meta backend change * add TODO to simplify logic * rpc: bump major version --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…-org#30042) The gather path attended over the selected latents with a plain matmul and softmax. It only ran with n_ubatch <= 16, and the flash attention backends now skip the masked rows through n_kv_max, so the scatter path covers every case. Drop the gather flag, the gathered attention branch and gather_mla_rows. set_input_kpool always maps padding to the n_kv sentinel, and the slot mask becomes sel_mask since only the scatter reads it.
* model: K2 Horizon gguf conversion code
* model: loading hparams and tensors in k2-horizon.cpp
* model: K2 Horizon compute graph
* model: K2 Horizon compute graph adjustment and registering tokenizers
* model: K2 Horizon chat template and accomodate safetensors naming
* unicode : add the K2-Horizon pre-tokenizer splitter
The K2-Horizon regex had no arm in unicode_regex_split_custom and fell through to the
general std::regex fallback, which fails two ways.
On MSVC std::regex rejects \p{...}, so no K2-Horizon GGUF loads on Windows at all:
llama-quantize, llama-imatrix and llama-perplexity all abort with
regex_error(error_escape) before a token is produced.
Where the fallback does compile it is still wrong. unicode_regex_split collapses each
codepoint to a single byte naming its Unicode category before matching, and U+200C/U+200D
are category Control, which has no entry in k_ucat_cpt, so both become the 0xD0 fallback
byte. The literal and alternatives in K2's regex can then never match and
every ZWNJ or ZWJ ends a letter run.
The splitter is the existing llama3 one with a single rule widened, since K2's regex
differs from llama3's only in that a letter run also takes marks, ZWNJ and ZWJ.
tests/test-unicode.cpp gains a case for this: it fails before the change with
[Amy] [ZWNJ khaham] and passes after with the run intact.
* tests: expand K2 Horizon unicode splitter coverage
* unicode: handle K2 Horizon case folding and empty input
Assisted-by: Codex
* jinja : support sequence indices in selectattr and rejectattr
Assisted-by: Codex
* model : add K2 Horizon dense and MoVA support
Includes the K2 Horizon implementation from ifm-ai/llama.cpp with converter, tensor-parallel and model save/reload fixes.
Assisted-by: Codex
* chat : support K2 Horizon reasoning and tool calls
Assisted-by: Codex
* conversion: remove obsolete K2 Aurora alias
Assisted-by: Codex
* k2-horizon: enforce response schemas and load YaRN betas
Constrain final JSON after reasoning, accept flexible JSON tool envelopes,
enforce XML dialects, and handle repeated or alternate thinking markers.
Load YaRN beta metadata instead of retaining the default values.
Add schema, streaming, continuation, and model reload regressions. Validate
CUDA and CPU builds and 0.9B, 4B, and MoVA conversation/tool round trips.
Assisted-by: Codex
* renaming template fixture
* adressing cisc follows ups
* desloppify the parser / adress aldehir comments
* clean test-chat
* remove fallback : model trained mostly on high anyway
* fix k2 attn_v_exp tn splitting and metal fusion baseline
* k2-horizon : forward expand views before sums
* k2-horizon: copy embds before group norm to fix TP
* disable tesnor parallelism
---------
Co-authored-by: Ryandito Diandaru <ryandito.diandaru@mbzuai.ac.ae>
Co-authored-by: WestWaters <mario.papaleo2013@gmail.com>
Co-authored-by: Natani L. Mayday <71436458+TaskPuppyNatani@users.noreply.github.com>
Co-authored-by: West <100190545+WestWaters@users.noreply.github.com>
Co-authored-by: aaryamonvikram <aaryamonvikram@gmail.com>
Co-authored-by: aaryamonvikram <96529820+aaryamonvikram@users.noreply.github.com>
* vocab : implement PLaMo-3 tokenizer pre-segmentation The PLaMo-3 tokenizer inserts hard boundaries before running the Unigram DP, around <|plamo:...|>-looking text, and around runs of at least 4 identical characters or 2 spaces. Without them llama.cpp tokenizes code indentation and repeated punctuation differently from the reference. Reproduce the two re.sub() passes in llm_tokenizer_plamo2 by encoding each segment independently. * add vocab type "plamo3" * Update src/llama-vocab.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * misc change --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…-org#28782) * ggml-cuda: use per-thread stream for buffer-init padding memset * ci : re-enable test-backend-ops -j for ROCm
The XIELU CUDA kernel template is already generic over the element type; only the F32/F16 type assertion and the else-if dispatch were missing. Add the nv_bfloat16 branch to the launcher, and drop the temporary supports_op gate in ggml-cuda.cu that rejected BF16+XIELU. test-backend-ops gains two BF16 cases ([10,5,4,3] and [512,16,1,1]). docs/ops/CUDA.csv and docs/ops.md are regenerated; the F32 xIELU row flips from no to yes as well, i.e. the previous record was stale. Tested: - Mac CPU: xIELU F32/F16/BF16, 6/6 - Mac Metal: existing F32/F16, 4/4; BF16 still unsupported - RTX 4090 CUDA: xIELU F32/F16/BF16, 6/6 - RTX 4090 CUDA BF16-only: 2/2 - git diff --check passes
…gml-org#30067) * hex-cpy: replace more paths with dma and simplify l2flush * hex-cpy: use DMA in all sametype paths * hex-cpy: rewrite the rest of the copy paths (diff type) to use dma * hex-concat: use dma for multi-dev path which also removes the need for l2-line alignment * hex-concat: proper support for mdev splitting * hex-concat: cleanup ctx and kern params usage * hex-cpy: cleanup contex and remove left-over non-dma checks * hex-cpy: clean dma_cpy naming * hex-cpy: proper kernel params and kernel selection * hex-build: resolve left-over rebase conflicts * hex-cpy: update dev guide to clarify 128 byte alignment requirement * hex-concat: make sure we go through mdev barrier * hex-concat: make sure to flush dma-queue * hex-cpy/concat: cleanup kparams and vtcm layout handling * hex-cpy: remove dead check for contig (routed to diff kernel) and update comments * hex-dev: update developer guide based on latest changes * hex-cpy: safe skip of noop copies * hex-dup: route DUP to CPY
…9171) * sycl: accelerate GLM MLA prefill with MKL flash attention GLM-4.7 Flash uses an MLA shape with 576-wide Q/K heads, a 512-wide V head, GQA 20, and F16 KV. The SYCL dispatcher rejects this shape because the normal MKL flash-attention gate requires matching K/V widths and caps the head dimension at 512, so prompt processing falls back to the substantially slower TILE kernel. Admit only the validated 576/576/512, GQA-20 F16 shape to the existing MKL pipeline. Keep all other mismatched K/V shapes on their current fallback paths. Handle GLM's V cache as a narrower strided view of K rows. Select the strided F16 descriptor when row stride is padded, and alias K/V dequantization buffers only when their logical widths match. Restrict the stride exception to a real V view sharing K's row stride. Add the exact 576/512, GQA-20 prompt-path backend test. On an Intel Arc Pro B70 at master e613ef2, pp8192 improves from 432.80 to 1292.29 tok/s (2.99x, +198.6%). tg256 remains unchanged within noise at 45.67 versus 45.65 tok/s. The exact MLA test passes and debug output confirms MKL dispatch. * sycl: store MKL flash attention scores in F16 Keep the QK GEMM output in F16 instead of F32. The online softmax still converts each score to F32 for its max, exponent, and sum, so the per-element math is unchanged apart from score rounding, and the F32 matrix was being written only to be consumed as F16 probabilities. The F32 score matrix is the largest flash-attention intermediate on this path; storing it as F16 halves its size and traffic. This builds on the coalesced softmax loads from 1aa2954, which read each score row cooperatively, so the smaller dtype pays off. Measured on an Intel Arc Pro B70 with the dispatch from the previous commit, -ngl 999 -b 4096 -ub 1024 -ctk f16 -ctv f16 -fa on: pp8192 1583.9 -> 1657.1 tok/s (+4.6%) pp64000 610.0 -> 684.5 tok/s (+12.2%) pp131072 ~354 -> 402.1 tok/s (+13.5%) tg256 at 8k context is unchanged (32.64), and the FLASH_ATTN_EXT suite shows no new failures. The exact GLM MLA backend cases pass against CPU. Adjust the ~354 baseline figure if you prefer citing only measured pairs (the 131k dispatch-only point came from the equivalent maintained build). Optionally add Assisted-by: <tool name> per the contribution guidelines since AI contributed to the change. * Revert "sycl: store MKL flash attention scores in F16" This reverts commit 265f974.
* hexagon: support tiled Q4_K GET_ROWS Assisted-by: OpenCode * properly reject Q4_K views Assisted-by: OpenCode * hexagon: support tiled Q6_K GET_ROWS Assisted-by: OpenCode
* model : add LiquidAI/d1-omni-600M decision model Assisted-by: Claude Opus 5.5 * mtmd : keep conformer GLU sigmoid on CUDA Assisted-by: Claude Opus 5.5 * server : take d1omni audio through images and input_audio, scope memory-less lfm2 to non-causal Assisted-by: Claude Opus 5.5 * common : rename decision type d1omni to lfm2-d1-omni, server : make images an alias of files Assisted-by: Claude Opus 5.5
* hexagon: Q6_K weight dequant speedup Assisted-by: OpenCode * unroll by another factor of 2 Assisted-by: OpenCode
* hex-bufs: add support for alloc_buffer_n * hex-bufs: add support for splitting large tensors into separate buffers * hex-bufs: update GGML_HEXAGON_MBUF to accept three values dyn,static,total * hex-bufs: bump dyn. default to 512MB since 128MB causes perf regressions with big MOEs * hex-run: add --no-embd-offload option to simplify command lines on devices that need it * Update scripts/snapdragon/run.py Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com> --------- Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com>
…9507) All were duplicates of the values in ggml/src/ggml-sycl/presets.hpp
* Update convert_hf_to_gguf: support Qwen3.5 embedding models * Behavior-preserving refactor for conventions. * conversion: simplify pooling comment
Co-authored-by: cwriter <cwriter@localhost>
* ggml-cuda: assign two GDN state columns per warp * ggml-cuda: use 4 GDN state columns per warp at S_v=128 * ggml-cuda: default cols_per_warp=4 * ggml-cuda: address GDN review nits
* model : support classifier_activation for rerankers Assisted-by: Claude Opus 5.5 * model : map classifier gelu to gelu_erf and accept tanh Assisted-by: Claude Opus 5.5 * model : default act_cls to tanh, ModernBERT falls back to gelu_erf Assisted-by: Claude Opus 5.5
…ggml-org#29608) Co-authored-by: cwriter <cwriter@localhost>
* sycl: fuse the delta-net alpha gate (add + unary + mul) * tests: cover the fused add + unary + mul chain * sycl: give the fused alpha gate a flat path and pin the node skip
…ggml-org#29781) * cuda : support arbitrary striding for unary ops on f16, f32, and bf16 * Remove added newline --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…gml-org#30107) The bucket search in topk_nary_search.comp started from the range [0, 0xFF800000), which ends just below the ordered-uint mapping of +inf, so +inf and NaN were never counted. A workgroup block with fewer than k countable values left the ballot empty and the shader read uninitialized shared state (hang/device lost on NVIDIA, wrong indices on AMD), and a few +inf in a block were selected without being counted, dropping real top values. Map NaN to -inf on input, start from [0, 0xFFFFFFFF) so every value is counted, and clamp the top bucket's end (2^32) instead of wrapping to 0. The k = 1 path compared float bits as signed integers, which orders negative values backwards; compare floats instead. Add test_top_k_inf to test-backend-ops: negative values, fewer than k +inf and many -inf, for k = 1, 10, 40. Assisted-by: Claude Opus 5.5
…org#26004) * server : preserve context checkpoints across slot save/restore Append the checkpoints after the packed server_tokens payload added in ggml-org#26640 and count them in n_written / n_read, so a restored slot can still roll back to a checkpoint instead of re-processing the whole prompt. * server : drop draft checkpoint data that does not match the draft context Restoring a slot saved with a different draft KV cache type aborted in load_dft(). Test-load one draft checkpoint on restore and drop the draft data if it does not fit, instead of crashing. Adds a regression test. Co-authored-by: Igor Okulist <okigan@gmail.com> * server : harden the checkpoint appendix of slot save files Bound each blob size by the bytes left in the file before allocating, open the file with UTF-8 paths on Windows like the llama state payload, fall back to full prompt re-processing when a checkpoint restored from a slot file fails to load, and replace the 1024 count cap by keeping the last n_ctx_checkpoints while reading. * server : report an incomplete checkpoint appendix as a failed slot save Return an error to the client when the appendix cannot be written, like a failed payload write, and make the oversized-blob test declare a size that cannot be allocated, so an unbounded allocation fails the test. * server : reject an empty target state in the checkpoint appendix A saved checkpoint always holds a target state, an empty blob would roll back without restoring anything. Also log with the slot id, and load the draft test model from the HF cache instead of a second download. * common : return bool from checkpoint load_tgt / load_dft A checkpoint restored from a slot file falls back to full prompt re-processing when it fails to load, a checkpoint created in memory still aborts. --------- Co-authored-by: Igor Okulist <okigan@gmail.com>
…-org#29453) * CUDA: fix CCCL version guard breaking on major version rollover The guard compared the major and minor components independently: CCCL_MAJOR_VERSION >= 3 && CCCL_MINOR_VERSION >= 1 Minor resets to 0 whenever a new major series is cut, so on CCCL 4.x this evaluates as 4 >= 3 && 0 >= 1, i.e. false. STRIDED_ITERATOR_AVAILABLE stops being defined and argsort silently falls back to the init_offsets path. Nothing warns and the build still succeeds, so the regression is a quiet performance loss rather than a compile error. CCCL already exposes the version as a single packed integer in MMMmmmpp form, which is what its own version header uses: CCCL_VERSION = MAJOR * 1000000 + MINOR * 1000 + PATCH so 3.4.3 is 3004003 and ">= 3.1" is a plain ">= 3001000". One comparison, with no component arithmetic left to get wrong. Checked against a hand-written "version >= 3.1" reference over 2.9.9, 3.0.0, 3.1.0, 3.1.99, 3.2.0, 3.4.3, 3.9.9, 3.99.99, 4.0.0, 4.2.7 and 5.0.0: no divergences. The old guard disagreed at 4.0.0 and 5.0.0. Verified on RTX 4070 (sm_89), CUDA 13.4, CCCL 3.4.3: - cmake --build build --config Release: exit 0 - test-backend-ops test -o ARGSORT -b CUDA0: 98/98 passed, CUDA0 OK Note that a passing regression test does not on its own prove the guard is still taken, since the fallback path passes too. Preprocessing the real translation unit confirms the strided-iterator branch is the one compiled in: counting_iterator is present, init_offsets is not. Signed-off-by: Heitor <heitorgm@outlook.com> * Update ggml/src/ggml-cuda/argsort.cu * Apply suggestion from @ORippler --------- Signed-off-by: Heitor <heitorgm@outlook.com> Co-authored-by: Oliver Simons <osimons@nvidia.com>
* llama : fix DFlash output head sharing Assisted-by: Codex * dflash : read tied output weights from GGUF metadata Assisted-by: Codex * llama : share tied word embedding metadata Assisted-by: Codex * llama : remove DFlash embedding head fallback Assisted-by: Codex
79 upstream commits, 50569eb..a11f57b. No carried patch landed upstream. Conflicts: - ggml/src/ggml-cuda/gated_delta_net.cu: upstream ggml-org#30087 (4 state columns per warp) rewrote the recurrent kernel that our ggml-org#22587 row-per-warp port replaced. Switched to upstream's kernel and dropped the ggml-org#22587 port: the file is now upstream + ggml-org#26001 chunked prefill + the local chunked/rollback split, rebuilt by 3-way merging ggml-org#30087 onto the pre-port file. The rollback tail still calls launch_gated_delta_net (n_seqs = 1), so it picks up ggml-org#30087's CTA sizing unchanged. - ggml/src/ggml-cuda/mmq.cu: upstream ggml-org#29953 (MMQ OOB fix) moved J selection to the host and pads src1 for J_best. Took upstream's padding; the fused ggml-org#28702 gate/up path pads for its own largest tile (GGML_CUDA_MMQ_GATE_UP_SWIGLU_J_MAX) rounded up to the largest load chunk, since its J set can exceed J_best. args take J_best, x_gate set after. - ggml/src/ggml-cuda/mmq.cuh: mmq_args takes upstream's J_best in place of ncols_opt; x_gate kept after it. - tests/test-llama-archs.cpp: kept upstream's host_experts_test and our mtp parameter of get_gguf_ctx (ggml-org#26827 test). - tools/server/server-context.cpp: kept the cost-routing slot state with upstream's comment rename; checkpoint loads take upstream's GGML_ASSERT around load_tgt/load_dft with our spec_ckpt_flags (on-device checkpoints). Fixups in auto-merged files (did not compile after the merge): - tools/server/server-context.cpp: upstream ggml-org#26004 writes checkpoints to slot save files as std::vector; added common_shared_state_buffer overloads of ckpt_read_buf/ckpt_write_buf for our ggml-org#27451 shared checkpoint payloads. - src/models/qwen35.cpp, src/models/qwen35moe.cpp: upstream ggml-org#30097 replaced the local mtp_only with nextn_flags(); derive mtp_only from nf.trunk for the ggml-org#29143 d2t trim and the local embedded trim.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merges 79 upstream commits, ggml-org
50569eb87..a11f57ba9.Merge with "Create a merge commit" (no squash, no rebase), or the next sync re-conflicts everything.
Dropped patches
IMPROVEMENT_LEDGER.mditem 2, so the GATED_DELTA_NET perf comparison below decides whether to re-port it.No carried patch landed upstream.
Conflicts
ggml/src/ggml-cuda/gated_delta_net.cuggml/src/ggml-cuda/mmq.cuggml/src/ggml-cuda/mmq.cuhmmq_argsuses upstream'sJ_bestinstead ofncols_opt;x_gatestays after it.tests/test-llama-archs.cpphost_experts_testand ourmtpparameter (ggml-org#26827 test).tools/server/server-context.cppGGML_ASSERTwith ourspec_ckpt_flags.Auto-merged but did not compile; fixed in the merge commit:
tools/server/server-context.cpp: upstream server : preserve context checkpoints across slot save/restore ggml-org/llama.cpp#26004 saves checkpoints to slot files asstd::vector. Addedcommon_shared_state_bufferoverloads for our server : share checkpoint state and harden prompt cache OOM handling ggml-org/llama.cpp#27451 shared checkpoints.src/models/qwen35.cpp,src/models/qwen35moe.cpp: upstream llama: share the nextn tensor flags between models ggml-org/llama.cpp#30097 replaced the localmtp_onlywithnextn_flags().mtp_onlyis now derived fromnf.trunkfor the d2t trims.Checks
sync.sh verify: no fork file lost its delta; size changes are the CUDA: row-per-warp kernel for GATED_DELTA_NET ggml-org/llama.cpp#22587 drop and the fixups above. Leak check clean.ctest -L main: 53/54 pass. The failure,test-tokenizers-ggml-vocabs, happens because git-lfs is missing locally, not because of the merge.test-backend-ops -b CUDA0 -o MUL_MAT,MUL_MAT_Q_GATE_UP_SWIGLU_DOWN,GATED_DELTA_NET,GATED_DELTA_NET_CACHE_FUSION,SSM_CONVall OKtest-backend-ops perf -b CUDA0 -o GATED_DELTA_NETvs master (ggml-cuda: assign four GDN state columns per warp ggml-org/llama.cpp#30087 vs CUDA: row-per-warp kernel for GATED_DELTA_NET ggml-org/llama.cpp#22587)bench_qwen38.py comparemaster vs sync: pp/tg and spec acceptance within noisedraft_n_accepted> 0