Skip to content

Sync upstream master (2026-10-08, ggml-org a11f57ba9) - #18

Merged
benjigill merged 81 commits into
masterfrom
sync/upstream-2026-10-08
Oct 8, 2026
Merged

benjigill merged 81 commits into
masterfrom
sync/upstream-2026-10-08

Conversation

@benjigill

Copy link
Copy Markdown
Owner

Merges 79 upstream commits, ggml-org 50569eb87..a11f57ba9.

Merge with "Create a merge commit" (no squash, no rebase), or the next sync re-conflicts everything.

Dropped patches

No carried patch landed upstream.

Conflicts

File Resolution
ggml/src/ggml-cuda/gated_delta_net.cu Upstream ggml-org#30087 kernel; ggml-org#22587 port dropped. Rebuilt with a 3-way merge of ggml-org#30087 onto the pre-port file (upstream + ggml-org#26001 + local rollback split).
ggml/src/ggml-cuda/mmq.cu Upstream ggml-org#29953 (MMQ OOB fix): J is now picked on the host and src1 is padded for J_best. The fused ggml-org#28702 gate/up path pads for its own largest tile (64) rounded up to the largest load chunk, because its J set can exceed J_best.
ggml/src/ggml-cuda/mmq.cuh mmq_args uses upstream's J_best instead of ncols_opt; x_gate stays after it.
tests/test-llama-archs.cpp Kept upstream's host_experts_test and our mtp parameter (ggml-org#26827 test).
tools/server/server-context.cpp Kept the cost-routing state; checkpoint loads use upstream's GGML_ASSERT with our spec_ckpt_flags.

Auto-merged but did not compile; fixed in the merge commit:

Checks

tboinovski1 and others added 30 commits October 5, 2026 17:42
* hexagon: ssm-conv double-buffered DMA for prefill and decode restructuring

* hex-ssm-conv: remove divs from loops and fix trace events

* hex-dma: improved SSM_CONV dma pipeline and streamlined dma_queue

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* llama: re-pool each shared k-pool rep once

With shared cells every pool is re-pooled, and since the pooled keys
are always scattered, the pools a seq_cp shares between sequences
wrote the same rep row from several scatter entries, a data race on
the CPU backend. Mark each rep once: the sharing sequences read the
same row through pool_cells.

* llama: assert whole-sequence seq_cp in the hybrid idx memory

The recurrent state is always copied whole whatever the range, and a
k-pool cell shared by a partial copy could carry two pool groupings
with a single pooled row. Every caller copies whole sequences, so
reject partial ranges instead of supporting them.

* llama: drop the k-pool cache_safe mode

With whole-sequence seq_cp, sequences sharing cells share their pools
too, so the pooled row of a shared rep is valid for all of them. Mark
each rep once in every ubatch instead of re-pooling everything while
cells are shared, which removes the sharing scan and the stale-all
workarounds in seq_rm, state_read and state_drop. seq_cp now only
stales the destination.
* ggml-openvino: skip unselected graph branches and support DUP

Upstream ggml-org#29622 adds a mixed token/embd branch to every input
embedding graph through ggml_build_forward_select(). Its nodes are
not flagged for compute, but the backend translated them anyway,
and the DUP in that branch was unsupported, so the scheduler split
the graph and passed the embeddings across the split with a fixed
token count. The first single-token decode then failed
(test-thread-safety on CPU and GPU).

Build the OV model from the compute nodes only, and translate a
same-type contiguous DUP like CONT so the graph stays on one backend.

* ggml-openvino: make inp_scale_rows token dim dynamic

ggml-org#29622 also moves the per-token embedding scale (gemma3, gemma3n,
gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim
and pad it per chunk on the static (NPU) path.

* ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights

Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU
plugin fails to compile that form for some row counts with "clFinish,
error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on
the MUL_MAT cases added in ggml-org#29869 (e.g. m=1000, n=2, k=1024). Model
weights use a u4 zero point and are not affected.

Report these cases as unsupported on GPU until the plugin is fixed.
Op tests check support before allocating, so the check matches unbound
weights only; model loading probes with a dummy buffer and keeps its
weights on the GPU.

* ggml-openvino: create FILL in the output type

translate_fill always built an f32 constant, so an f16 FILL produced
f32 data and the copy back overran the f16 output buffer. Use the
output type for the constant.

* ggml-openvino: reject CONCAT with a quantized type

Quantized inputs are dequantized when translated, so the backend cannot
write a quantized CONCAT output. Report it as unsupported, as for CPY
to a quantized type.

* ggml-openvino: handle the single recurrent state gather of build_rs

ggml-org#29856 changed build_rs to gather all recurrent states with one GET_ROWS
on the s_copy leaf and take the ubatch and extra states as views of it.
The stateful path matched only the previous form, a GET_ROWS per view of
s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU
("is_axis_valid(axis, r)" in a Concat).

For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the
active-state gather, keep the rank-4 layout of reshapes that read a view
of it, and map the copy of the empty extra-state view to the single-slot
remainder writeback. Do not warn about the dynamic dim of empty views.

* openvino: align eltwise operand ranks to work around a GPU-plugin defect

* openvino: match the MoE fusion on the rank-3 stateful graph

* ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks

The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract
whose operand ranks differ. In gemma-3 the lower-rank operand of the
post-attention residual add is the norm output, and unsqueezing it makes
the GPU plugin compute the layer wrongly: gemma-3 returns empty answers
on GPU with stateful execution.

Skip the rewrite when the lower-rank operand is an RMS norm output.

* docs : update OpenVINO validated models

---------

Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>
…-org#30034)

With GGML_BACKEND_DL=ON, backends must be loaded explicitly before
creating models.

Assisted-by: Codex
* ggml: refactor selective expert copying to user code

* tests: enroll two models into selective expert copy test

* tests: use deepseek2 as test model

* improve comment in ggml-backend.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* cont: fix whitespace

* cont : better comments

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…gml-org#30020)

The speculative MTP init enables NextN extraction on the target and draft
contexts after both were created and their schedulers reserved. With
unmasked extraction the trunk graph keeps every token through the last
layer instead of cropping to the output rows, so the first decode
reallocates to that batch's shape and the next, wider batch trips
GGML_SCHED_DEBUG_REALLOC. Invalidate the reserve when the flags change so
the next compute re-reserves with the new graph shape.

Assisted-by: Claude
…0040)

* add text_config as fallback

* remove redundant llm_config check

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…#30017)

* mimo2 : always emit h_nextn

the other nextn-capable models set it unconditionally

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

* models : consolidate nextn row cropping into shared helpers

- replace the duplicated crop conditions and the per-model flags (narrow_early,
  crop_before_ffn, crop_last_layer, emit_h_nextn) with two helpers on llm_graph_context:
  crop_before_nextn() / crop_after_nextn()
- models that only tested embeddings_nextn_masked now share the same condition, so they
  crop the last layer before the nextn capture whenever extraction is off
- t_h_nextn is now set unconditionally in mimo2, qwen4exp and deepseek4 (as in the other
  nextn-capable models); host-side reads stay gated by cparams.embeddings_nextn

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
A 1.0 loader (e.g. Android 8.1) has no vkEnumerateInstanceVersion, so
backend init called a null pointer. Treat it like any loader under 1.2.

Fixes ggml-org#29871.

Assisted-by: Claude Opus 5.5
* ggml-cuda: chunk large BF16/FP16 to F32 conversions

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* ggml-cuda: respect dst stride in chunked cuBLAS matmul

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* ggml: fix CLAMP on non-contiguous views (CPU, CUDA)

CUDA clamped ggml_nelements values flat and ignored the view strides.
CPU addressed row j as j*nb01 and ignored nb02/nb03. Both now follow the
strides of dims 1..3; CUDA supports_op requires contiguous rows, like
Metal. test_clamp gains a non-contiguous view case.

* cuda: clamp kernel uses fastdiv for the view strides
* rpc: allow -sm tensor

* fix flush for apple rdma

* move graph_uids to rpc_dispatcher

* cont : fix conflict

* cont: stop spinning dispatcher thread

* remove meta backend change

* add TODO to simplify logic

* rpc: bump major version

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…-org#30042)

The gather path attended over the selected latents with a plain
matmul and softmax. It only ran with n_ubatch <= 16, and the flash
attention backends now skip the masked rows through n_kv_max, so
the scatter path covers every case.

Drop the gather flag, the gathered attention branch and
gather_mla_rows. set_input_kpool always maps padding to the n_kv
sentinel, and the slot mask becomes sel_mask since only the
scatter reads it.
* model: K2 Horizon gguf conversion code

* model: loading hparams and tensors in k2-horizon.cpp

* model: K2 Horizon compute graph

* model: K2 Horizon compute graph adjustment and registering tokenizers

* model: K2 Horizon chat template and accomodate safetensors naming

* unicode : add the K2-Horizon pre-tokenizer splitter

The K2-Horizon regex had no arm in unicode_regex_split_custom and fell through to the
general std::regex fallback, which fails two ways.

On MSVC std::regex rejects \p{...}, so no K2-Horizon GGUF loads on Windows at all:
llama-quantize, llama-imatrix and llama-perplexity all abort with
regex_error(error_escape) before a token is produced.

Where the fallback does compile it is still wrong. unicode_regex_split collapses each
codepoint to a single byte naming its Unicode category before matching, and U+200C/U+200D
are category Control, which has no entry in k_ucat_cpt, so both become the 0xD0 fallback
byte. The literal ‌ and ‍ alternatives in K2's regex can then never match and
every ZWNJ or ZWJ ends a letter run.

The splitter is the existing llama3 one with a single rule widened, since K2's regex
differs from llama3's only in that a letter run also takes marks, ZWNJ and ZWJ.

tests/test-unicode.cpp gains a case for this: it fails before the change with
[Amy] [ZWNJ khaham] and passes after with the run intact.

* tests: expand K2 Horizon unicode splitter coverage

* unicode: handle K2 Horizon case folding and empty input

Assisted-by: Codex

* jinja : support sequence indices in selectattr and rejectattr

Assisted-by: Codex

* model : add K2 Horizon dense and MoVA support

Includes the K2 Horizon implementation from ifm-ai/llama.cpp with converter, tensor-parallel and model save/reload fixes.

Assisted-by: Codex

* chat : support K2 Horizon reasoning and tool calls

Assisted-by: Codex

* conversion: remove obsolete K2 Aurora alias

Assisted-by: Codex

* k2-horizon: enforce response schemas and load YaRN betas

Constrain final JSON after reasoning, accept flexible JSON tool envelopes,
enforce XML dialects, and handle repeated or alternate thinking markers.
Load YaRN beta metadata instead of retaining the default values.

Add schema, streaming, continuation, and model reload regressions. Validate
CUDA and CPU builds and 0.9B, 4B, and MoVA conversation/tool round trips.

Assisted-by: Codex

* renaming template fixture

* adressing cisc follows ups

* desloppify the parser / adress aldehir comments

* clean test-chat

* remove fallback : model trained mostly on high anyway

* fix k2 attn_v_exp tn splitting and metal fusion baseline

* k2-horizon : forward expand views before sums

* k2-horizon: copy embds before group norm to fix TP

* disable tesnor parallelism

---------

Co-authored-by: Ryandito Diandaru <ryandito.diandaru@mbzuai.ac.ae>
Co-authored-by: WestWaters <mario.papaleo2013@gmail.com>
Co-authored-by: Natani L. Mayday <71436458+TaskPuppyNatani@users.noreply.github.com>
Co-authored-by: West <100190545+WestWaters@users.noreply.github.com>
Co-authored-by: aaryamonvikram <aaryamonvikram@gmail.com>
Co-authored-by: aaryamonvikram <96529820+aaryamonvikram@users.noreply.github.com>
* vocab : implement PLaMo-3 tokenizer pre-segmentation

The PLaMo-3 tokenizer inserts hard boundaries before running the Unigram
DP, around <|plamo:...|>-looking text, and around runs of at least 4
identical characters or 2 spaces. Without them llama.cpp tokenizes code
indentation and repeated punctuation differently from the reference.

Reproduce the two re.sub() passes in llm_tokenizer_plamo2 by encoding
each segment independently.

* add vocab type "plamo3"

* Update src/llama-vocab.cpp

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

* misc change

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…-org#28782)

* ggml-cuda: use per-thread stream for buffer-init padding memset

* ci : re-enable test-backend-ops -j for ROCm
The XIELU CUDA kernel template is already generic over the element
type; only the F32/F16 type assertion and the else-if dispatch were
missing. Add the nv_bfloat16 branch to the launcher, and drop the
temporary supports_op gate in ggml-cuda.cu that rejected BF16+XIELU.

test-backend-ops gains two BF16 cases ([10,5,4,3] and [512,16,1,1]).
docs/ops/CUDA.csv and docs/ops.md are regenerated; the F32 xIELU row
flips from no to yes as well, i.e. the previous record was stale.

Tested:
- Mac CPU: xIELU F32/F16/BF16, 6/6
- Mac Metal: existing F32/F16, 4/4; BF16 still unsupported
- RTX 4090 CUDA: xIELU F32/F16/BF16, 6/6
- RTX 4090 CUDA BF16-only: 2/2
- git diff --check passes
…gml-org#30067)

* hex-cpy: replace more paths with dma and simplify l2flush

* hex-cpy: use DMA in all sametype paths

* hex-cpy: rewrite the rest of the copy paths (diff type) to use dma

* hex-concat: use dma for multi-dev path which also removes the need for l2-line alignment

* hex-concat: proper support for mdev splitting

* hex-concat: cleanup ctx and kern params usage

* hex-cpy: cleanup contex and remove left-over non-dma checks

* hex-cpy: clean dma_cpy naming

* hex-cpy: proper kernel params and kernel selection

* hex-build: resolve left-over rebase conflicts

* hex-cpy: update dev guide to clarify 128 byte alignment requirement

* hex-concat: make sure we go through mdev barrier

* hex-concat: make sure to flush dma-queue

* hex-cpy/concat: cleanup kparams and vtcm layout handling

* hex-cpy: remove dead check for contig (routed to diff kernel) and update comments

* hex-dev: update developer guide based on latest changes

* hex-cpy: safe skip of noop copies

* hex-dup: route DUP to CPY
…9171)

* sycl: accelerate GLM MLA prefill with MKL flash attention

GLM-4.7 Flash uses an MLA shape with 576-wide Q/K heads, a 512-wide
V head, GQA 20, and F16 KV. The SYCL dispatcher rejects this shape
because the normal MKL flash-attention gate requires matching K/V
widths and caps the head dimension at 512, so prompt processing falls
back to the substantially slower TILE kernel.

Admit only the validated 576/576/512, GQA-20 F16 shape to the existing
MKL pipeline. Keep all other mismatched K/V shapes on their current
fallback paths.

Handle GLM's V cache as a narrower strided view of K rows. Select the
strided F16 descriptor when row stride is padded, and alias K/V
dequantization buffers only when their logical widths match. Restrict
the stride exception to a real V view sharing K's row stride.

Add the exact 576/512, GQA-20 prompt-path backend test.

On an Intel Arc Pro B70 at master e613ef2, pp8192 improves from
432.80 to 1292.29 tok/s (2.99x, +198.6%). tg256 remains unchanged
within noise at 45.67 versus 45.65 tok/s. The exact MLA test passes
and debug output confirms MKL dispatch.

* sycl: store MKL flash attention scores in F16

Keep the QK GEMM output in F16 instead of F32. The online softmax still
converts each score to F32 for its max, exponent, and sum, so the
per-element math is unchanged apart from score rounding, and the F32
matrix was being written only to be consumed as F16 probabilities.

The F32 score matrix is the largest flash-attention intermediate on this
path; storing it as F16 halves its size and traffic. This builds on the
coalesced softmax loads from 1aa2954, which read each score row
cooperatively, so the smaller dtype pays off.

Measured on an Intel Arc Pro B70 with the dispatch from the previous
commit, -ngl 999 -b 4096 -ub 1024 -ctk f16 -ctv f16 -fa on:

pp8192   1583.9 -> 1657.1 tok/s (+4.6%)
pp64000   610.0 ->  684.5 tok/s (+12.2%)
pp131072  ~354  ->  402.1 tok/s (+13.5%)

tg256 at 8k context is unchanged (32.64), and the FLASH_ATTN_EXT suite
shows no new failures. The exact GLM MLA backend cases pass against CPU.
Adjust the ~354 baseline figure if you prefer citing only measured pairs (the 131k dispatch-only point came from the equivalent maintained build). Optionally add Assisted-by: <tool name> per the contribution guidelines since AI contributed to the change.

* Revert "sycl: store MKL flash attention scores in F16"

This reverts commit 265f974.
kurquhar and others added 28 commits October 7, 2026 17:04
* hexagon: support tiled Q4_K GET_ROWS

Assisted-by: OpenCode

* properly reject Q4_K views

Assisted-by: OpenCode

* hexagon: support tiled Q6_K GET_ROWS

Assisted-by: OpenCode
* model : add LiquidAI/d1-omni-600M decision model

Assisted-by: Claude Opus 5.5

* mtmd : keep conformer GLU sigmoid on CUDA

Assisted-by: Claude Opus 5.5

* server : take d1omni audio through images and input_audio, scope memory-less lfm2 to non-causal

Assisted-by: Claude Opus 5.5

* common : rename decision type d1omni to lfm2-d1-omni, server : make images an alias of files

Assisted-by: Claude Opus 5.5
* hexagon: Q6_K weight dequant speedup

Assisted-by: OpenCode

* unroll by another factor of 2

Assisted-by: OpenCode
* hex-bufs: add support for alloc_buffer_n

* hex-bufs: add support for splitting large tensors into separate buffers

* hex-bufs: update GGML_HEXAGON_MBUF to accept three values dyn,static,total

* hex-bufs: bump dyn. default to 512MB since 128MB causes perf regressions with big MOEs

* hex-run: add --no-embd-offload option to simplify command lines on devices that need it

* Update scripts/snapdragon/run.py

Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com>

---------

Co-authored-by: Jhen-Jie Hong <iainst0409@gmail.com>
…9507)

All were duplicates of the values in ggml/src/ggml-sycl/presets.hpp
* Update convert_hf_to_gguf: support Qwen3.5 embedding models

* Behavior-preserving refactor for conventions.

* conversion: simplify pooling comment
Co-authored-by: cwriter <cwriter@localhost>
* ggml-cuda: assign two GDN state columns per warp

* ggml-cuda: use 4 GDN state columns per warp at S_v=128

* ggml-cuda: default cols_per_warp=4

* ggml-cuda: address GDN review nits
* model : support classifier_activation for rerankers

Assisted-by: Claude Opus 5.5

* model : map classifier gelu to gelu_erf and accept tanh

Assisted-by: Claude Opus 5.5

* model : default act_cls to tanh, ModernBERT falls back to gelu_erf

Assisted-by: Claude Opus 5.5
* sycl: fuse the delta-net alpha gate (add + unary + mul)

* tests: cover the fused add + unary + mul chain

* sycl: give the fused alpha gate a flat path and pin the node skip
…ggml-org#29781)

* cuda : support arbitrary striding for unary ops on f16, f32, and bf16

* Remove added newline

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…gml-org#30107)

The bucket search in topk_nary_search.comp started from the range
[0, 0xFF800000), which ends just below the ordered-uint mapping of +inf,
so +inf and NaN were never counted. A workgroup block with fewer than k
countable values left the ballot empty and the shader read uninitialized
shared state (hang/device lost on NVIDIA, wrong indices on AMD), and a few
+inf in a block were selected without being counted, dropping real top
values.

Map NaN to -inf on input, start from [0, 0xFFFFFFFF) so every value is
counted, and clamp the top bucket's end (2^32) instead of wrapping to 0.

The k = 1 path compared float bits as signed integers, which orders
negative values backwards; compare floats instead.

Add test_top_k_inf to test-backend-ops: negative values, fewer than k
+inf and many -inf, for k = 1, 10, 40.

Assisted-by: Claude Opus 5.5
…org#26004)

* server : preserve context checkpoints across slot save/restore

Append the checkpoints after the packed server_tokens payload added in ggml-org#26640
and count them in n_written / n_read, so a restored slot can still roll back to
a checkpoint instead of re-processing the whole prompt.

* server : drop draft checkpoint data that does not match the draft context

Restoring a slot saved with a different draft KV cache type aborted in
load_dft(). Test-load one draft checkpoint on restore and drop the draft
data if it does not fit, instead of crashing. Adds a regression test.

Co-authored-by: Igor Okulist <okigan@gmail.com>

* server : harden the checkpoint appendix of slot save files

Bound each blob size by the bytes left in the file before allocating, open the
file with UTF-8 paths on Windows like the llama state payload, fall back to full
prompt re-processing when a checkpoint restored from a slot file fails to load,
and replace the 1024 count cap by keeping the last n_ctx_checkpoints while reading.

* server : report an incomplete checkpoint appendix as a failed slot save

Return an error to the client when the appendix cannot be written, like a
failed payload write, and make the oversized-blob test declare a size that
cannot be allocated, so an unbounded allocation fails the test.

* server : reject an empty target state in the checkpoint appendix

A saved checkpoint always holds a target state, an empty blob would roll back
without restoring anything. Also log with the slot id, and load the draft test
model from the HF cache instead of a second download.

* common : return bool from checkpoint load_tgt / load_dft

A checkpoint restored from a slot file falls back to full prompt re-processing
when it fails to load, a checkpoint created in memory still aborts.

---------

Co-authored-by: Igor Okulist <okigan@gmail.com>
…-org#29453)

* CUDA: fix CCCL version guard breaking on major version rollover

The guard compared the major and minor components independently:

    CCCL_MAJOR_VERSION >= 3 && CCCL_MINOR_VERSION >= 1

Minor resets to 0 whenever a new major series is cut, so on CCCL 4.x
this evaluates as 4 >= 3 && 0 >= 1, i.e. false. STRIDED_ITERATOR_AVAILABLE
stops being defined and argsort silently falls back to the
init_offsets path. Nothing warns and the build still succeeds, so the
regression is a quiet performance loss rather than a compile error.

CCCL already exposes the version as a single packed integer in
MMMmmmpp form, which is what its own version header uses:

    CCCL_VERSION = MAJOR * 1000000 + MINOR * 1000 + PATCH

so 3.4.3 is 3004003 and ">= 3.1" is a plain ">= 3001000". One
comparison, with no component arithmetic left to get wrong.

Checked against a hand-written "version >= 3.1" reference over 2.9.9,
3.0.0, 3.1.0, 3.1.99, 3.2.0, 3.4.3, 3.9.9, 3.99.99, 4.0.0, 4.2.7 and
5.0.0: no divergences. The old guard disagreed at 4.0.0 and 5.0.0.

Verified on RTX 4070 (sm_89), CUDA 13.4, CCCL 3.4.3:

  - cmake --build build --config Release: exit 0
  - test-backend-ops test -o ARGSORT -b CUDA0: 98/98 passed, CUDA0 OK

Note that a passing regression test does not on its own prove the guard
is still taken, since the fallback path passes too. Preprocessing the
real translation unit confirms the strided-iterator branch is the one
compiled in: counting_iterator is present, init_offsets is not.

Signed-off-by: Heitor <heitorgm@outlook.com>

* Update ggml/src/ggml-cuda/argsort.cu

* Apply suggestion from @ORippler

---------

Signed-off-by: Heitor <heitorgm@outlook.com>
Co-authored-by: Oliver Simons <osimons@nvidia.com>
* llama : fix DFlash output head sharing

Assisted-by: Codex

* dflash : read tied output weights from GGUF metadata

Assisted-by: Codex

* llama : share tied word embedding metadata

Assisted-by: Codex

* llama : remove DFlash embedding head fallback

Assisted-by: Codex
79 upstream commits, 50569eb..a11f57b. No carried patch landed upstream.

Conflicts:
- ggml/src/ggml-cuda/gated_delta_net.cu: upstream ggml-org#30087 (4 state columns
  per warp) rewrote the recurrent kernel that our ggml-org#22587 row-per-warp port
  replaced. Switched to upstream's kernel and dropped the ggml-org#22587 port: the file
  is now upstream + ggml-org#26001 chunked prefill + the local chunked/rollback split,
  rebuilt by 3-way merging ggml-org#30087 onto the pre-port file. The rollback tail
  still calls launch_gated_delta_net (n_seqs = 1), so it picks up ggml-org#30087's
  CTA sizing unchanged.
- ggml/src/ggml-cuda/mmq.cu: upstream ggml-org#29953 (MMQ OOB fix) moved J selection
  to the host and pads src1 for J_best. Took upstream's padding; the fused
  ggml-org#28702 gate/up path pads for its own largest tile
  (GGML_CUDA_MMQ_GATE_UP_SWIGLU_J_MAX) rounded up to the largest load chunk,
  since its J set can exceed J_best. args take J_best, x_gate set after.
- ggml/src/ggml-cuda/mmq.cuh: mmq_args takes upstream's J_best in place of
  ncols_opt; x_gate kept after it.
- tests/test-llama-archs.cpp: kept upstream's host_experts_test and our mtp
  parameter of get_gguf_ctx (ggml-org#26827 test).
- tools/server/server-context.cpp: kept the cost-routing slot state with
  upstream's comment rename; checkpoint loads take upstream's GGML_ASSERT
  around load_tgt/load_dft with our spec_ckpt_flags (on-device checkpoints).

Fixups in auto-merged files (did not compile after the merge):
- tools/server/server-context.cpp: upstream ggml-org#26004 writes checkpoints to slot
  save files as std::vector; added common_shared_state_buffer overloads of
  ckpt_read_buf/ckpt_write_buf for our ggml-org#27451 shared checkpoint payloads.
- src/models/qwen35.cpp, src/models/qwen35moe.cpp: upstream ggml-org#30097 replaced
  the local mtp_only with nextn_flags(); derive mtp_only from nf.trunk for the
  ggml-org#29143 d2t trim and the local embedded trim.
@benjigill
benjigill merged commit 5dd7480 into master Oct 8, 2026
@benjigill
benjigill deleted the sync/upstream-2026-10-08 branch October 8, 2026 18:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.