Repository navigation
Conversation
* hex-concat: reduce pkts in gather/transpose hot loop gather directly into dst buffer, use special instruction for gather sync * hex-concat: use fastdiv replace calls to sw divide with fastpath * hex-concat: optimize DMA-HVX pipeline and add transpose helpers
* model : support classifier_pooling for ModernBERT rerankers Assisted-by: Claude Opus 5.5 * model : read classifier pooling type in load_hparams Write classifier.pooling_type from _try_set_pooling_type whenever the config has classifier_pooling, and read it in llama_model_base::load_hparams. ModernBERT falls back to mean when it is unspecified. Assisted-by: Claude Opus 5.5 * conversion : only accept cls and mean for classifier_pooling Assisted-by: Claude Opus 5.5 * model : rename classifier_pooling_type to pooling_type_cls Assisted-by: Claude Opus 5.5
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
GGML_PAD(nbytes, alignment) wraps to 0 when nbytes is within (alignment - 1) of SIZE_MAX, which silently bypassed the size overflow guard in gguf_init_from_reader. Reject the tensor before padding when nbytes + (alignment - 1) would overflow. Adds a test-gguf handcrafted case (F32, ne = [4, 2^30-1, 2^30+1, 1]) whose ggml_nbytes = 2^64 - 16 lands in the wrap window. Fails on master, passes with the guard.
* hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU) Widens ggml_hexagon_supported_activations() to accept F16 (src0/dst/src1 must agree on type), and adds F16 per-thread worker functions in act-ops.c mirroring the existing F32 workers, backed by new HVX f16 kernels (hvx_sigmoid_f16_aa, hvx_tanh_f16_aa, hvx_mul_mul_f16_aa, hvx_min_scalar_f16 family). SILU, GELU, GELU_QUICK, GEGLU, and SWIGLU are verified correct on-device (QRD8850) via test-backend-ops CPU-diffed correctness tests. SWIGLU_OAI's F16 path is code-complete and builds clean on host + all 4 DSP arch variants (v73/v75/v79/v81), but has no F16 test-case coverage in test-backend-ops and is therefore unverified on-device in this change. * hex-ops: align macros * hex-ops: minor formatting --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
cont: fix code style Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx * Add initial MTP support * Merge branch optimizations. Reduce allocated compute buffer size, speed up long context decode, fla, and slight MTP improvements. * Review driven changes, remove env vars, protect tensors * Strip MTP for initial PR * Clean up after mtp strip * Clean up after mtp strip * Update speculative.cpp * Update llama-context.h * Clean up after mtp strip * Fix tokenizer ignore merges * Improve quantization protection selection * Refactor mhc helpers, graph base * Lint Fixes * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * Skip glm5-next in model saver, fix CRLF * Skip glm5-next in sweep * Remove T4 fallback * Review cleanup * Review suggestions * Defer separate MTP gguf handling to MTP PR, drop filter * Repad n_head_kv * kpool init apply * Order by descending score * Drop guard * read kpool from hparams, clarify kpool cache flags, remove kpool_build_state(nullptr) * Add glm5-next support to model saver and add arch test fixture * Review cleanup * Kpool pooled caching clarify * Add multi stream support * Finish Rebase * Sparse FA fir DSA prefill * Const * Update llama-model.cpp to fix rebase error * gguf-py : merge tensor map entries for HC tensors * model : use build_gdn_l2_norm in GLM5_NEXT implementation * chore : remove trailing whitespace * model : use new OP precision setting API in GLM5_NEXT implementation * mtmd : use ggml_swiglu_clamp in GLM5V and apply the image token limit The two clamps around swiglu_split are what ggml_swiglu_clamp already does, so the clamp bounds collapse back to one value. GLM5V also never called set_limit_image_tokens(), so --image-max-tokens had no effect. Assisted-by: Claude Opus 5 (cherry picked from commit 46d18e1) * llama : keep the GLM5-Next k-pool layout across ubatches The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed by the positions relative to the sequence's first one, so the layout now lives on the memory and a ubatch only appends to it. A sequence edit no longer stales every pooled key either, only the ones at or after the edited position, which makes a tail seq_rm free. The pooling subgraph is built unconditionally so the graph shape no longer changes every kpool tokens, and the pool axis is folded into rows before soft_max, which otherwise exceeds the CUDA gridDim.y limit past n_kv 262144. Assisted-by: Claude Opus 5 (cherry picked from commit 5d1c40b) * model : write the GLM5-Next recurrent rollback checkpoints The conv state and the delta net state were only written to the live row, so a rollback restored whatever the checkpoint rows happened to hold. Take the same route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv groups. That also drops a state view that assumed contiguous rows. Enroll the arch in test-recurrent-state-rollback, which catches this under its garbage-filled cache pass. Assisted-by: Claude Opus 5 (cherry picked from commit 5ace37e) * llama: fix PR ggml-org#27773 test-save-load-state restore failure Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence. Assisted-by: Codex * llama: fix PR ggml-org#27773 gpu-rocm graph reallocation Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error. Assisted-by: Codex * llama : fix GLM5-Next k-pool layout staleness after edits and shared teardown Two defects in the cross-ubatch k-pool layout added by the k-pool commit: 1. Wrong results. An edited sequence only rebuilt its pool layout when its cell count changed, so if the first ubatch after an edit added back exactly as many cells as were removed, the stale position-to-cell list survived. With a unified cache and more than one sequence, where another sequence takes the freed cells, the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN). Rebuild whenever the sequence is stale, not only on a size mismatch. 2. Slowdown. "shared" mode was assumed to end only with an edit that forces a rebuild, but sharing also ends when the other sequence is removed. The survivor kept shared = true, pinning cache_safe off and re-pooling every pool on every ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence so one rebuild re-derives sharing and cache_safe returns to 1. Assisted-by: Claude Opus 5 * llama : fix build_attn_mha stream stride for non-contiguous q build_attn_mha split the batch into streams with a stream stride of q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2], when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the stride was then n_head times too large and every stream s >= 1 read another head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was wrong for every stream past the first. Unified KV and decode were unaffected (n_stream == 1, and decode takes the gather path). Other MLA models concat rope so q is contiguous and the computed value is unchanged for them. Compute the stride from the token dimension, which is identical for a contiguous q. Assisted-by: Claude Opus 5 * llama : re-derive GLM5-Next k-pool sharing on state_read/state_drop The shared-cell teardown added to seq_rm (stale every sequence when the layout has shared cells, so a survivor does not keep shared = true and pin cache_safe off) was missing from the other paths that can free shared cells: state_read and state_drop staled only the one sequence. Apply the same re-derivation there and correct the comment that claimed sharing ends only via an edit or seq_rm. Assisted-by: Claude Opus 5 * quant : drop duplicate GLM5-Next hc_ filter The hc_ name filter was listed twice in the GLM5_NEXT protection block. Assisted-by: Claude Opus 5 * glm5-next: scope K-pool cache access to indexed operations * glm5-next: keep K-pool access in hybrid index memory * glm5-next: keep mHC graph builders model-local * glm5-next: mark only touched pools per ubatch --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com> Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
* add models backend check * t4-medium for faster build
The MUSA vendor header never defined __CUDA_ARCH__, so every architecture test in the shared ggml-cuda sources evaluated to 0. Kernel bodies gated on the architecture therefore compiled to nothing, for example the q8_0 -> f16 dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to an empty body in host code. Report the newest architecture like the HIP backend does and exclude the NVIDIA-only features explicitly, as they are not usable on MUSA. Define it for device passes only: CUB uses defined(__CUDA_ARCH__) to detect device compilation, which is also how nvcc behaves. Drop the now-redundant defined(__CUDA_ARCH__) checks in the architecture comparisons: __CUDA_ARCH__ is undefined in host passes for CUDA and MUSA, and HIP defines it for every pass, so both forms select the same branch.
ggml-org#27773 adds the glm5-next arch without its rows in the Metal fusion baseline, so test-fusion --check fails on it. The rows come from test-fusion --record on an M5 Max, and --check passes 270/270.
sihanyu03
force-pushed
the
test-llama-archs-causal-attn
branch
from
September 30, 2026 08:58
3099613 to
730803e
Compare
…l-org#28381) * openvino: serve GET_ROWS on a weight view from the base Constant Resolve view_src when collecting weight Constants so a view over a quantized weight no longer becomes a dynamic typed Parameter, and fold the row offset of the view into the gather indices instead of slicing the dequantization subgraph. * openvino: lift the quantized GET_ROWS view rejection The supports_op rejection of a quantized src0 view with a nonzero offset keeps the vs0 GET_ROWS cases of ggml-org#28253 away from OpenVINO. The weight view now resolves to the base Constant with the row offset folded into the gather indices, so the rejection goes away.
…re plumbing (ggml-org#29582) * ui : type-safe API types, fetch helpers and download-ready models store plumbing Assisted-by: pi:GLM-5.3-Flash * ui : document the model list index pairing, fix an em-dash Assisted-by: pi:zai-org/GLM-5.3-Flash * Update tools/ui/src/lib/components/app/chat/index.ts Co-authored-by: Pascal <admin@serveurperso.com> --------- Co-authored-by: Pascal <admin@serveurperso.com>
…ml-org#27946) * ui : model id grammar for sidecars, quants and capability parsing Extend the shared model id parser with sidecar tokens (draft variants and auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add the tools capability to ModelCapabilities; the selector option row picks it up from the model's declared capabilities. Assisted-by: pi:GLM-5.3-Flash * ui : escape sidecar tokens in the regex alternation Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : Hugging Face Hub data layer Add HuggingFaceService and its constants/enums/types: GGUF repo search, file tree and model detail fetching, quant/sidecar filename analysis, shard-set collapsing and the llama.app catalog feed, plus an orgOf() helper on the model name utils. Assisted-by: pi:GLM-5.3-Flash * ui : strip provider tilde prefix from hub avatar urls Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : trim redundant comments in the HF data layer service Per review: drop JSDoc that restates the method name and inline comments that restate the code; keep only comments carrying non-obvious context. Assisted-by: pi:zai-org/GLM-5.3-Flash * ui : harden the HF data layer error typing, cover the helpers in tests Carries the HTTP status on retryable fetch errors instead of matching the message text. Marks expand-dependent catalog fields optional and documents the data/models index pairing. Adds table tests for the pure helpers. Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model memory-fit estimation Replace the raw runtime-memory estimate with the app's compatibility check: the smallest real Mac memory tier that fits a model file, budgeted as RAM x 0.75 minus fixed overhead with headroom on the file size. The constants move to lib; the unused runtime-memory estimate is dropped. browser-info's OS detection is exported for reuse. Assisted-by: pi:GLM-5.3-Flash * ui : cover the memory-fit and tool-use heuristics in tests Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model download pipeline Track HuggingFace downloads end to end: the server download/cancel endpoints, a status manager fed by the /models/sse download progress events, and a models-discover store holding the catalog and detail state for the discover view. Downloaded and in-flight entries are excluded from the loadable model list. Assisted-by: pi:GLM-5.3-Flash * ui : route sidecar tag lookup through the sidecars util, validate the paused list Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : shared model display primitives Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio order) out of ModelId and reuse it there, add the shared DialogConfirmDownload for destructive download actions, the discover org avatar with dark-mode inversion and the thin download progress bar, and rework ModelId badges to take thinking/tool-use support directly. Assisted-by: pi:GLM-5.3-Flash * ui : remember hub avatars that failed to load Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : render shared model row hints as native titles Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash * ui : fix badge guard for draft sidecars, keep parameter precision hasBadges now counts draft sidecar badges, so a sidecar-only model still renders. Billions keep one decimal for hub counts and stay bare for whole values. Avatar failures track the org instead of the instance, and the download progress bar no longer pulses while determinate. Assisted-by: pi:zai-org/GLM-5.3-Flash
) This commit adds an optional --add-bos token command line option to the run-org-model.py script. The motivation for this is that there are models, for example Gemma4, that explicitely set the add_bos value to true in llama-vocab.cpp even if the original model does not set this value to True. It would be nice to be able to force the models to agree on the bos token so that logit verification can proceed. Refs: ggml-org#21500
* cpu: accept BF16 in src1 of mul_mat ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in src1. The CPU backend refused that combination, so it was reported as unsupported on every backend and never compared against anything. Widen BF16 into the F32 work buffer, next to the existing packing of F32 into vec_dot_type. This is the arithmetic the Metal mat vec kernel already uses, both operands promoted to float and accumulated in float, so the two agree exactly rather than approximately. Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus three mul_mat cases with BF16 in src1. * vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16 supports_op only checked the src1 type for non contiguous tensors, so a contiguous BF16 src1 was accepted and the pipeline lookup asserted. The only BF16 src1 path is the BF16 x BF16 multiply, every other src0 type now reports the op as unsupported and the scheduler keeps it on the CPU. The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16 mat vec variants of the Metal backend, which land separately.
…g#29675) * ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) * ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops Assisted-by: Claude Opus 5.5 * CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build * ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
* llama: llama_prefetch_rows * llama: support row prefetch on Windows Apply the Windows port contributed by @praneshgo unchanged. Source: ggml-org#29599 (comment) * avoid exposing llama-mmap in model code, route via llama-impl * add windows check, only prefetch in lazy mode * cont : clean-up * cont : fix build * cont : clarify padding token for gemma4 --------- Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…l-org#29744) The fixture recycles its two blocks over 8 cache slots, so the fp16 error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU jobs of Models Backend. Two l-cycles keep every branch of the cycle loop and halve the error.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* support coerced array attributes * add tests
* convert : update to support dflash * cont : fix Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…ml-org#29722) * cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast On Windows the simple input reader sends CTRL_C_EVENT to every process attached to the console when stdin reaches EOF, killing unrelated processes such as a supervising agent. The CLI only stopped on EOF because of that self inflicted SIGINT; on POSIX, and with the advanced reader, it spins forever printing prompts. Drop the broadcast so both platforms just return an empty read, and treat an empty read as EOF in the chat loop and the model selection, since a submitted line always ends with a newline. * cli: keep the newline of a trailing "/" and stop mtmd-cli on EOF A lone "/" came back as an empty read and was taken for EOF, and mtmd-cli only stopped on EOF through the removed broadcast.
Co-authored-by: Vishal Singh <numeric-id+vishalMCE@users.noreply.github.com>
* ggml: fix integer overflow guard for zero-element tensors * ggml: validate number of elements in tensor to prevent integer overflow * ggml: fix error print
The sparse indexer mask is built with a set_rows scatter. Padded pools, absent sequences and missing tail cells all pointed to the same n_kv sentinel row, and invisible pools picked by top_k to fill the selection overlap the tail cells of the token, so several CPU threads wrote the same element (ThreadSanitizer data race in the sanitize CI). Allocate the slot mask for both selection paths and route every dead slot to its own dump row n_kv + slot. Live slots address disjoint cells, so the scatter indices of a token are unique.
* migrate the rest * test-thread-safety * rm common_batch_staged
* llama: properly handle KV on training * improve
After the device decode, flip causal_attn off, decode n_ubatch/2 then n_ubatch tokens. Both have the same node count, so a shape that depends on the flag makes the second reallocate at an unchanged graph size, which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
sihanyu03
force-pushed
the
test-llama-archs-causal-attn
branch
from
September 30, 2026 17:03
730803e to
c353025
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR adds a check to
test-llama-archsthat the graph shape does not depend oncausal_attn, which ggml-org#28751 assumes.After each row's existing decode, the check sets
causal_attnoff and decodesn_ubatch/2and thenn_ubatchtokens. If the graph changes with the flag, the first decode re-plans the compute buffers for the smaller batch (the graph size changed, so this re-plan counts as expected, withunexpected = false). The second decode ofn_ubatchtokens has the same graph shape as the first but larger tensors, so it needs a reallocation. This time the graph hasn't changed from the previous decode, so the re-plan isunexpected = trueand aborts underGGML_SCHED_NO_REALLOC.Since ggml-org#28751, every arch the test covers passes.
Additional information
qwen4expwith "unexpected graph reallocation", on CPU and CUDA. The same revert without this check passes.GGML_CUDA_DEVICES=1..4 ./build/bin/test-llama-archs -s 1) on 2xH100, and with all devices virtual on one GPU, and on CPU.Requirements