Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueComment |
ngxson
force-pushed
the
xsn/llama_batch_ext
branch
from
September 12, 2026 09:07
045b5f9 to
e67da54
Compare
ngxson
force-pushed
the
xsn/llama_batch_ext_2
branch
from
September 12, 2026 09:14
90f4a89 to
ca70c5d
Compare
…#29075) * key the fa-vec tuned table by family instead of SKU * fall back to baseline for untuned fa-vec gpu families
* jinja : parse unary +/- before variables Lexer already emits unary_operator for -n / +n, and runtime executes unary -. Parse them at multiplicative precedence so slices like items[:-n] and GigaChat indent[:-indent_factor] work. * jinja : keep filters/tests outside unary operands Unary +/- must bind only the primary/postfix operand so -n|abs is (-n)|abs, not -(n|abs). Add unary + and filter/test regression coverage. Signed-off-by: sinksilk <785976238@qq.com> --------- Signed-off-by: sinksilk <785976238@qq.com>
* server: wake up sleeping server correctly * server: wake up sleeping server correctly (local aliases removed)
* CUDA: enable sparse-fa for dsv4 prefill (again) * CUDA: unroll the query loop of the sparse mask scan The query loop of flash_attn_mask_to_sparse_indices has a runtime trip count, which keeps the unrolled scan over the values of a lane from issuing its loads together. Template the kernel on ncols1 so the loop is bounded at compile time: batch one decodes compile to straight line code and the scan drops from 46 to 17 us at 49k columns on sparse decode shapes. * CUDA: pick the out of bounds check of the sparse mask scan in host code The query loop of the ncols1 == 8 scan keeps a runtime bound and an early exit, so it does not unroll past its first iteration. Template the kernel on whether the last group of queries is partial, decided on the host from n_queries, and hoist the column bound out of the loop: the loop becomes straight line code and the batched sparse op at 49k context drops from 586 to 244 us. --------- Co-authored-by: Pascal <admin@serveurperso.com>
ggml_conv_1d_dw builds its im2col as f32 when the kernel is bf16, then multiplies the two, so a depthwise convolution over bf16 weights asks for kernel_mul_mv_f32_bf16, which was never instantiated. The base, the _4 and the _short families are filled in next to their bf16 neighbours, inside the same runtime guard, so a device without bf16 support is unaffected.
…org#29320) Restore get_cache_directory() as fs::path as string() can be lossy on Windows Partially reverts ggml-org#29125 Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…29297) Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
…ecific backend (ggml-org#27372) * tests: add backend option to test-llama-archs * Update tests/test-llama-archs.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * remove extra space --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Resolve the target arch with get_model_architecture so vision targets (e.g. Lfm2VlForConditionalGeneration) map to their text model for the vocab. Fix double rope reorder for LFM2/LFM2.5 DSpark drafters
ggml-org#29336) make-release-desc.sh now emits "Changelog since [vX.Y.Z](<repo>/releases/tag/vX.Y.Z)" instead of a plain version string, so the release notes link back to the previous release. The repo URL is derived from the origin remote (SSH or HTTPS); if it cannot be resolved (local run without origin), the title falls back to plain text. Assisted-by: pi:llama.cpp/Qwen3.8-27B
…de (ggml-org#29316) * test-save-load-state : print a per-model results table in --models mode in --models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with one column per test, colored PASS/FAIL/SKIP cells, row by row. - run_save_load_tests_for_model returns a test_suite with a dynamic std::vector<test_status> and continues past failures: tests 3-5 are SKIPped when the baseline (test 1) fails, model init failure skips all - per-test token dumps, test headers and PASS lines are demoted to LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode - the table header/rows derive their columns from test_names; the model name is printed and flushed before the suite runs so the model currently in flight is always visible - single-model output and exit codes are unchanged Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : print example usage on -h add a print_usage callback passed to common_params_parse, so -h/--help also shows example commands for the tool-specific --models option and the -lv verbosity level Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : remove comments ref: ggml-org#29316 Assisted-by: pi:llama.cpp/Qwen3.8-27B
…n) (ggml-org#29325) Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* model : fold Ling 3.0 VL into the BailingMoeV3 architecture Assisted-by: Scout * model : keep shared NORM rope list intact when gating bailingmoe3 on mrope sections --------- Co-authored-by: aetherbird <aetherbird@users.noreply.github.com>
* cuda : add conv3d with implicit GEMM * cuda : refine conv3d implicit GEMM and handle empty kernels
…#29328) * Enable coopmat support for Vulkan backend * Fixed the mul_mat_s * Removed the debug statement
* vulkan: handle misalignment in conv_2d and conv_3d * fix test-backend-ops print
…gml-org#27952) * vulkan: add int8 coopmat quantized matmul shader * apply scales inline * use scalar sums * probe and directly access coopmat values instead of going through shmem * add q8_0 support * add BK_STEP to shader, default to 2 * use larger workgroups * double buffering * preload scales * coopmat load first, then wmma * use float for scales * add faster RDNA int->float conversion * workgroup scheduling for cache proximity * clean up * use wave32 * restructure for vgpr use * skip computation for inactive tiles * only force subgroup size 32 on AMD RDNA * use BK_STEP 4 * fix compilation * move quant-specific prefetch function out of main file * add q4_1, q5_0, q5_1 support * restructure mmq cm1 functions * enable mul_mat_id support * fix segfault * fix mul_mat_id bug * support iq4_nl and mxfp4 * remove elem row/col fast path, invalid for RDNA4 * use shmem arrays for LUTs * use 4-byte loads where possible * add q3_k, q4_k, q5_k, q6_k and nvfp4 support * fix l warptile * improve performance * improve performance * improvements * dedup b scales * merge shmem arrays * undo uint8_t, gate to RDNA3/4 * add RDNA4 architecture, use for hardcoded coopmat elem thread access, set BK_STEP back to 4 * improve offset application * clean up * fix iq4_nl and nvfp4 performance * rdna4 tuning * use BK_STEP 2 on MUL_MAT_ID * adapt to upstream changes * fix shmem support function, clean up comments * fix warptile logic Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> * vulkan: add IQ4_XS support to the coopmat1 integer matmul shader (ggml-org#28440) Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8 and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling. Assisted-by: OpenAI Codex Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * avoid compiling f16 acc shader variants --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…29514) * ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows * cont : metal paravirtual device is not compatible
* RPC: use RDMA completion queue to not spin * add TODO for apple RDMA
…rg#29516) Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…zes (ggml-org#28907) * HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes * CI: hip-quality-check: ignore spills for very large mfma mma kernels
The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through to the default case and run as a dense GEMM against the materialized rotation tensor, correct but O(n^2) instead of O(n log n). fwht_kernel_wide runs one row per work-group instead of per sub-group, so each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below the sub-group width still shuffle; those up to the work-group width go through work-group local memory; the rest stay in registers. Same butterfly and sign convention as the existing narrow kernel. ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a GPU-labeled platform to exist and throws before any op-level test can run, so test-backend-ops could not be exercised on this box (a GPU-less pod) even via the CPU device. Verified instead with a standalone harness: the same kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1, OpenCL CPU backend), checked against an independent recursive-doubling Hadamard reference, cross-validated by first running the existing unmodified narrow kernel through the identical harness and confirming it passes (rules out a reference-convention bug before trusting a pass on the new code). Random-input results for all four widths, single- and multi-row: N=1024 NT=256 rows=1 max_abs_err=1.7e-07 max_rel_err=4.9e-04 PASS N=2048 NT=256 rows=1 max_abs_err=1.9e-07 max_rel_err=2.0e-04 PASS N=4096 NT=256 rows=1 max_abs_err=2.0e-07 max_rel_err=1.4e-04 PASS N=8192 NT=256 rows=1 max_abs_err=2.5e-07 max_rel_err=3.8e-03 PASS N=1024 NT=256 rows=7 max_abs_err=2.4e-07 max_rel_err=1.0e-03 PASS N=2048 NT=256 rows=5 max_abs_err=3.0e-07 max_rel_err=9.4e-04 PASS N=4096 NT=256 rows=3 max_abs_err=2.7e-07 max_rel_err=1.7e-03 PASS N=8192 NT=256 rows=2 max_abs_err=2.5e-07 max_rel_err=1.9e-03 PASS This covers the kernel algorithm itself; it does not exercise the ggml dispatch/supports_op integration end to end, which needs a real GPU (or a SYCL GPU plugin) to get past backend registration. test-backend-ops build is verified: fwht.cpp recompiles with zero warnings as part of ggml-sycl.
* add support for dict builtin * add tests
* bump ty to 0.0.84 * fix assertion bug caught by ty
Recent PLaMo-3 models use YaRN, while some earlier PLaMo-3 models do not. The recent PLaMo-3 store their YaRN settings as flat config keys (rope_scaling_factor, initial_context_length) and build the dict at runtime in Plamo3Config.rope_parameters. The current converter misses these settings and writes plain RoPE metadata to GGUF. Mirror the runtime settings into rope_parameters so the corresponding rope.scaling.* is written to GGUF.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
- register --rpc unconditionally and call llama_supports_rpc() only from its handler - print server "initialization ..." log after args are parsed Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
…(ie. Qwen3 and Qwen3-VL) (ggml-org#28876) * server : allow splitting RANK pooling for causal LLM rerankers Rerank models fall into two categories: bidirectional cross-encoders (BERT, etc.) that require all tokens in a single physical batch, and causal LLMs repurposed as rerankers (Qwen3, Qwen3-VL) that can use chunked prefill like any other decoder. Previously the server rejected all RANK-pooling inputs larger than n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL arch checks to determine last-token pooling. This broke long-document and multimodal reranking for causal models. Fix: expose llama_get_causal_attn(ctx) so the server can check the effective runtime attention type (reflecting any --attention override or set_causal_attn call). Also expose llama_model_is_causal(model) for querying the static architectural property from GGUF metadata. can_split() now permits chunked prefill for RANK pooling when the context is causal. The graph builder's inline arch check is replaced with the same cparams.causal_attn predicate, removing the duplication. Assisted-by: Opencode/Qwen3.8-27B * remove unused llama_model_is_causal, fix whitespace Assisted-by: opencode --------- Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
…che (ggml-org#28956) * vulkan: read the batch stride of an in place src0 from nb[2] A dim01 contiguous tensor can still be a view whose batches are strided by more than ne[1] rows, the first rows of a KV cache for example. Both the mat-vec and the matrix paths read such a tensor in place but passed ne00*ne01 as the batch stride, so every head past the first read the wrong rows. The same applies to src1. The stride now comes from nb[2] whenever the tensor is used in place; the value is unchanged for a contiguous tensor. test-backend-ops gets an m_v parameter on test_mul_mat, the number of rows of a in memory, and two cases at the shapes of a decoder self attention over a cache. * vulkan: size the in place A and B ranges by their strided extent The matrix path bound src0 and src1 to the shader with a range of elements times type size, which ends before the batches of a strided view. Pipelines with bounded access read zero past that range, so the same view that the mat-vec path already handles gave wrong results on Intel and on NVIDIA without coopmat2. The range now comes from ggml_nbytes when the tensor is read in place. * vulkan: address review from jeffbolznv Bind the in place A and B of the matrix path with ggml_vk_subbuffer, which spans to the end of the buffer, so a strided view is in range without computing its extent. mul_mat_id reads the batch stride of an in place src0 and src1 with the same helper as mul_mat. test_mul_mat_id gets an m_v parameter, the number of rows of as in memory, and a case whose experts are strided by more rows than it uses. * vulkan: read the batch stride of an in place src0 in mul_mat_vec_id The single token path of mul_mat_id passed ne00*ne01 as the batch stride of A, so a strided expert view read the wrong rows. The stride now comes from ggml_vk_batch_stride like the other three paths, and src1 follows the same rule. test_mul_mat_id gets a single token case over the strided view. * vulkan: address review from jeffbolznv The batch stride of an in place tensor is taken from nb[2] as nb[2] / type_size * block_size, which holds when nb[2] is padded and not a multiple of nb[1]. A test_mul_mat case with a padded batch stride covers it. * vulkan: keep the A and B ranges exact in mul_mm The quantized A loads of mul_mm carry no row bound and rely on the descriptor range to read zeros past the last row of a partial tile. Binding A and B up to the end of the buffer let those tiles read the leftovers of a previous node and hung the NVFP4 mul_mm on NVIDIA without coopmat2. The range is the strided extent of a tensor read in place and the staged size otherwise.
* tests : init ggml for test-recurrent-state-rollback * cont : same for test-save-load-state * cont : add to test-state-restore-fragmented + add TODOs
* can reproduce the issue vlad sees * fix fma issue * drop volatile * fix volatile runtime task * add arm flag if needed * fix hsum compile error * fix syntax in quants * strengthen sve probing * make the syntax fixes one liners * remove debug code * formatting * remove macro for float * drive down gcc instruction count * support armec * fix CI comments address CI comments fix cross compile issue remove warning fix style and fix fma probing fix style * add documentation * update documentation
…gml-org#28751) * context : do not re-reserve the scheduler when toggling causal_attn `llama_context::set_causal_attn()` marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive `sched_reserve()` passes per image. This is especially slow for multi-image or video inputs. The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below). The re-reserve is unnecessary in this case because `causal_attn` only changes the values written to KQ mask, not tensor shapes or any other buffer sizes. Note: `causal_attn` is a graph reuse key (`llm_graph_params` via `cparams`), so a new graph is built regardless of `sched_need_reserve`, so this doesn't change the graph rebuilding behaviour. llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images, cache_prompt=false, prompt_ms median of 3 (before -> after): | images | config | H200 before -> after | RTX 4090 before -> after | |-|-|-|-| | 1 | `-c 8192 -ub 512` | 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) | | 24 | `-c 8192 -ub 512` | 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) | | 24 | `-c 32768 -ub 2048` | 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) | Generated output remains identical before and after. * qwen4exp : make the indexer bias shape independent of causal_attn The block/cell bias path was selected on cparams.causal_attn, so the causal and non-causal graphs differed in tensor shapes and ops. With the re-reserve removed (previous commit), a runtime flip resulted in reallocating the compute buffers, which would fail under GGML_SCHED_NO_REALLOC. This commit selects the block path from the mask shape only, independent of causal_attn. causal_attn is instead passed to set_input_qsa. causal_attn is fixed per graph as it's part of the reuse key. Causal values are unchanged. Non-causal values now follow the reference rule, where every visible block competes on score and only unpooled cells are always selected. * context : state the causal_attn shape rule in the comment * cont : add TODOs --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* metal: support left and circular padding in GGML_OP_PAD Align Metal with CPU, CUDA and Vulkan: shift the source coordinates by the left paddings, wrap them around with the same wrap_around when circular, and read the source through nb00, which also fixes a right padding of a permuted source. A test case covers it. Drop the f32_4 kernel: its selection is disabled as slower, and it fails two pad cases once enabled. * metal: use a function constant for the circular pad variant Address review from ggerganov: replace the bool template with FC_PAD, as FC_upscale_aa does, so the pad kernel is compiled once and specialized per pipeline.
…ims on x86 (ggml-org#29423) * ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 * add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail * ggml-cpu: fix FA softcap handling for padded KV tiles
* tests : use llama_context_ptr in test-recurrent-state-rollback Replace raw llama_context pointers with llama_context_ptr and drop the manual llama_free calls and cleanup lambda. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : run test-recurrent-state-rollback over all dummy models Add a --models DIR mode that mirrors test-save-load-state: iterate every dummy model, report PASS/FAIL/SKIP in a table and fail only when a model fails. Register a single ctest entry with ARGS --models instead of the four per-model registrations. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * cont : fix typo * metal : allow fusing 0-element nodes to keep graph packing shape-independent The fusion packing in ggml_metal_fusion_max excluded 0-element tensors and the topk_moe/moe_reduce checks rejected n_tokens == 0, so graphs decoding batches with no outputs packed differently from the worst-case reserved graph. The Metal optimizer then reordered the nodes differently and ggml_gallocr_needs_realloc failed on the layout mismatch, forcing an unexpected graph re-reserve (caught by GGML_SCHED_DEBUG_REALLOC). Treat empty tensors like their non-empty counterparts: match them in the pattern sequence and only reject genuinely malformed shapes. Fused kernels dispatch zero threadgroups for empty graphs, which is a legal no-op. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : run test_multi_seq_split_replay as a separate test test_multi_seq_split_replay was invoked at the end of test_rollback, so its result was folded into the rollback status and it only ran when the rollback part passed. Give it its own test_status return, run both tests independently over both cache fills via a shared run_tests helper, and report them as separate rollback / split replay columns in the --models table with per-test summaries. The exit code fails when either test fails. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : loosen the split replay nmse bound to 1e-4 test-generate-models seeds its weights from std::random_device, and some generated lfm2 models drift up to ~1.7e-5 nmse on the split replay due to rounding noise, tripping the previous 1e-5 bound. Raise the bound to 1e-4 so the random generations stop flaking. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : reuse run_tests_for_model in single-model mode The single-model path duplicated the model init and the non-recurrent check from run_tests_for_model; route it through the shared helper instead. Model load failures now return FAIL rather than SKIP so that --model with a broken file still exits non-zero, and the helper loads with model_only like the --models loop does since the tests create their own contexts. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
…sor (ggml-org#29471) * Fix: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor * Clang formatting
Supersedes ggml-org#29158 Signed-off-by: Adrien Gallouët <angt@huggingface.co>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Additional information
Requirements