Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
113 commits
Select commit Hold shift + click to select a range
ee3ecce
metal : key the fa-vec tuned table by family instead of SKU (#29075)
forforever73 Sep 23, 2026
4e416ee
jinja : parse unary +/- before variables (#29244)
cs-fisha Sep 23, 2026
42916d8
server: fix token counting API crash on sleep (#29309)
willweimike Sep 23, 2026
dc9879c
CUDA: enable sparse-fa for dsv4 prefill (again) (#29298)
am17an Sep 23, 2026
9575389
metal: add the missing f32 x bf16 mul_mv variants (#28741)
ServeurpersoCom Sep 23, 2026
bddf826
common : keep HF cache dir as path, expose UTF-8 only for logs (#29320)
angt Sep 23, 2026
66fba63
CUDA: add a reserve to avoid spurious warning on older GCC builds (#2…
am17an Sep 23, 2026
e4e2f62
ggml : bump version to 0.25.1 (ggml/1637)
ggerganov Sep 23, 2026
177cd8c
sync : ggml
ggerganov Sep 23, 2026
7fe450e
llama.cpp : bump version to 0.5.0 (#29333)
ggerganov Sep 23, 2026
fee39dd
opencl: add A8 Q6_K non-MoE dp4a binary kernel (#29057)
shaofeiqi Sep 23, 2026
6e60f35
ci : use hf-jobs-cpu-xl runner in server sanitize workflow (#29297)
ggerganov Sep 23, 2026
d2e5458
tests: add `-b/--backend` option to test-llama-archs for testing a sp…
yomaytk Sep 23, 2026
b9ae43a
server: allow preset to set log file (#29334)
ngxson Sep 23, 2026
bd4f514
convert : allow vision target for DFlash/Dspark (#29339)
tdakhran Sep 23, 2026
013b31c
scripts : make-release-desc - link previous release in changelog titl…
ggerganov Sep 24, 2026
9710a32
hexagon: reject MUL_MAT_ID when src1 precision is F32 (#29348)
jhen0409 Sep 24, 2026
4c5957c
test-save-load-state : print a per-model results table in --models mo…
ggerganov Sep 24, 2026
2b70583
server,common : fix the GCC 12 stringop-overread false positive (agai…
angt Sep 24, 2026
f830688
model : add Ling 3.0 VL support (#29151)
aetherbird Sep 24, 2026
53ed051
cuda : add conv3d with implicit GEMM (#29137)
leejet Sep 24, 2026
3423f94
vulkan: tune KHR cooperative matrix support for Adreno GPUs (#29328)
Raman-Raje Sep 24, 2026
6b790a9
vulkan: handle misalignment in conv_2d and conv_3d (#29365)
0cc4m Sep 24, 2026
70c4e15
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (…
0cc4m Sep 24, 2026
70596c4
ci : use hf-jobs-cpu-performance, disable pytest workers (#29369)
ggerganov Sep 24, 2026
308883b
server : change default pytest workers to 4 (#29376)
danbev Sep 24, 2026
fc343a8
llama: add llama_batch_ext (#24669)
ngxson Sep 24, 2026
945064f
ui : fix missing svg use and animation elements in preview and downlo…
nandan2003 Sep 24, 2026
d954d62
adapt common
ngxson Aug 17, 2026
38b68c5
add common_batch
ngxson Aug 17, 2026
9ce186d
wip
ngxson Aug 17, 2026
8212c78
test: flush status (#28352)
kurquhar Sep 24, 2026
9c8978b
wip: spec
ngxson Sep 10, 2026
70e4407
cont
ngxson Sep 10, 2026
abb6af2
common_speculative_process
ngxson Sep 11, 2026
7c0fca3
server_batch to use common_batch
ngxson Sep 11, 2026
8ce1b95
rm some stale calls
ngxson Sep 11, 2026
c10598d
migrate mtmd
ngxson Sep 12, 2026
a72e04a
cuda : add F16 kernel support for CONV_2D_DW (#29064)
crion99 Sep 24, 2026
97a418b
hexagon: support I32 CPY and CONT (#29379)
kurquhar Sep 24, 2026
07fc586
hexagon: dynamic quantizer improvements (#29395)
kurquhar Sep 24, 2026
5cf3a35
llama-grammar: fix numeric truncation for token_id parsing (#29382)
apach301 Sep 24, 2026
a02c7f5
hexagon: handle multi-sequence in concat_2d (#29344)
jhen0409 Sep 24, 2026
bced459
sync : ggml (#29396)
ggerganov Sep 24, 2026
cdc0642
metal : optimize sparse FA + clean-up (#29377)
ggerganov Sep 24, 2026
84e76d8
metal : fix graph capture and handle empty graphs (#29390)
ggerganov Sep 24, 2026
5f1235f
handle imrope, handle return val of add()/add_embd()
ngxson Sep 24, 2026
ed319fe
hexagon: use DMA for contiguous dim1 CONCAT (#29404)
kurquhar Sep 25, 2026
4de0926
hexagon: add q5_k quant type support (#29123)
jhen0409 Sep 25, 2026
f805c57
llama : fix tensor split for fused qkv with uneven K/V head sizes (#2…
am17an Sep 25, 2026
1ab7e5a
CUDA: fuse RMS_NORM + SCALE into one kernel (#29393)
InflexCZE Sep 25, 2026
f9af9be
musa: fix PH1 (MTT S5000) operator failures and build issues (#29193)
yeahdongcn Sep 25, 2026
cd74ef6
[SYCL] support sparse FA (#28796)
arthw Sep 25, 2026
66963a8
rpc: include nb in the get_alloc_size cache key and floor the result …
Jesssullivan Sep 25, 2026
d028c69
HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m…
IMbackK Sep 25, 2026
e9f824d
llama : add `llama_prec_policy` + model-driven W4A4 path (#24364)
ynankani Sep 25, 2026
5a75f14
metal : split fa kernels into per-dtype libraries (#29329)
ggerganov Sep 25, 2026
e351231
metal: FWHT kernels for block widths above 512 (#29095)
bri-prism Sep 25, 2026
27b20ba
common : extract shared unicode path/string helpers (#29415)
angt Sep 25, 2026
d81aef1
gguf-py : TemplateProcessing has final word on add_special_token (#29…
CISC Sep 25, 2026
b248f4a
gguf-py : ByteLevel processing defaults bos/eos to False (#29422)
CISC Sep 25, 2026
e85e15c
Fixing the vulkan build issue of legacy GLSLC version that has no coo…
sliu39 Sep 25, 2026
924144f
Merge branch 'master' into xsn/llama_batch_ext_2
ngxson Sep 25, 2026
28ce6ed
add spec zeros vector
ngxson Sep 25, 2026
a25c986
opencl: add bin kernel `kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_…
shaofeiqi Sep 25, 2026
fcc8915
mtmd: fix mel preprocessor in LFM2 audio (#29403)
ykhrustalev Sep 25, 2026
4b1a27f
common,rpc : simplify fs_create_directory_with_parents() (#29432)
angt Sep 25, 2026
171e884
vendor : update cpp-httplib to 0.58.0 (#29407)
cabelo Sep 25, 2026
4e74811
hexagon: find software divide calls using binary inspection tool (#29…
trivikram-reddy1 Sep 26, 2026
9f70b2c
opencl: add A8 Q8_0 non-MoE dp4a binary kernel (#29439)
shaofeiqi Sep 26, 2026
d834d44
ggml-cpu: tiled mul_mat for k-quants (#27851)
jbooth Sep 26, 2026
965f897
polished Readme and llama-bench (#28968)
truecoder34 Sep 26, 2026
a1de614
jinja : support noncall test statements with arg (#29443)
CISC Sep 26, 2026
08618ff
llama : fix K/V and recurrent state cleanup after failed restores (#2…
CHIPMUNK-T0T Sep 26, 2026
86a24a1
jinja : fix compile error (#29468)
CISC Sep 26, 2026
81bc6b8
jinja : implement sameas test (#29448)
CISC Sep 26, 2026
2145525
Revert "Change max context length for auto-fitting with unified KV (#…
gaugarg-nv Sep 26, 2026
fcb3074
server : fix wake_fd warning on Windows (#29479)
angt Sep 26, 2026
6f856c7
cuda: add F16 input to the FWHT (#29096)
bri-prism Sep 26, 2026
694ec23
musa: build the docker images from the PH1 MUSA SDK image (#29481)
yeahdongcn Sep 26, 2026
9588757
cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (#28717)
anavp-nvidia Sep 26, 2026
2b129cc
hexagon: support for backend sampler (#29502)
max-krasnyansky Sep 27, 2026
7ac59a6
hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (#29511)
kurquhar Sep 27, 2026
85ca3b5
hrm : fix layer placement of `z_l_init` weight (#29512)
ggerganov Sep 27, 2026
187664b
llama-bench : fix OOB access of hf_file (#29515)
angt Sep 27, 2026
7fb2b08
ci : enable GGML_SCHED_DEBUG_REALLOC=1 for ctest workflows (#29514)
ggerganov Sep 27, 2026
d7fb90e
RPC: use RDMA completion channel to not spin (#29440)
am17an Sep 27, 2026
da6c28e
common : throw instead of abort on grammar without llguidance (#29516)
angt Sep 27, 2026
cea7462
vulkan: fix argsort kernel selection for Adreno (#29469)
0cc4m Sep 27, 2026
2ebd9ae
HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch si…
IMbackK Sep 27, 2026
36d7b08
CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#2…
animeshsri14 Sep 27, 2026
c829670
sycl: FWHT kernels for block widths above 512 (#29243)
bri-prism Sep 27, 2026
c9064dd
opencl: refine bin kernel loading condition (#29503)
lhez Sep 27, 2026
33c923d
jinja : add support for dict builtin (#29477)
CISC Sep 27, 2026
6fd50a4
ci : bump ty to 0.0.84 (#29529)
CISC Sep 27, 2026
9adc7f4
convert : export YaRN scaling parameters for PLaMo-3 (#29528)
tokinasin Sep 27, 2026
136887b
common : make string_split<T> throw on invalid values (#29518)
angt Sep 27, 2026
a97cce8
common : avoid side effects around params parsing (#29537)
ggerganov Sep 27, 2026
4da6337
server : allow RANK pooling batch splitting for causal LLM rerankers …
timothywang21 Sep 27, 2026
5262471
vulkan: fix wrong results when a mul_mat reads a slice of a larger ca…
ServeurpersoCom Sep 28, 2026
81ef10e
tests : fix ggml init (#29554)
ggerganov Sep 28, 2026
0c6a6a7
Enables Windows ARM64 build with MSVC cl.exe (#28362)
sarahwu185 Sep 28, 2026
ed7ac35
context : do not re-reserve the scheduler when toggling causal_attn (…
sihanyu03 Sep 28, 2026
4364bf7
metal: support left and circular padding in GGML_OP_PAD (#29561)
ServeurpersoCom Sep 28, 2026
c2a9e16
HIP: fix template skip for DKQ > 256 mfma kernels (#29559)
IMbackK Sep 28, 2026
03a667a
vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (#…
fxgsell Sep 28, 2026
f916130
ci : ignore more vgpr spills in > 256 DQK fattn kernels (#29571)
IMbackK Sep 28, 2026
6f767fe
ggml-cpu: enable tiled flash attention for non-vector-multiple head d…
SongXiaoXi Sep 28, 2026
d77dd08
tests : refactor test-recurrent-state-rollback (#29426)
ggerganov Sep 28, 2026
f00a64c
webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_te…
jbooth Sep 28, 2026
6c7a87f
common : fix HF cache paths on Windows (#29475)
angt Sep 28, 2026
2f651e6
add warning on zero fill path
ngxson Sep 28, 2026
7bbb184
Merge branch 'master' into xsn/llama_batch_ext_2
ngxson Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 2 additions & 3 deletions .devops/musa.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,10 +1,9 @@
ARG UBUNTU_VERSION=22.04
# This needs to generally match the container host's environment.
ARG MUSA_VERSION=rc4.3.0
# Target the MUSA build image
ARG BASE_MUSA_DEV_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-devel-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_DEV_CONTAINER=registry.mthreads.com/mcconline/inference/pytorch:2.9.1.post1-py3.10-musa5.2.0-mp31-devel-ubuntu${UBUNTU_VERSION}-amd64

ARG BASE_MUSA_RUN_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_RUN_CONTAINER=${BASE_MUSA_DEV_CONTAINER}

ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
Expand Down
4 changes: 3 additions & 1 deletion .github/workflows/build-apple.yml
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ concurrency:
env:
GGML_NLOOP: 3
GGML_N_THREADS: 1
GGML_SCHED_DEBUG_REALLOC: 1
LLAMA_ARG_LOG_COLORS: 1
LLAMA_ARG_LOG_PREFIX: 1
LLAMA_ARG_LOG_TIMESTAMPS: 1
Expand Down Expand Up @@ -98,7 +99,8 @@ jobs:
id: cmake_test
run: |
cd build
ctest -L main -E "test-llama-archs" --verbose --timeout 900
# ref: https://github.com/ggml-org/llama.cpp/pull/19802#issuecomment-4013704023
ctest -L main -E "test-llama-archs|test-save-load-state" --verbose --timeout 900

macos-latest-x64:
runs-on: macos-15-intel
Expand Down
1 change: 1 addition & 0 deletions .github/workflows/build-cpu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ concurrency:
env:
GGML_NLOOP: 3
GGML_N_THREADS: 1
GGML_SCHED_DEBUG_REALLOC: 1
LLAMA_ARG_LOG_COLORS: 1
LLAMA_ARG_LOG_PREFIX: 1
LLAMA_ARG_LOG_TIMESTAMPS: 1
Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/build-cuda-ubuntu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,7 @@ jobs:

musa:
runs-on: ubuntu-22.04
container: mthreads/musa:rc4.3.0-devel-ubuntu22.04-amd64
container: registry.mthreads.com/mcconline/inference/pytorch:2.9.1.post1-py3.10-musa5.2.0-mp31-devel-ubuntu22.04-amd64

steps:
- name: Clone
Expand All @@ -156,7 +156,7 @@ jobs:
id: depends
run: |
apt-get update
apt-get install -y build-essential git cmake libssl-dev jq
apt-get install -y build-essential git cmake libssl-dev jq python3-venv

- name: ccache
uses: ggml-org/ccache-action@v1.2.24
Expand All @@ -178,8 +178,8 @@ jobs:
run: |
cmake -B build -S . \
-DGGML_MUSA=ON \
-DMUSA_ARCHITECTURES=21
time cmake --build build --config Release -j $(nproc)
-DMUSA_ARCHITECTURES=31
cmake --build build --config Release -j $(nproc)

- name: ccache-buckets-save
if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/master' }}
Expand Down
1 change: 1 addition & 0 deletions .github/workflows/build-vulkan.yml
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ concurrency:
env:
GGML_NLOOP: 3
GGML_N_THREADS: 1
GGML_SCHED_DEBUG_REALLOC: 1
LLAMA_ARG_LOG_COLORS: 1
LLAMA_ARG_LOG_PREFIX: 1
LLAMA_ARG_LOG_TIMESTAMPS: 1
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/python-type-check.yml
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ jobs:
uses: actions/setup-python@v6
with:
python-version: "3.11"
pip-install: -r requirements/requirements-all.txt ty==0.0.78
pip-install: -r requirements/requirements-all.txt ty==0.0.84
# - name: Type-check with Pyright
# uses: jakebailey/pyright-action@v2
# with:
Expand Down
6 changes: 3 additions & 3 deletions .github/workflows/server-sanitize.yml
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ concurrency:

jobs:
server:
runs-on: hf-jobs-cpu-upgrade
runs-on: hf-jobs-cpu-performance

strategy:
matrix:
Expand Down Expand Up @@ -116,12 +116,12 @@ jobs:
run: |
source .venv/bin/activate
cd tools/server/tests
PYTEST_WORKERS=1 ./tests.sh
PYTEST_WORKERS=4 ./tests.sh

- name: Slow tests
id: server_integration_tests_slow
if: ${{ (github.event.schedule || github.event.inputs.slow_tests == 'true') && matrix.build_type == 'Release' }}
run: |
source .venv/bin/activate
cd tools/server/tests
PYTEST_WORKERS=1 SLOW_TESTS=1 ./tests.sh
PYTEST_WORKERS=4 SLOW_TESTS=1 ./tests.sh
1 change: 1 addition & 0 deletions .pi/gg/SYSTEM.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ General:
- Don't try to build or run the code unless you are explicitly asked to do so
- Use the `gh` CLI tool when querying PRs, issues, or other GitHub resources
- When [MODEL] is needed, first try to get it from the `PI_MODEL_NAME` env var before asking the user
- Never read the `AGENTS.md` file

Coding:
- When in doubt, always refer to the CONTRIBUTING.md file of the project
Expand Down
4 changes: 2 additions & 2 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@ include(CheckIncludeFileCXX)

### llama.cpp version
set(LLAMA_VERSION_MAJOR 0)
set(LLAMA_VERSION_MINOR 4)
set(LLAMA_VERSION_PATCH 1)
set(LLAMA_VERSION_MINOR 5)
set(LLAMA_VERSION_PATCH 0)
set(LLAMA_VERSION_BASE "${LLAMA_VERSION_MAJOR}.${LLAMA_VERSION_MINOR}.${LLAMA_VERSION_PATCH}")

# whether this is a development/nightly build
Expand Down
2 changes: 1 addition & 1 deletion CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@
/ggml/src/ggml-cann/ @ggml-org/ggml-cann
/ggml/src/ggml-common.h @ggerganov
/ggml/src/ggml-cpu/ @ggerganov
/ggml/src/ggml-cpu/iqp.* @bartowski1182
/ggml/src/ggml-cpu/tiled/ @jbooth @bartowski1182
/ggml/src/ggml-cpu/spacemit/ @alex-spacemit
/ggml/src/ggml-cuda/ @ggml-org/ggml-cuda
/ggml/src/ggml-cuda/vendors/hip.h @IMbackK
Expand Down
2 changes: 1 addition & 1 deletion ci/README-MUSA.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ docker run --privileged -it \
-v $HOME/llama.cpp/ci-cache:/ci-cache \
-v $HOME/llama.cpp/ci-results:/ci-results \
-v $PWD:/ws -w /ws \
mthreads/musa:rc4.3.0-devel-ubuntu22.04-amd64
registry.mthreads.com/mcconline/inference/pytorch:2.9.1.post1-py3.10-musa5.2.0-mp31-devel-ubuntu22.04-amd64
```

Inside the container, execute the following commands:
Expand Down
4 changes: 2 additions & 2 deletions ci/run.sh
Original file line number Diff line number Diff line change
Expand Up @@ -158,8 +158,8 @@ if [ ! -z ${GG_BUILD_WEBGPU} ]; then
fi

if [ ! -z ${GG_BUILD_MUSA} ]; then
# Use qy1 by default (MTT S80)
MUSA_ARCH=${MUSA_ARCH:-21}
# Use ph1 by default (MTT S5000)
MUSA_ARCH=${MUSA_ARCH:-31}
CMAKE_EXTRA="${CMAKE_EXTRA} -DGGML_MUSA=ON -DMUSA_ARCHITECTURES=${MUSA_ARCH}"
fi

Expand Down
19 changes: 10 additions & 9 deletions common/arg.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -2673,16 +2673,17 @@ common_params_context common_params_parser_init(common_params & params, llama_ex
params.video_ffmpeg_bin_dir = value;
}
).set_examples(mmproj_examples).set_env("LLAMA_ARG_VIDEO_FFMPEG_DIR"));
if (params.is_gen_docs || llama_supports_rpc()) {
add_opt(common_arg(
{"--rpc"}, "SERVERS",
"comma-separated list of RPC servers (host:port)",
[](common_params & params, const std::string & value) {
add_rpc_devices(value);
GGML_UNUSED(params);
add_opt(common_arg(
{"--rpc"}, "SERVERS",
"comma-separated list of RPC servers (host:port)",
[](common_params & params, const std::string & value) {
if (!llama_supports_rpc()) {
throw std::invalid_argument("RPC not supported in this build");
}
).set_env("LLAMA_ARG_RPC"));
}
add_rpc_devices(value);
GGML_UNUSED(params);
}
).set_env("LLAMA_ARG_RPC"));
add_opt(common_arg(
{"-lm", "--load-mode"}, "MODE",
"model loading mode (default: auto)\n"
Expand Down
Loading