Name and Version
8183 (66d65ec), compiled this morning. Running on a Mac Mini M4 Pro 64GB.
$ llama-server --version
ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices
ggml_metal_library_init: using embedded metal library
ggml_metal_library_init: loaded in 0.006 sec
ggml_metal_rsets_init: creating a residency set collection (keep_alive = 180 s)
ggml_metal_device_init: GPU name: MTL0
ggml_metal_device_init: GPU family: MTLGPUFamilyApple9 (1009)
ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
ggml_metal_device_init: simdgroup reduction = true
ggml_metal_device_init: simdgroup matrix mul. = true
ggml_metal_device_init: has unified memory = true
ggml_metal_device_init: has bfloat = true
ggml_metal_device_init: has tensor = false
ggml_metal_device_init: use residency sets = true
ggml_metal_device_init: use shared buffers = true
ggml_metal_device_init: recommendedMaxWorkingSetSize = 55662.79 MB
version: 8183 (66d65ec29)
built with AppleClang 17.0.0.17000604 for Darwin arm64
Operating systems
Mac
Which llama.cpp modules do you know to be affected?
llama-server
Command line
llama-server --model Qwen3.5-35B-A3B-UD-Q8_K_XL.gguf --mmproj mmproj-F16-Qwen3.5-35B-A3B.gguf --ctx-size 32768 --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0.0
Problem description & steps to reproduce
I was asked by the warmup message to submit this issue. The model functions, but this message appears in the startup of llama-server:
warmup: flash attention is enabled
warmup: *****************************************************************
warmup: WARNING: the CLIP graph uses unsupported operators by the backend
warmup: the performance will be suboptimal
warmup: list of unsupported ops (backend=MTL0):
warmup: UPSCALE: type = f32, ne = [92 92 1152 1]
warmup: flash attention is enabled
warmup: please report this on github as an issue
warmup: ref: #16837 (comment)
warmup: *****************************************************************
The performance was ok (25 t/s over a small generation of 5k tokens), but noticeably slower than GLM-4.7 Flash which is about 40 t/s).
First Bad Commit
Relatively new model. My first time using it.
Relevant log output
Logs
Name and Version
8183 (66d65ec), compiled this morning. Running on a Mac Mini M4 Pro 64GB.
Operating systems
Mac
Which llama.cpp modules do you know to be affected?
llama-server
Command line
Problem description & steps to reproduce
I was asked by the warmup message to submit this issue. The model functions, but this message appears in the startup of
llama-server:warmup: flash attention is enabled
warmup: *****************************************************************
warmup: WARNING: the CLIP graph uses unsupported operators by the backend
warmup: the performance will be suboptimal
warmup: list of unsupported ops (backend=MTL0):
warmup: UPSCALE: type = f32, ne = [92 92 1152 1]
warmup: flash attention is enabled
warmup: please report this on github as an issue
warmup: ref: #16837 (comment)
warmup: *****************************************************************
The performance was ok (25 t/s over a small generation of 5k tokens), but noticeably slower than GLM-4.7 Flash which is about 40 t/s).
First Bad Commit
Relatively new model. My first time using it.
Relevant log output
Logs