Skip to content

Misc. bug: Unsupported operation (backend=MTL0): UPSCALE: type = f32, ne = [92 92 1152 1] #20011

Description

@michaelmarziani

Name and Version

8183 (66d65ec), compiled this morning. Running on a Mac Mini M4 Pro 64GB.

$ llama-server --version
ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices
ggml_metal_library_init: using embedded metal library
ggml_metal_library_init: loaded in 0.006 sec
ggml_metal_rsets_init: creating a residency set collection (keep_alive = 180 s)
ggml_metal_device_init: GPU name:   MTL0
ggml_metal_device_init: GPU family: MTLGPUFamilyApple9  (1009)
ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4  (5002)
ggml_metal_device_init: simdgroup reduction   = true
ggml_metal_device_init: simdgroup matrix mul. = true
ggml_metal_device_init: has unified memory    = true
ggml_metal_device_init: has bfloat            = true
ggml_metal_device_init: has tensor            = false
ggml_metal_device_init: use residency sets    = true
ggml_metal_device_init: use shared buffers    = true
ggml_metal_device_init: recommendedMaxWorkingSetSize  = 55662.79 MB
version: 8183 (66d65ec29)
built with AppleClang 17.0.0.17000604 for Darwin arm64

Operating systems

Mac

Which llama.cpp modules do you know to be affected?

llama-server

Command line

llama-server --model Qwen3.5-35B-A3B-UD-Q8_K_XL.gguf --mmproj mmproj-F16-Qwen3.5-35B-A3B.gguf --ctx-size 32768 --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0.0

Problem description & steps to reproduce

I was asked by the warmup message to submit this issue. The model functions, but this message appears in the startup of llama-server:

warmup: flash attention is enabled
warmup: *****************************************************************
warmup: WARNING: the CLIP graph uses unsupported operators by the backend
warmup: the performance will be suboptimal
warmup: list of unsupported ops (backend=MTL0):
warmup: UPSCALE: type = f32, ne = [92 92 1152 1]
warmup: flash attention is enabled
warmup: please report this on github as an issue
warmup: ref: #16837 (comment)
warmup: *****************************************************************

The performance was ok (25 t/s over a small generation of 5k tokens), but noticeably slower than GLM-4.7 Flash which is about 40 t/s).

First Bad Commit

Relatively new model. My first time using it.

Relevant log output

Logs

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions