Skip to content

Eval bug: mtmd CUDA flash attention gives wrong image embeddings on Qwen3.5 / Qwen3-VL #29629

Description

@andrewleech

Name and Version

llama.cpp 4df29be (master, 2026-08-16), CUDA 12.8, NVIDIA driver 580.x on Linux x86_64. Built with
-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=70-real -DGGML_CUDA_FORCE_MMQ=ON -DGGML_NATIVE=ON, plus a second build with
86-real.

Operating systems

Linux

GGML backends

CUDA

Hardware

AMD Ryzen 9 5900X
Tesla V100-SXM2-16GB (SM70) and RTX A2000 (SM86).

Models

OvisOCR2: https://huggingface.co/Abiray/OvisOCR2-GGUF
Qwen3.5: https://huggingface.co/unsloth/Qwen3.5-4B-GGUF

Problem description & steps to reproduce

I have been trialling OCR with OvisOCR on a table-heavy electronics datasheet, swapping between inference engines to narrow down the cause of some differences in innacuracies in the output.

On llama.cpp with CUDA flash attention on, the vision embeddings are 3.5% off compared with CUDA FA-off on a page rendered at 150 DPI; the worst token drops to 0.74 cosine on the V100 (0.64 on the A2000). CPU FA-on agrees with CUDA FA-off to about 0.2%. The gap grows with image size and can change greedy output.

Observed behaviour

Page 6 of TI's SN74LVC1G17 datasheet (electrical characteristics table), rendered at 150 DPI (1275×1650, 2,080 image tokens). Each row compares against CUDA FA-off with the same model and mmproj:

Model / backend (FA on) cos mean worst token cos mean relative L2 p99 max
Qwen3.5-4B Q8_0 + F32 mmproj, CPU 0.999993 0.9936 0.0019 0.0149 0.113
Qwen3.5-4B Q8_0 + F32 mmproj, V100 (SM70) 0.998832 0.743 0.0345 0.261 1.577
Qwen3.5-4B locally converted F16 + F32 mmproj, A2000 (SM86) 0.998627 0.638 0.0369 0.245 1.611
OvisOCR2 F16 + F16 mmproj, CPU 0.999997 0.9987 0.00143 0.0149 0.0509
OvisOCR2 F16 + F16 mmproj, V100 (SM70) 0.999104 0.820 0.0241 0.229 1.514

The V100 and CPU Qwen rows use Unsloth's ready-made GGUFs; the A2000 row predates that check and uses Qwen's weights converted with the stock converter. The OvisOCR2 files are ready-made GGUFs.

I also checked pages 7 and 17 on the V100: mean relative errors were 2.7% and 3.0%, with worst-token cosine 0.86 and 0.74 respectively.

I originally hit this with OvisOCR2, a Qwen3.5 document OCR fine-tune. On a 2592×3680 page (37,260 patches), CUDA FA-on embeddings were 14.7% off against HF transformers fp32, with a worst-token cosine of 0.46. I can't share those test pages, so the recipe below uses the public TI datasheet. The OvisOCR2 runs helped narrow it down:

  • MTMD_DEBUG_GRAPH: patch_bias, inp_pos_emb and ln1-0 match HF to four decimal places. The first difference is attn_out-0 (max absolute difference 0.047); by layer_out-11 it reaches 0.40.
  • F32 vs F16 mmproj and language weights barely move the result. The text-only path matches HF too (median top-5 logprob difference 0.0016 over 200 tokens).
  • Injecting llama.cpp's image embeddings into HF reproduces llama.cpp's next-token logprobs within 0.01 nats. HF's own embeddings choose a different greedy token.
  • build_attn in clip.cpp casts K/V to F16 and requests GGML_PREC_F32. CPU FA-on uses the same cast and stays close to the reference. The CUDA FA accumulation looks suspect for this long, non-causal, head-dim-64 attention (8.5k–37k keys). The CUDA tile and MMA kernels both contain half-precision weighted-V accumulators, even for an F32 precision request.
  • FA-off isn't usable for me at the larger document size: the 37k-patch graph requests about 64 GB of compute buffer and cudaMalloc fails.

Steps to reproduce

  1. Download the model and projector for OvisOCR2, Qwen3.5-4B, or both:
curl -fL -o OvisOCR2-F16.gguf https://huggingface.co/Abiray/OvisOCR2-GGUF/resolve/main/OvisOCR2-F16.gguf
curl -fL -o ovis-mmproj-F16.gguf https://huggingface.co/Abiray/OvisOCR2-GGUF/resolve/main/mmproj-F16.gguf
curl -fL -o Qwen3.5-4B-Q8_0.gguf https://huggingface.co/unsloth/Qwen3.5-4B-GGUF/resolve/main/Qwen3.5-4B-Q8_0.gguf
curl -fL -o qwen-mmproj-F32.gguf https://huggingface.co/unsloth/Qwen3.5-4B-GGUF/resolve/main/mmproj-F32.gguf
  1. Render page 6 at 150 DPI with pymupdf 1.28.2. My PDF SHA-256 is c76fa723fac4502423967a1aa087dd669106b52b8de620c024b26e56f5d59309; the PNG SHA-256 is 319476c76ec1d04cb250910b2b681fd73431ae8839a1c369fded85edc2fcda04.
python -m venv .venv && . .venv/bin/activate
pip install pymupdf==1.28.2 numpy
curl -fL -o sn74lvc1g17.pdf https://www.ti.com/lit/ds/symlink/sn74lvc1g17.pdf
python -c 'import pymupdf; doc = pymupdf.open("sn74lvc1g17.pdf"); doc[5].get_pixmap(matrix=pymupdf.Matrix(150/72, 150/72)).save("ti-p6-150.png")'
  1. Save this as dump_mtmd_embd.cpp and build it against libmtmd:
dump_mtmd_embd.cpp
// Dump llama.cpp (mtmd) image embeddings for one image to a raw float32 file.
// Usage: dump_mtmd_embd MODEL.gguf MMPROJ.gguf IMAGE.png OUT.f32 [fa=auto|on|off]
// Output layout: int32 n_tokens, int32 n_embd, then n_tokens*n_embd float32 (token-major).
#include "llama.h"
#include "mtmd.h"
#include "mtmd-helper.h"

#include <cstdio>
#include <cstring>
#include <string>

int main(int argc, char ** argv) {
    if (argc < 5) {
        fprintf(stderr, "usage: %s MODEL MMPROJ IMAGE OUT [fa=auto|on|off]\n", argv[0]);
        return 2;
    }
    std::string fa = argc > 5 ? argv[5] : "auto";
    llama_backend_init();
    llama_model_params mp = llama_model_default_params();
    mp.n_gpu_layers = 0;  // only the vision tower is needed; the text model supplies n_embd
    llama_model * model = llama_model_load_from_file(argv[1], mp);
    if (!model) return 1;
    mtmd_context_params cp = mtmd_context_params_default();
    cp.use_gpu = true;
    cp.print_timings = false;
    cp.image_max_tokens = 16384;
    cp.flash_attn_type = fa == "on" ? LLAMA_FLASH_ATTN_TYPE_ENABLED : fa == "off" ? LLAMA_FLASH_ATTN_TYPE_DISABLED : LLAMA_FLASH_ATTN_TYPE_AUTO;
    mtmd_context * ctx = mtmd_init_from_file(argv[2], model, cp);
    if (!ctx) return 1;
    mtmd_helper_bitmap_wrapper bw = mtmd_helper_bitmap_init_from_file(ctx, argv[3], false);
    if (!bw.bitmap) return 1;
    std::string prompt = mtmd_default_marker();
    mtmd_input_text text = {prompt.c_str(), prompt.size(), false, true};
    mtmd_input_chunks * chunks = mtmd_input_chunks_init();
    const mtmd_bitmap * bitmaps[] = {bw.bitmap};
    if (mtmd_tokenize(ctx, chunks, &text, bitmaps, 1) != 0) return 1;
    for (size_t i = 0; i < mtmd_input_chunks_size(chunks); i++) {
        const mtmd_input_chunk * chunk = mtmd_input_chunks_get(chunks, i);
        if (mtmd_input_chunk_get_type(chunk) != MTMD_INPUT_CHUNK_TYPE_IMAGE) continue;
        if (mtmd_encode_chunk(ctx, chunk) != 0) return 1;
        int32_t n_tok = (int32_t) mtmd_input_chunk_get_n_tokens(chunk);
        int32_t n_embd = llama_model_n_embd_inp(model);
        FILE * f = fopen(argv[4], "wb");
        fwrite(&n_tok, sizeof n_tok, 1, f);
        fwrite(&n_embd, sizeof n_embd, 1, f);
        fwrite(mtmd_get_output_embd(ctx), sizeof(float), (size_t) n_tok * n_embd, f);
        fclose(f);
        fprintf(stderr, "wrote %d x %d embeddings (fa=%s) to %s\n", n_tok, n_embd, fa.c_str(), argv[4]);
        return 0;
    }
    fprintf(stderr, "no image chunk\n");
    return 1;
}
g++ -O2 -std=c++17 dump_mtmd_embd.cpp -Iinclude -Iggml/include -Itools/mtmd -Lbuild/bin \
    -lmtmd -lllama -lggml -lggml-base -Wl,--allow-shlib-undefined -Wl,-rpath,$PWD/build/bin -o dump_mtmd_embd

--allow-shlib-undefined lets libggml-cuda resolve driver symbols at runtime.
4. Dump the embeddings with FA off, FA on, and CPU FA on:

./dump_mtmd_embd OvisOCR2-F16.gguf ovis-mmproj-F16.gguf ti-p6-150.png ovis-faoff.f32 off
./dump_mtmd_embd OvisOCR2-F16.gguf ovis-mmproj-F16.gguf ti-p6-150.png ovis-faon.f32 on
MTMD_BACKEND_DEVICE=CPU ./dump_mtmd_embd OvisOCR2-F16.gguf ovis-mmproj-F16.gguf ti-p6-150.png ovis-cpu.f32 on

Repeat with the ready-made Qwen3.5-4B pair to reproduce the same error independently:

./dump_mtmd_embd Qwen3.5-4B-Q8_0.gguf qwen-mmproj-F32.gguf ti-p6-150.png qwen-faoff.f32 off
./dump_mtmd_embd Qwen3.5-4B-Q8_0.gguf qwen-mmproj-F32.gguf ti-p6-150.png qwen-faon.f32 on
MTMD_BACKEND_DEVICE=CPU ./dump_mtmd_embd Qwen3.5-4B-Q8_0.gguf qwen-mmproj-F32.gguf ti-p6-150.png qwen-cpu.f32 on
  1. Save this as compare.py and compare the dumps (numpy required):
compare.py
# Compare dump_mtmd_embd outputs pairwise. Usage: python compare.py REF.f32 OTHER.f32 [...]
import sys
import numpy as np

def load(p):
    n, d = np.fromfile(p, dtype=np.int32, count=2)
    return np.fromfile(p, dtype=np.float32, offset=8).reshape(n, d).astype(np.float64)

ref = load(sys.argv[1])
for p in sys.argv[2:]:
    x = load(p)
    cos = (x * ref).sum(1) / (np.linalg.norm(x, axis=1) * np.linalg.norm(ref, axis=1))
    rel = np.linalg.norm(x - ref, axis=1) / np.linalg.norm(ref, axis=1)
    print(f"{p} vs {sys.argv[1]}: cos mean={cos.mean():.6f} min={cos.min():.6f} | "
          f"rel-err mean={rel.mean():.4f} p99={np.quantile(rel, .99):.4f} max={rel.max():.4f}")
python compare.py ovis-faoff.f32 ovis-cpu.f32 ovis-faon.f32
python compare.py qwen-faoff.f32 qwen-cpu.f32 qwen-faon.f32

On my V100, CUDA FA-on mean relative L2 error against FA-off is 2.41% for OvisOCR2 (worst-token cosine 0.820) and 3.45% for Qwen3.5-4B (worst-token cosine 0.743). CPU FA-on is at 0.14% and 0.19% respectively.

First Bad Commit

Not bisected. Vision FA was introduced in #16837] in November 2025; I haven't checked whether this error was present from the start. Earlier CUDA FA fixes addressed numerical issues #16540, #17746, #17875, but I haven't found this long-image embedding error reported there.

Relevant log output

Logs OvisOCR2 F16 + mmproj F16 on the V100, TI p6 at 150 DPI, llama.cpp `4df29be4f`. The dumper's output for the three runs (filtered to the backend, FA and encode lines):
# FA on (CUDA)
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 16144 MiB):
  Device 0: Tesla V100-SXM2-16GB, compute capability 7.0, VMM: yes, VRAM: 16144 MiB
clip_ctx: CLIP using CUDA0 backend
warmup: flash attention is enabled
image_tokens->nx = 40
image_tokens->ny = 52
clip_encode: copying image 1/1 to input buffer (nx=1280, ny=1664)
clip_encode: output embedding shape [1024, 2080, 1]
wrote 2080 x 1024 embeddings (fa=on) to ovis-faon.f32

# FA off (CUDA)
clip_ctx: CLIP using CUDA0 backend
warmup: flash attention is disabled
wrote 2080 x 1024 embeddings (fa=off) to ovis-faoff.f32

# FA on (CPU backend)
clip_ctx: CLIP using CPU backend
warmup: flash attention is enabled
wrote 2080 x 1024 embeddings (fa=on) to ovis-cpu.f32

The comparison, with CUDA FA-off as the reference:

$ python compare.py ovis-faoff.f32 ovis-cpu.f32 ovis-faon.f32
ovis-cpu.f32 vs ovis-faoff.f32: cos mean=0.999997 min=0.998708 | rel-err mean=0.0014 p99=0.0149 max=0.0509
ovis-faon.f32 vs ovis-faoff.f32: cos mean=0.999104 min=0.820091 | rel-err mean=0.0241 p99=0.2310 max=1.5139

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions