Skip to content

server: support image+text input for embeddings (Qwen3-VL-Embedding) - #18665

Closed
ngxson wants to merge 1 commit into
ggml-org:masterfrom
ngxson:xsn/qwen3_vl_embd
Closed

ngxson wants to merge 1 commit into
ggml-org:masterfrom
ngxson:xsn/qwen3_vl_embd

Conversation

@ngxson

@ngxson ngxson commented Jan 7, 2026 •

Copy link
Copy Markdown
Collaborator

Target support: https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B

Important

the original Qwen3-VL-Embedding model is missing 1_Pooling, I don't think it's actually ready to be used unless Qwen team fixed it (I already reached out to them, but got no responses)

But currently, the model is missing 1_Pooling, so it cannot be correctly converted to GGUF

This PR aims to support mixed text+image (and maybe audio input for models supporting it) using OAI-compat content-like schema:

{
    "input": [
        {
            "type": "text",
            "text": "mixed text and image input"
        },
        {
            "type": "image",
            "image_url": {
                "url": "https://huggingface.co/ggml-org/tinygemma3-GGUF/resolve/main/test/11_truck.png"
            }
        }
    ]
}

@ggerganov

Copy link
Copy Markdown
Member

When you convert the model, try to add --sentence-transformers-dense-modules:

llama.cpp/convert_hf_to_gguf.py

Lines 10974 to 10981 in 294b2b4

parser.add_argument(
"--sentence-transformers-dense-modules", action="store_true",
help=("Whether to include sentence-transformers dense modules."
"It can be used for sentence-transformers models, like google/embeddinggemma-300m"
"Default these modules are not included.")
)

@CISC

CISC commented Jan 7, 2026

Copy link
Copy Markdown
Member

IIRC they forgot to add 1_Pooling initially on other embedding models too, since this one is not public yet maybe ask about it?

@ngxson ngxson changed the title server: support image+text input for embeddings (Qwen3-VL-Embedding) server: support image+text input for embeddings Jan 7, 2026
@ngxson

ngxson commented Jan 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

Oh sorry I didn't notice that it's private 😅 temporary closing this to keep it under the radar

@ngxson ngxson closed this Jan 7, 2026
@ngxson ngxson reopened this Jan 8, 2026
@ngxson ngxson changed the title server: support image+text input for embeddings server: support image+text input for embeddings (Qwen3-VL-Embedding) Jan 8, 2026
@Tokimorphling

Copy link
Copy Markdown

I've fixed the Qwen3-VL-Embedding issues in llama.cpp and verified the fix with regression tests. Check out the code here: https://github.com/Tokimorphling/qwen3-vl-embedding

The implementation of the Qwen3-VL series in llama.cpp seems to be problematic.

@ngxson ngxson closed this Apr 22, 2026
@ethanmc22

ethanmc22 commented Jul 4, 2026 •

Copy link
Copy Markdown

Are there any plans to merge official support for this model? Having a multimodal embedding model would open up a lot of doors with llama.cpp

@alpaim

alpaim commented Jul 4, 2026

Copy link
Copy Markdown

@ethanmc22
You can run these models already on current master, with no PR needed.

You just need a few extra steps on both the server and client sides:

Server side (one env var plus the existing flags):

LLAMA_MEDIA_MARKER=<__media__> llama-server \
 -m Qwen.Qwen3-VL-Embedding-2B.Q4_K_M.gguf \
 --mmproj mmproj-Qwen.Qwen3-VL-Embedding-2B.f16.gguf \
 --embedding --pooling last --embd-normalize 2

Without LLAMA_MEDIA_MARKER set, the server picks a random marker at startup, so a hardcoded <__media__> in your prompt will not match and every multimodal request will return HTTP 500 with Failed to tokenize prompt. The current marker is also exposed at GET /props under the media_marker field, so a client can fetch it dynamically, but pinning the env var is cleaner.

Client side, send the same shape /completions already accepts; it works on /v1/embeddings too:

{
    "input": [
        {
            "prompt_string": "Describe this image. <__media__>",
            "multimodal_data": [
                "<base64>"
            ]
        }
    ]
}

I'm using this exact quant: https://huggingface.co/DevQuasar/Qwen.Qwen3-VL-Embedding-2B-GGUF

@ethanmc22

Copy link
Copy Markdown

@alpaim

Thank you so much!! I just tried this out and it works unbelievably well!!!

I use local models for mechanical engineering work, a lot of it is done in pdf's, graphs, tables, hand writing, sketches, etc, which cant be converted to text cleanly and needs native vision to work but uploading whole 50 page pdfs as images is far too slow and burns 100k+ tokens.

Just tried it out with a quick python script and it works perfectly for retrieval tasks with all of my handwritten notes, lookup graphs/tables, pdfs, etc.

Hopefully the qwen3vl reranker gets support soon aswell, i wasnt able to get this working unfortunatly.

@timothywang21

Copy link
Copy Markdown
Contributor

Hi @ngxson I've implemented native image + text support for both the Embedding and Reranker endpoints. This was in support of using Qwen3VL-Embedding and Reranker models for RAG workflows. Related to #25921

https://github.com/timothywang21/llama.cpp-Qwen3VL-Support

Changes I made

  1. Added multimodal (image + text) embedding support with OAI compatible standards on /v1/embedding
  2. Added multimodal (image + text) rerank support with Jina/TEI compatible standards on /v1/rerank (and the other aliases)
  3. Allow batch splitting (aka. batching) for causal decoder rerankers like the Qwen3/Qwen3VL family
  • (Rationale: previous reranking models were bidirectional so all the inputs need to be within one batch to feed into the model, Qwen 3 rerankers do not need this requirement since they are just repurposed LLMs).
  1. Made embedding and rerank tasks stateless - KV prefix reuse (n_past > 0) is now disabled between discrete calls to the /v1/embeddings and /v1/rerank endpoints; every request re-processes from token 0.
  • This fixed a really pesky CUDA memory bug that showed up when multimodal was supported that skips re encoding vision chunks, which caused stale pointers that faulted the CUDA backend.
  1. Updated server README.

Let me know how you want to proceed.

timothywang21 added a commit to timothywang21/llama.cpp-Qwen3VL-Support that referenced this pull request Sep 28, 2026
…ing)

## Overview

Right now, multimodal support for embeddings is possible but requires janky workarounds like ggml-org#18665 (comment)

This PR consolidates these workarounds by implementing native multimodal support as well as introducing an OpenAI-style API for the /v1/embeddings endpoint. This new OAI-style API is the defacto standard that is used by Openrouter and other providers for multimodal embeddings. The previous legacy API is still supported.

**Legacy API (still supported):**
```json
{
  "input": {
    "prompt_string": "Describe this image<__media__>",
    "multimodal_data": ["iVBORw0KGgoAAAANSUhEUg..."]
  }
}
```

**OAI-compatible API (new), multiple content arrays in one request:**
```json
{
  "input": [
    {
      "content": [
        { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..." } },
        { "type": "text", "text": "Describe this image" }
      ]
    },
    {
      "content": [
        { "type": "text", "text": "This is a second content array you can pass in one request" }
      ]
    }
  ]
}
```

**Incorrect (bare content array - rejected with 400):**
```json
{
  "input": [
    { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..." } },
    { "type": "text", "text": "Describe this image" }
  ]
}
```

Assisted-by: Opencode Qwen3.8 27B
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants