Skip to content

server : support typed content (vision/audio/video) input for /v1/embeddings endpoint - #29556

Merged
ngxson merged 5 commits into
ggml-org:masterfrom
timothywang21:qwen3vl-embedding-support
Sep 28, 2026
Merged

ngxson merged 5 commits into
ggml-org:masterfrom
timothywang21:qwen3vl-embedding-support

Conversation

@timothywang21

Copy link
Copy Markdown
Contributor

Overview

Right now, multimodal support for embeddings is possible but requires janky workarounds like #18665 (comment)

This PR consolidates these workarounds by implementing native multimodal support as well as introducing an OpenAI-style API for the /v1/embeddings endpoint. This new OAI-style API is the defacto standard that is used by Openrouter and other providers for multimodal embeddings. The previous legacy API is still supported.

Legacy API (still supported):

{
  "input": {
    "prompt_string": "Describe this image<__media__>",
    "multimodal_data": ["iVBORw0KGgoAAAANSUhEUg..."]
  }
}

OAI-compatible API (new), multiple content arrays in one request:

{
  "input": [
    {
      "content": [
        { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..." } },
        { "type": "text", "text": "Describe this image" }
      ]
    },
    {
      "content": [
        { "type": "text", "text": "This is a second content array you can pass in one request" }
      ]
    }
  ]
}

Incorrect (bare content array - rejected:

{
  "input": [
    { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..." } },
    { "type": "text", "text": "Describe this image" }
  ]
}

Code Changes

  1. Implemented tokenize_oai_content_array, which accepts the OpenAI-style wrapped content array format for multimodal embedding requests. Each {"content": [...]} object is one input that produces one embedding; text parts are concatenated and image_url parts are decoded via handle_media then spliced with process_mtmd_prompt.
  • NOTE: The legacy formats (plain string, token arrays, mixed arrays, and the {prompt_string, multimodal_data} object) continue to work unchanged via tokenize_input_prompts. Bare content arrays (the unwrapped shape) are rejected with a migration message.
  1. Disabled KV prefix reuse for stateless embedding/rerank tasks so that repeated inputs do not incorrectly share cached KV across requests. Each new request is done from scratch without reusing previous context from other requests. This was done for logical correctness as well as resolving a really pesky bug that I believe was caused by stale pointers from the multimodal inputs when it gets reused by another embedding/rerank request.

  2. Updated /v1/embeddings README.

Additional information

This is the second PR in support of adding full multimodal support for embedding and reranker endpoints so that models like the Qwen3-VL embedding and rerank can be fully utilized. See #28876

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Assisted-by: Opencode Qwen3.8 27B but all code was verified by me

…ing)

Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.

The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.

Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.

Assisted-by: Opencode Qwen3.8 27B
@timothywang21
timothywang21 requested a review from a team as a code owner September 28, 2026 07:08
@github-actions github-actions Bot added documentation Improvements or additions to documentation server labels Sep 28, 2026
@timothywang21
timothywang21 force-pushed the qwen3vl-embedding-support branch 2 times, most recently from d8a819a to 6ef7118 Compare September 28, 2026 07:36
@ngxson

ngxson commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

thanks. this is something I planned from #25093 (comment)

I will take over this PR for now, please don't push to it

@timothywang21

Copy link
Copy Markdown
Contributor Author

Got it. I have the multimodal rerank PR ready to go when this is finished.

@ngxson ngxson changed the title server : support multimodal input for /v1/embeddings endpoint (ie. Qwen3-VL) server : support typed content (vision/audio/video) input for /v1/embeddings endpoint Sep 28, 2026
@ngxson
ngxson merged commit 680a036 into ggml-org:master Sep 28, 2026
17 checks passed
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Sep 29, 2026
ggml-opencl.cpp: this branch's layout kept, the three upstream OpenCL changes
(ggml-org#29401 q5_K bin kernels, ggml-org#29439 q8_0 dp4a bin kernel, ggml-org#29503 bin kernel
loading condition) replayed onto it. The q5_K bin layout and the q8_0 dp4a bin
GEMM are opt-in here (GGML_OPENCL_Q5_K_BIN=1, GGML_OPENCL_Q8_0_BIN_DP4A=1):
by default they would take the spec/MTP verify widths from the cooperative-K
and narrow kernels. q5_K bin is limited to 2-D weights, q8_0 bin to N > 16.

common/speculative, server-context: this branch's llama_batch implementation
kept (tree drafting puts several seq_ids on one token, which common_batch
cannot hold); a common_batch overload of common_speculative_process converts
for the new callers, the mtmd post-decode callback takes the new embd batch,
and ggml-org#28876, ggml-org#29556 and ggml-org#29648 are applied to server-context.
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…eddings endpoint (ggml-org#29556)

* server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding)

Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.

The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.

Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.

Assisted-by: Opencode Qwen3.8 27B

* clean up comments and docs

* refactor

* add tests

* support video and audio inp

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
thom-dev-fr added a commit to thom-dev-fr/llama.cpp that referenced this pull request Oct 2, 2026
Upstream ggml-org#29556 accepts typed content (image, audio, video) in embeddings
inputs. Let embeddings and embeddings_openai resolve attachment:name in
those content arrays, as chat does, so local callers embed media without
base64. Add test-engine-vision-embeddings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
thom-dev-fr added a commit to thom-dev-fr/llama.cpp that referenced this pull request Oct 2, 2026
Upstream ggml-org#29556 accepts typed content (image, audio, video) in embeddings
inputs. Let embeddings and embeddings_openai resolve attachment:name in
those content arrays, as chat does, so local callers embed media without
base64. Add test-engine-vision-embeddings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…eddings endpoint (ggml-org#29556)

* server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding)

Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.

The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.

Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.

Assisted-by: Opencode Qwen3.8 27B

* clean up comments and docs

* refactor

* add tests

* support video and audio inp

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…eddings endpoint (ggml-org#29556)

* server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding)

Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.

The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.

Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.

Assisted-by: Opencode Qwen3.8 27B

* clean up comments and docs

* refactor

* add tests

* support video and audio inp

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
(cherry picked from commit 680a036)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants