Repository navigation
server : support typed content (vision/audio/video) input for /v1/embeddings endpoint - #29556
Merged
Merged
Conversation
…ing)
Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.
The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.
Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.
Assisted-by: Opencode Qwen3.8 27B
timothywang21
force-pushed
the
qwen3vl-embedding-support
branch
2 times, most recently
from
September 28, 2026 07:36
d8a819a to
6ef7118
Compare
Collaborator
|
thanks. this is something I planned from #25093 (comment) I will take over this PR for now, please don't push to it |
Contributor
Author
|
Got it. I have the multimodal rerank PR ready to go when this is finished. |
wanghqc
added a commit
to qualcomm/llama.cpp
that referenced
this pull request
Sep 29, 2026
ggml-opencl.cpp: this branch's layout kept, the three upstream OpenCL changes (ggml-org#29401 q5_K bin kernels, ggml-org#29439 q8_0 dp4a bin kernel, ggml-org#29503 bin kernel loading condition) replayed onto it. The q5_K bin layout and the q8_0 dp4a bin GEMM are opt-in here (GGML_OPENCL_Q5_K_BIN=1, GGML_OPENCL_Q8_0_BIN_DP4A=1): by default they would take the spec/MTP verify widths from the cooperative-K and narrow kernels. q5_K bin is limited to 2-D weights, q8_0 bin to N > 16. common/speculative, server-context: this branch's llama_batch implementation kept (tree drafting puts several seq_ids on one token, which common_batch cannot hold); a common_batch overload of common_speculative_process converts for the new callers, the mtmd post-decode callback takes the new embd batch, and ggml-org#28876, ggml-org#29556 and ggml-org#29648 are applied to server-context.
pierreguillot
pushed a commit
to Ircam-Partiels/llama.cpp
that referenced
this pull request
Oct 1, 2026
…eddings endpoint (ggml-org#29556) * server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding) Accept the OpenAI-style wrapped content array format for multimodal embedding requests. Each {"content": [...]} object is one input that produces one embedding; text parts are concatenated and image_url parts are decoded via handle_media then spliced with process_mtmd_prompt. The legacy formats (plain string, token arrays, mixed arrays, and the {prompt_string, multimodal_data} object) continue to work unchanged via tokenize_input_prompts. Bare content arrays (the unwrapped shape) are rejected with a migration message. Also disables KV prefix reuse for stateless embedding/rerank tasks so that repeated inputs do not incorrectly share cached KV across requests. Assisted-by: Opencode Qwen3.8 27B * clean up comments and docs * refactor * add tests * support video and audio inp --------- Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
thom-dev-fr
added a commit
to thom-dev-fr/llama.cpp
that referenced
this pull request
Oct 2, 2026
Upstream ggml-org#29556 accepts typed content (image, audio, video) in embeddings inputs. Let embeddings and embeddings_openai resolve attachment:name in those content arrays, as chat does, so local callers embed media without base64. Add test-engine-vision-embeddings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
thom-dev-fr
added a commit
to thom-dev-fr/llama.cpp
that referenced
this pull request
Oct 2, 2026
Upstream ggml-org#29556 accepts typed content (image, audio, video) in embeddings inputs. Let embeddings and embeddings_openai resolve attachment:name in those content arrays, as chat does, so local callers embed media without base64. Add test-engine-vision-embeddings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
…eddings endpoint (ggml-org#29556) * server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding) Accept the OpenAI-style wrapped content array format for multimodal embedding requests. Each {"content": [...]} object is one input that produces one embedding; text parts are concatenated and image_url parts are decoded via handle_media then spliced with process_mtmd_prompt. The legacy formats (plain string, token arrays, mixed arrays, and the {prompt_string, multimodal_data} object) continue to work unchanged via tokenize_input_prompts. Bare content arrays (the unwrapped shape) are rejected with a migration message. Also disables KV prefix reuse for stateless embedding/rerank tasks so that repeated inputs do not incorrectly share cached KV across requests. Assisted-by: Opencode Qwen3.8 27B * clean up comments and docs * refactor * add tests * support video and audio inp --------- Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 7, 2026
…eddings endpoint (ggml-org#29556) * server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding) Accept the OpenAI-style wrapped content array format for multimodal embedding requests. Each {"content": [...]} object is one input that produces one embedding; text parts are concatenated and image_url parts are decoded via handle_media then spliced with process_mtmd_prompt. The legacy formats (plain string, token arrays, mixed arrays, and the {prompt_string, multimodal_data} object) continue to work unchanged via tokenize_input_prompts. Bare content arrays (the unwrapped shape) are rejected with a migration message. Also disables KV prefix reuse for stateless embedding/rerank tasks so that repeated inputs do not incorrectly share cached KV across requests. Assisted-by: Opencode Qwen3.8 27B * clean up comments and docs * refactor * add tests * support video and audio inp --------- Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> (cherry picked from commit 680a036)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Right now, multimodal support for embeddings is possible but requires janky workarounds like #18665 (comment)
This PR consolidates these workarounds by implementing native multimodal support as well as introducing an OpenAI-style API for the /v1/embeddings endpoint. This new OAI-style API is the defacto standard that is used by Openrouter and other providers for multimodal embeddings. The previous legacy API is still supported.
Legacy API (still supported):
OAI-compatible API (new), multiple content arrays in one request:
Incorrect (bare content array - rejected:
Code Changes
Disabled KV prefix reuse for stateless embedding/rerank tasks so that repeated inputs do not incorrectly share cached KV across requests. Each new request is done from scratch without reusing previous context from other requests. This was done for logical correctness as well as resolving a really pesky bug that I believe was caused by stale pointers from the multimodal inputs when it gets reused by another embedding/rerank request.
Updated /v1/embeddings README.
Additional information
This is the second PR in support of adding full multimodal support for embedding and reranker endpoints so that models like the Qwen3-VL embedding and rerank can be fully utilized. See #28876
Requirements