Skip to content

model: support embeddinggemma2 (text+vision+audio) - #30054

Merged
ngxson merged 1 commit into
masterfrom
xsn/embd_gemma2
Oct 6, 2026
Merged

ngxson merged 1 commit into
masterfrom
xsn/embd_gemma2

Conversation

@ngxson

@ngxson ngxson commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

Support https://huggingface.co/google/embeddinggemma-2

GGUF: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF

Requirements

@ngxson
ngxson requested review from CISC and ggerganov as code owners October 6, 2026 16:02
@github-actions github-actions Bot added model Model specific testing Everything test related conversion labels Oct 6, 2026
@ngxson
ngxson merged commit 4fbc76d into master Oct 6, 2026
18 of 21 checks passed
@CISC CISC added the highlight Changes that will be highlighted in the next release notes label Oct 6, 2026
imaami pushed a commit to imaami/llama.cpp that referenced this pull request Oct 6, 2026
gagallo7 pushed a commit to gagallo7/qvac-fabric-llm.cpp that referenced this pull request Oct 8, 2026
[qvac-b11259: scale tokens after build_inp_embd, this base has no tok_scale parameter (ggml-org#29622)]

Assisted-by: Claude Code
(cherry picked from commit 4fbc76d)
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 8, 2026
CVeniamin added a commit to CVeniamin/ollama that referenced this pull request Oct 9, 2026
The pinned llama.cpp (b11351) predates ggml-org/llama.cpp#30054, which
adds the gemma-embedding2 architecture for EmbeddingGemma 2. Carry its
graph as an Ollama-owned source plus a small registration patch, the way
Laguna was carried. The only change from upstream is the input embedding
scale, written with the pinned build_inp_embd API as gemma4.cpp does.

Remove both files once LLAMA_CPP_VERSION includes that change.
aranlucas pushed a commit to aranlucas/LokalBot that referenced this pull request Oct 9, 2026
Pins the source build that adds EmbeddingGemma 2 support
(ggml-org/llama.cpp#30054). No v-tag after v0.6.0 contains it yet. The
source archive has no .git, so pass the build number explicitly and
--version reports build 11474.
brianreborn added a commit to brianreborn/familia that referenced this pull request Oct 10, 2026
graph.yaml declares runtimes (bin, tag, commit, load-tested archs). Each node names its runtime; the validator checks the binary reports the declared commit and fails if the node's GGUF arch was not verified on that runtime. Coder stays on b11374; embeddinggemma-2 (gemma-embedding2, unknown to b11374) runs on official b11539 (includes ggml-org/llama.cpp#30054), port 9951, --embeddings. Tested on miryam: 768-dim embeddings from the graph-rendered command while coder ran (RAM used 5.7/7.1 GiB).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion highlight Changes that will be highlighted in the next release notes model Model specific testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants