Skip to content

Support for EmbeddingGemma 2 (gemma-embedding2 architecture) #22

Description

@dhaern

EmbeddingGemma 2 (google/embeddinggemma-2) uses a new GGUF architecture, gemma-embedding2, and neither Ampere release can load it. Upstream added the architecture on 2026-10-06 in ggml-org/llama.cpp#30054, and ggml-org publishes official GGUF files at ggml-org/embeddinggemma-2-GGUF (BF16 and Q8_0, plus mmproj files for vision and audio).

Current behavior

llama-server -m embeddinggemma-2-Q8_0.gguf --embedding fails at model load on both releases:

v3.4.2 (build 7834)
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'gemma-embedding2'

v3.4.5 (build 8896)
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'gemma-embedding2'

Quantizing the official BF16 file to Q8R16 fails the same way on both:

$ llama-quantize embeddinggemma-2-BF16.gguf out.gguf Q8R16 2
llama_model_quantize: failed to quantize: unknown model architecture: 'gemma-embedding2'

The quantizer itself works. Requantizing the EmbeddingGemma 1 Q8_0 GGUF with --allow-requantize produces a valid Q8R16 file on both releases, so the only thing missing is the architecture.

With upstream llama.cpp b11454 the text GGUF loads on its own and returns embeddings (768 dimensions, 8192 context), without the mmproj files.

Request

Rebase onto an upstream build that includes #30054, or backport the architecture, and keep the Q8R16 quantizer and kernels working for it. Text embeddings are enough on my side.

Timings on Altra, for reference

OCI VM.Standard.A1.Flex (Ampere Altra, Neoverse-N1), 4 cores, Ubuntu 24.04. One document of about 440 tokens per request, 4 threads, three servers loaded at the same time and queried in rotating order, median of 5 rounds:

Build and format Model ms per document
Ampere v3.4.2, Q8R16 EmbeddingGemma 1 (300M) 604
upstream b11454, Q8_0 (repack on) EmbeddingGemma 1 (300M) 822
upstream b11454, Q8_0 (repack on) EmbeddingGemma 2 2472

Q8R16 is 1.36x faster than upstream's repacked Q8_0 on the first generation. EmbeddingGemma 2 on upstream runs about 3x slower than the first one, and I'd like to see how much of that the Q8R16 kernels recover.

Related

v3.4.5 crashes with SIGILL on Neoverse-N1 (#21, #20), so I'm pinned to v3.4.2. A release with this architecture is only usable for me if it also runs on N1. #12 and #17 are the earlier requests to update the upstream base.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions