EmbeddingGemma 2 (google/embeddinggemma-2) uses a new GGUF architecture, gemma-embedding2, and neither Ampere release can load it. Upstream added the architecture on 2026-10-06 in ggml-org/llama.cpp#30054, and ggml-org publishes official GGUF files at ggml-org/embeddinggemma-2-GGUF (BF16 and Q8_0, plus mmproj files for vision and audio).
Current behavior
llama-server -m embeddinggemma-2-Q8_0.gguf --embedding fails at model load on both releases:
v3.4.2 (build 7834)
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'gemma-embedding2'
v3.4.5 (build 8896)
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'gemma-embedding2'
Quantizing the official BF16 file to Q8R16 fails the same way on both:
$ llama-quantize embeddinggemma-2-BF16.gguf out.gguf Q8R16 2
llama_model_quantize: failed to quantize: unknown model architecture: 'gemma-embedding2'
The quantizer itself works. Requantizing the EmbeddingGemma 1 Q8_0 GGUF with --allow-requantize produces a valid Q8R16 file on both releases, so the only thing missing is the architecture.
With upstream llama.cpp b11454 the text GGUF loads on its own and returns embeddings (768 dimensions, 8192 context), without the mmproj files.
Request
Rebase onto an upstream build that includes #30054, or backport the architecture, and keep the Q8R16 quantizer and kernels working for it. Text embeddings are enough on my side.
Timings on Altra, for reference
OCI VM.Standard.A1.Flex (Ampere Altra, Neoverse-N1), 4 cores, Ubuntu 24.04. One document of about 440 tokens per request, 4 threads, three servers loaded at the same time and queried in rotating order, median of 5 rounds:
| Build and format |
Model |
ms per document |
| Ampere v3.4.2, Q8R16 |
EmbeddingGemma 1 (300M) |
604 |
| upstream b11454, Q8_0 (repack on) |
EmbeddingGemma 1 (300M) |
822 |
| upstream b11454, Q8_0 (repack on) |
EmbeddingGemma 2 |
2472 |
Q8R16 is 1.36x faster than upstream's repacked Q8_0 on the first generation. EmbeddingGemma 2 on upstream runs about 3x slower than the first one, and I'd like to see how much of that the Q8R16 kernels recover.
Related
v3.4.5 crashes with SIGILL on Neoverse-N1 (#21, #20), so I'm pinned to v3.4.2. A release with this architecture is only usable for me if it also runs on N1. #12 and #17 are the earlier requests to update the upstream base.
EmbeddingGemma 2 (google/embeddinggemma-2) uses a new GGUF architecture,
gemma-embedding2, and neither Ampere release can load it. Upstream added the architecture on 2026-10-06 in ggml-org/llama.cpp#30054, and ggml-org publishes official GGUF files at ggml-org/embeddinggemma-2-GGUF (BF16 and Q8_0, plus mmproj files for vision and audio).Current behavior
llama-server -m embeddinggemma-2-Q8_0.gguf --embeddingfails at model load on both releases:Quantizing the official BF16 file to Q8R16 fails the same way on both:
The quantizer itself works. Requantizing the EmbeddingGemma 1 Q8_0 GGUF with
--allow-requantizeproduces a valid Q8R16 file on both releases, so the only thing missing is the architecture.With upstream llama.cpp b11454 the text GGUF loads on its own and returns embeddings (768 dimensions, 8192 context), without the mmproj files.
Request
Rebase onto an upstream build that includes #30054, or backport the architecture, and keep the Q8R16 quantizer and kernels working for it. Text embeddings are enough on my side.
Timings on Altra, for reference
OCI VM.Standard.A1.Flex (Ampere Altra, Neoverse-N1), 4 cores, Ubuntu 24.04. One document of about 440 tokens per request, 4 threads, three servers loaded at the same time and queried in rotating order, median of 5 rounds:
Q8R16 is 1.36x faster than upstream's repacked Q8_0 on the first generation. EmbeddingGemma 2 on upstream runs about 3x slower than the first one, and I'd like to see how much of that the Q8R16 kernels recover.
Related
v3.4.5 crashes with SIGILL on Neoverse-N1 (#21, #20), so I'm pinned to v3.4.2. A release with this architecture is only usable for me if it also runs on N1. #12 and #17 are the earlier requests to update the upstream base.