Skip to content

model: Granite Speech Plus - #24818

Merged
ngxson merged 5 commits into
ggml-org:masterfrom
gabe-l-hart:GraniteSpeechPlus
Jun 23, 2026
Merged

ngxson merged 5 commits into
ggml-org:masterfrom
gabe-l-hart:GraniteSpeechPlus

Conversation

@gabe-l-hart

Copy link
Copy Markdown
Collaborator

Overview

This PR adds support for ibm-granite/granite-speech-4.1-2b-plus. This model is a slight tweak on the GraniteSpeechForConditionalGeneration architecture that extracts multiple encoder layers and concatenates them into the projected token stream.

Additional information

Testing

# Download speech plus from HF
uvx --with transformers hf download ibm-granite/granite-speech-4.1-2b-plus --local-dir ibm-granite/granite-speech-4.1-2b-plus

# Convert
uvx --with-requirements requirements/requirements-convert_hf_to_gguf.txt --extra-index-url https://download.pytorch.org/whl/cpu python convert_hf_to_gguf.py ibm-granite/granite-speech-4.1-2b-plus
uvx --with-requirements requirements/requirements-convert_hf_to_gguf.txt --extra-index-url https://download.pytorch.org/whl/cpu python convert_hf_to_gguf.py ibm-granite/granite-speech-4.1-2b-plus --mmproj

# Grab the sample audio
wget https://huggingface.co/ibm-granite/granite-speech-4.1-2b/blob/main/multilingual_sample.wav

# Build
mkdir -p build && cd build
cmake .. && make -j

# Test inference
./bin/llama-mtmd-cli -m ../ibm-granite/granite-speech-4.1-2b-plus/granite-speech-4.1-2B-plus-BF16.gguf --mmproj ../ibm-granite/granite-speech-4.1-2b-plus/mmproj-granite-speech-4.1-2b-plus-BF16.gguf --audio ../multilingual_sample.wav -p "transcribe this audio to a written format" --jinja

Requirements

AI Disclosure

AI used in planning and building. All code reviewed and refined by hand.

git-ai-stats
╔══════════════════════════════════════════════════════════╗
║           GIT AI USAGE ANALYSIS                          ║
╚══════════════════════════════════════════════════════════╝

📊 COMMITS BY AGENT

--- Aggregate ---
Commits                        |      Count
---------------------------------------------
IBM Bob, OpenCode + Qwen3.6-35b |          2
---------------------------------------------
TOTAL                          |          2

📊 COMMITS BY USAGE TYPE

--- Aggregate ---
Commits                        |      Count
---------------------------------------------
none                           |          0
draft                          |          1
full                           |          1
---------------------------------------------
TOTAL                          |          2

📈 LINES OF CODE BY AGENT

--- Aggregate ---
Agent                     |    Commits |  Additions |  Deletions
------------------------------------------------------------
IBM Bob, OpenCode + Qwen3.6-35b |          2 |         90 |          1
------------------------------------------------------------
TOTAL                     |          2 |         90 |          1

📈 LINES OF CODE BY USAGE TYPE

--- Aggregate ---
Usage Type           |    Commits |  Additions |  Deletions
-------------------------------------------------------
none                 |          0 |          0 |          0
draft                |          1 |         56 |          1
full                 |          1 |         34 |          0
-------------------------------------------------------
TOTAL                |          2 |         90 |          1

@gabe-l-hart
gabe-l-hart requested review from a team and CISC as code owners June 19, 2026 19:46
@github-actions github-actions Bot added examples python python script changes labels Jun 19, 2026
@gabe-l-hart
gabe-l-hart requested a review from ngxson June 19, 2026 19:48
Comment thread gguf-py/gguf/constants.py Outdated
Branch: GraniteSpeechPlus
AI-usage: full (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
Branch: GraniteSpeechPlus
AI-usage: draft (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
Comment thread tools/mtmd/models/granite-speech.cpp Outdated
Comment thread tools/mtmd/clip-model.h Outdated
@gabe-l-hart

Copy link
Copy Markdown
Collaborator Author

@ngxson more cleanup changes coming, will ping when review ready

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
@gabe-l-hart

Copy link
Copy Markdown
Collaborator Author

@ngxson Ok, ready for review now. Sorry to miss some of those naming mismatches.

Comment thread conversion/granite.py Outdated
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
@ngxson
ngxson merged commit a3900a6 into ggml-org:master Jun 23, 2026
26 of 27 checks passed
@gabe-l-hart
gabe-l-hart deleted the GraniteSpeechPlus branch June 23, 2026 12:25
Geminihaha pushed a commit to Geminihaha/llama.cpp that referenced this pull request Jun 25, 2026
* feat: Add conversion support for Granite Speech Plus

Branch: GraniteSpeechPlus
AI-usage: full (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* feat: Extend granite_speech to support plus multi-layer concatenation

Branch: GraniteSpeechPlus
AI-usage: draft (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(conversion): Fix plural naming for feature_layers for audio

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(mtmd): Align feature_layer usage and naming everywhere

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* style: Use fstring for log

Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>

---------

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
adrianhoehne pushed a commit to adrianhoehne/llama.cpp that referenced this pull request Jul 5, 2026
* feat: Add conversion support for Granite Speech Plus

Branch: GraniteSpeechPlus
AI-usage: full (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* feat: Extend granite_speech to support plus multi-layer concatenation

Branch: GraniteSpeechPlus
AI-usage: draft (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(conversion): Fix plural naming for feature_layers for audio

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(mtmd): Align feature_layer usage and naming everywhere

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* style: Use fstring for log

Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>

---------

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
platima added a commit to platima/llama.cpp-SpacemiT-K3 that referenced this pull request Jul 17, 2026
Port upstream a3900a6 (ggml-org#24818 "model: Granite Speech Plus") to the
SpacemiT K3 fork. The Plus variant concatenates selected encoder feature
layers with the final encoder output to form the QFormer projector input,
so proj_input_dim = n_embd * (feature_layers + 1). For the 4.1-2b-plus
mmproj feature_layers=[3] -> proj_input_dim=2048, matching the projector
cross_attn k/v weights [2048,1024].

Adapted to the fork's diverged mtmd lineage:
- clip-model.h: add ordered feature_layers vector + is_feature_layer()
  helper alongside the existing set-based vision_feature_layer (left
  untouched so vision/llava/granite4 paths are unaffected).
- clip-impl.h: add prefixed KEY_FEATURE_LAYERS "clip.%s.feature_layer".
- clip.cpp: parse clip.audio.feature_layer in the granite-speech loader
  (optional; empty for non-plus models) + log it.
- granite-speech.cpp: capture layer-0/intermediate/final outputs and
  ggml_concat them; reshape enc_windows with proj_input_dim. Non-plus
  path (empty feature_layers) is unchanged (proj_input_dim == n_embd).
- gguf-py + conversion/granite.py: constant, writer helper, and
  GraniteSpeechPlusMmprojModel for producing Plus GGUFs.

No K3 backend fix needed: the concat tensors are graph-internal, not
disk-loaded weights, so they never hit the IME2 q8_0 retype heuristic.

Verified on K3 (bf16 main + f16 mmproj, -t 8, -fa 1): audio run exits 0
and transcribes tools/mtmd/test-2.mp3 coherently ("The New York Times
from July 21st 1969 ... Men walk on moon ...").

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
platima added a commit to platima/llama.cpp-SpacemiT-K3 that referenced this pull request Jul 17, 2026
Stamp bump to 29 + docs for two multimodal models now validated on K3
hardware (both were code-present via the 0.1.6 mtmd lineage; patch 29
makes them run):

- DeepSeek-OCR-2 (deepseekocr2 vision): fixed a clip-warmup crash — our
  vision-retype was q8_0'ing the learned resample_query data tensor into
  the IME2 buffer, leaving its CPY with no backend (fixed in 458e725).
- Granite Speech 4.1 2B Plus (granite_speech audio): ported the ggml-org#24818
  multi-layer feature concat (fixed in 7016aee).

MODELS.md Matrix A/B updated; TODO.md patch-29 entry records both fixes
plus the deferred EAGLE3 port and the off-mainline Hy3/Cohere2 arches.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* feat: Add conversion support for Granite Speech Plus

Branch: GraniteSpeechPlus
AI-usage: full (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* feat: Extend granite_speech to support plus multi-layer concatenation

Branch: GraniteSpeechPlus
AI-usage: draft (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(conversion): Fix plural naming for feature_layers for audio

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(mtmd): Align feature_layer usage and naming everywhere

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* style: Use fstring for log

Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>

---------

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* feat: Add conversion support for Granite Speech Plus

Branch: GraniteSpeechPlus
AI-usage: full (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* feat: Extend granite_speech to support plus multi-layer concatenation

Branch: GraniteSpeechPlus
AI-usage: draft (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(conversion): Fix plural naming for feature_layers for audio

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(mtmd): Align feature_layer usage and naming everywhere

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* style: Use fstring for log

Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>

---------

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* feat: Add conversion support for Granite Speech Plus

Branch: GraniteSpeechPlus
AI-usage: full (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* feat: Extend granite_speech to support plus multi-layer concatenation

Branch: GraniteSpeechPlus
AI-usage: draft (Bob, OpenCode + Qwen3.6-35b)
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(conversion): Fix plural naming for feature_layers for audio

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* fix(mtmd): Align feature_layer usage and naming everywhere

Branch: GraniteSpeechPlus
AI-usage: none
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* style: Use fstring for log

Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>

---------

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

examples python python script changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants