Repository navigation
model: Granite Speech Plus - #24818
Merged
Merged
Conversation
ngxson
reviewed
Jun 20, 2026
Branch: GraniteSpeechPlus AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
Branch: GraniteSpeechPlus AI-usage: draft (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
gabe-l-hart
force-pushed
the
GraniteSpeechPlus
branch
from
June 22, 2026 16:21
50a246f to
136dca2
Compare
gabe-l-hart
commented
Jun 22, 2026
gabe-l-hart
commented
Jun 22, 2026
Collaborator
Author
|
@ngxson more cleanup changes coming, will ping when review ready |
Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>
gabe-l-hart
force-pushed
the
GraniteSpeechPlus
branch
from
June 22, 2026 16:47
136dca2 to
39c757a
Compare
Collaborator
Author
|
@ngxson Ok, ready for review now. Sorry to miss some of those naming mismatches. |
ngxson
approved these changes
Jun 22, 2026
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
CISC
approved these changes
Jun 23, 2026
Geminihaha
pushed a commit
to Geminihaha/llama.cpp
that referenced
this pull request
Jun 25, 2026
* feat: Add conversion support for Granite Speech Plus Branch: GraniteSpeechPlus AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Extend granite_speech to support plus multi-layer concatenation Branch: GraniteSpeechPlus AI-usage: draft (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(conversion): Fix plural naming for feature_layers for audio Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(mtmd): Align feature_layer usage and naming everywhere Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * style: Use fstring for log Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> --------- Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
adrianhoehne
pushed a commit
to adrianhoehne/llama.cpp
that referenced
this pull request
Jul 5, 2026
* feat: Add conversion support for Granite Speech Plus Branch: GraniteSpeechPlus AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Extend granite_speech to support plus multi-layer concatenation Branch: GraniteSpeechPlus AI-usage: draft (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(conversion): Fix plural naming for feature_layers for audio Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(mtmd): Align feature_layer usage and naming everywhere Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * style: Use fstring for log Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> --------- Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
platima
added a commit
to platima/llama.cpp-SpacemiT-K3
that referenced
this pull request
Jul 17, 2026
Port upstream a3900a6 (ggml-org#24818 "model: Granite Speech Plus") to the SpacemiT K3 fork. The Plus variant concatenates selected encoder feature layers with the final encoder output to form the QFormer projector input, so proj_input_dim = n_embd * (feature_layers + 1). For the 4.1-2b-plus mmproj feature_layers=[3] -> proj_input_dim=2048, matching the projector cross_attn k/v weights [2048,1024]. Adapted to the fork's diverged mtmd lineage: - clip-model.h: add ordered feature_layers vector + is_feature_layer() helper alongside the existing set-based vision_feature_layer (left untouched so vision/llava/granite4 paths are unaffected). - clip-impl.h: add prefixed KEY_FEATURE_LAYERS "clip.%s.feature_layer". - clip.cpp: parse clip.audio.feature_layer in the granite-speech loader (optional; empty for non-plus models) + log it. - granite-speech.cpp: capture layer-0/intermediate/final outputs and ggml_concat them; reshape enc_windows with proj_input_dim. Non-plus path (empty feature_layers) is unchanged (proj_input_dim == n_embd). - gguf-py + conversion/granite.py: constant, writer helper, and GraniteSpeechPlusMmprojModel for producing Plus GGUFs. No K3 backend fix needed: the concat tensors are graph-internal, not disk-loaded weights, so they never hit the IME2 q8_0 retype heuristic. Verified on K3 (bf16 main + f16 mmproj, -t 8, -fa 1): audio run exits 0 and transcribes tools/mtmd/test-2.mp3 coherently ("The New York Times from July 21st 1969 ... Men walk on moon ..."). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
platima
added a commit
to platima/llama.cpp-SpacemiT-K3
that referenced
this pull request
Jul 17, 2026
Stamp bump to 29 + docs for two multimodal models now validated on K3 hardware (both were code-present via the 0.1.6 mtmd lineage; patch 29 makes them run): - DeepSeek-OCR-2 (deepseekocr2 vision): fixed a clip-warmup crash — our vision-retype was q8_0'ing the learned resample_query data tensor into the IME2 buffer, leaving its CPY with no backend (fixed in 458e725). - Granite Speech 4.1 2B Plus (granite_speech audio): ported the ggml-org#24818 multi-layer feature concat (fixed in 7016aee). MODELS.md Matrix A/B updated; TODO.md patch-29 entry records both fixes plus the deferred EAGLE3 port and the off-mainline Hy3/Cohere2 arches. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
zommiommy
pushed a commit
to zommiommy/llama.cpp
that referenced
this pull request
Aug 18, 2026
* feat: Add conversion support for Granite Speech Plus Branch: GraniteSpeechPlus AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Extend granite_speech to support plus multi-layer concatenation Branch: GraniteSpeechPlus AI-usage: draft (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(conversion): Fix plural naming for feature_layers for audio Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(mtmd): Align feature_layer usage and naming everywhere Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * style: Use fstring for log Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> --------- Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
zbrad
pushed a commit
to zbrad/llama.cpp
that referenced
this pull request
Sep 10, 2026
* feat: Add conversion support for Granite Speech Plus Branch: GraniteSpeechPlus AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Extend granite_speech to support plus multi-layer concatenation Branch: GraniteSpeechPlus AI-usage: draft (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(conversion): Fix plural naming for feature_layers for audio Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(mtmd): Align feature_layer usage and naming everywhere Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * style: Use fstring for log Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> --------- Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
* feat: Add conversion support for Granite Speech Plus Branch: GraniteSpeechPlus AI-usage: full (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Extend granite_speech to support plus multi-layer concatenation Branch: GraniteSpeechPlus AI-usage: draft (Bob, OpenCode + Qwen3.6-35b) Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(conversion): Fix plural naming for feature_layers for audio Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix(mtmd): Align feature_layer usage and naming everywhere Branch: GraniteSpeechPlus AI-usage: none Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * style: Use fstring for log Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> --------- Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR adds support for ibm-granite/granite-speech-4.1-2b-plus. This model is a slight tweak on the
GraniteSpeechForConditionalGenerationarchitecture that extracts multiple encoder layers and concatenates them into the projected token stream.Additional information
Testing
Requirements
AI Disclosure
AI used in planning and building. All code reviewed and refined by hand.
git-ai-stats