Skip to content

model : add Ling 3.0 VL support - #29151

Merged
CISC merged 2 commits into
ggml-org:masterfrom
aetherbird:ling3-vl-only
Sep 24, 2026
Merged

CISC merged 2 commits into
ggml-org:masterfrom
aetherbird:ling3-vl-only

Conversation

@aetherbird

@aetherbird aetherbird commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Adds support for Ling-3.0-flash-VL, the vision-language variant of Ling 3.0 Flash (124B total / 5.1B active, hybrid KDA + gated MLA, 512-expert MoE). The text backbone is identical to bailingmoe3; the release adds a 27-block vision tower (qwen3_moe_vit family), a norm-only 2x2 merger, a two-layer projector, and uses M-RoPE with sections [8, 12, 12] shared between text and vision positions.

Additional information

  • test-llama-archs --arch bailingmoe3 passes (MoE fixture round-trip).
  • Converted the full model to BF16 GGUF (917 tensors, 248.9 GB) and mmproj (334 tensors) from the released safetensors.
  • End-to-end image + video inference on llama-server successful.

Requirements

  • I have read and agree with the contributing guidelines.
  • AI usage disclosure: Yes, I used AI to help write and check the code. I personally, thoroughly reviewed the CONTRIBUTING.md file.

@github-actions github-actions Bot added model Model specific testing Everything test related mtmd Related to multimodal functionality (video/image/audio) conversion labels Sep 19, 2026
@aetherbird
aetherbird force-pushed the ling3-vl-only branch 3 times, most recently from 34677b8 to 91ae657 Compare September 19, 2026 23:06
@aetherbird
aetherbird marked this pull request as ready for review September 19, 2026 23:19
@aetherbird
aetherbird requested review from a team, CISC and ggerganov as code owners September 19, 2026 23:19
Comment thread tools/mtmd/models/ling3vl.cpp Outdated
Comment thread gguf-py/gguf/constants.py Outdated
Comment thread conversion/bailingmoe3vl.py Outdated
Comment thread conversion/bailingmoe3vl.py Outdated
Comment thread conversion/bailingmoe3vl.py Outdated
Comment thread conversion/bailingmoe3vl.py Outdated
@aetherbird
aetherbird force-pushed the ling3-vl-only branch 2 times, most recently from 9b7299e to 7fcea1b Compare September 22, 2026 01:37
@aetherbird aetherbird changed the title model : add Ling 3.0 VL (BailingMoeV3VL) support model : fold Ling 3.0 VL into the BailingMoeV3 architecture Sep 22, 2026
@aetherbird

aetherbird commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

Reworked per review in 7fcea1b. "bailingmoe3vl" has been removed, and VL now converts under bailingmoe3 in conversion/bailingmoe3.py (text model subclasses BailingMoeV3Model, vision tower subclasses Qwen3VLVisionModel).

  • mrope_section is written as rope.dimension_sections, M-RoPE is selected via use_mrope() when sections are present, and text-only files will keep NORM rope.
  • Projector type renamed to ling3vl.
  • The mtmd graph uses build_vit with a rope callback instead of a hand-written ViT loop.
  • num_shared_experts fallback has been moved into BailingMoeV3Model and the num_nextn_predict_layers default has been removed.

Validation: test-llama-archs --arch bailingmoe3 passes on the M-RoPE fixture. End-to-end: a converted GGUF produces byte-identical greedy output (text and image prompts) before and after the fold.

@aetherbird
aetherbird requested a review from ngxson September 22, 2026 01:59
Comment thread conversion/bailingmoe3.py Outdated

@CISC CISC left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fix mrope CI failure. :)

Comment thread conversion/bailingmoe3.py Outdated
@ngxson

ngxson commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

hmm, did your agent accidentally changed the title of this PR?

@aetherbird

aetherbird commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

hmm, did your agent accidentally changed the title of this PR?

No. I re-titled this PR to match the intention of the rework that you and CISC requested.

@aetherbird

Copy link
Copy Markdown
Contributor Author

BailingMoeV3VLModel was missing the model_arch declaration that TextModel.__init_subclass__ requires. This has been fixed in 0d9cb8e.

The converter has now been validated end-to-end on a synthetic VL checkpoint.

@ngxson

ngxson commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

"fold" sounds weird. original title is already clear enough

many PRs in the past already doing exactly what you are doing here. you should have a look at them first

@aetherbird aetherbird changed the title model : fold Ling 3.0 VL into the BailingMoeV3 architecture model : add Ling 3.0 VL support Sep 22, 2026
@aetherbird

Copy link
Copy Markdown
Contributor Author

"fold" sounds weird. original title is already clear enough

many PRs in the past already doing exactly what you are doing here. you should have a look at them first

Changed.

@aetherbird

Copy link
Copy Markdown
Contributor Author
  • Re-ran the converter against the released inclusionAI/Ling-3.0-flash-VL checkpoint.
  • Re-validated video input decode (the model described video input correctly).
  • Re-validated that a 76k token prompt processes and generates cleanly.

@CISC

CISC commented Sep 23, 2026

Copy link
Copy Markdown
Member

@aetherbird

Copy link
Copy Markdown
Contributor Author

@aetherbird re #29151 (review), see f.ex. https://github.com/ggml-org/llama.cpp/actions/runs/35756040134/job/106940559614?pr=29151#step:3:1262

Fixed in 8ef43a4 by giving bailingmoe3 its own case after the list's return LLAMA_ROPE_TYPE_NORM.

The bailingmoe3 mrope gate was appended to the shared NORM fall-through list in llama_model_rope_type(), so every arch in that list inherited the use_mrope() check. deepseek32 got M-RoPE positions and its ggml_rope_ext call hit the assert.

Verified on a build of the PR merged with current master: test-generate-models generates all fixtures with no asserts and test-llama-archs --arch bailingmoe3 still passes.

@CISC
CISC merged commit f830688 into ggml-org:master Sep 24, 2026
27 of 28 checks passed
@aetherbird
aetherbird deleted the ling3-vl-only branch September 26, 2026 21:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion model Model specific mtmd Related to multimodal functionality (video/image/audio) testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants