Skip to content

mtmd : read LFM2 tiling params from GGUF metadata - #25524

Open
michael-dm wants to merge 1 commit into
ggml-org:masterfrom
michael-dm:lfm2-configurable-tiling
Open

michael-dm wants to merge 1 commit into
ggml-org:masterfrom
michael-dm:lfm2-configurable-tiling

Conversation

@michael-dm

@michael-dm michael-dm commented Jul 10, 2026 •

Copy link
Copy Markdown

Overview

Read LFM2/LFM2.5-VL image tiling parameters from GGUF metadata instead of hardcoded constants. The converter now writes min_tiles, max_tiles, tile_size and max_pixels_tolerance from processor_config.json, following the existing InternVL preproc_min_tiles / preproc_max_tiles pattern. do_image_splitting=false maps to min_tiles=max_tiles=1, which skips tiling.

Motivation: fine-tunes trained with do_image_splitting=false currently get forced tiling (19 chunks instead of 1 for a 1376x768 input), which breaks parity with the HF pipeline. GGUFs converted before this change keep the exact current behavior via defaults (2/10/512/2.0).

Additional information

Related: #21658 took a runtime CLI-flag approach to the same problem and was closed as too use-case specific. This does it at conversion time via metadata, consistent with how other projectors handle preprocessing params.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - code written with AI assistance (Claude), reviewed and submitted by me.

Write min_tiles, max_tiles, tile_size and max_pixels_tolerance from
processor_config.json when converting LFM2-VL mmproj, read them in clip.cpp
with the previous constants as defaults. do_image_splitting=false maps to
min_tiles=max_tiles=1 (no tiling).

Assisted-by: Claude Fable 5
Claude-Session: https://claude.ai/code/session_01Xyxm13w4R2MjaLzwpiHDbi
@michael-dm
michael-dm requested review from a team and CISC as code owners July 10, 2026 12:20
@github-actions github-actions Bot added mtmd Related to multimodal functionality (video/image/audio) conversion labels Jul 10, 2026
@CISC

CISC commented Aug 2, 2026

Copy link
Copy Markdown
Member

This should probably be done for DeepseekOCR2 too, otherwise LGTM, up to @ngxson though, needs rebase.

@ngxson

ngxson commented Aug 2, 2026

Copy link
Copy Markdown
Collaborator

@tdakhran could you have a look on this?

@tdakhran

tdakhran commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

@ngxson looks good to me, it keeps the default behavior unchanged

@tdakhran

Copy link
Copy Markdown
Contributor

this is very timely, https://huggingface.co/LiquidAI/LFM2.5-VL-3B has min_tiles=1 while https://huggingface.co/LiquidAI/LFM2.5-VL-1.6B has min_tiles=2.

@michael-dm , could you please rebase to resolve conflicts?
@ngxson , is anything else required to merge?

Comment thread tools/mtmd/mtmd-image.cpp
Comment on lines +948 to +949
const bool needs_tiling = hparams.preproc_max_tiles > 1
&& (original_size.width > tile_size * max_pixels_tolerance || original_size.height > tile_size * max_pixels_tolerance);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

note this will be modified in #27057

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion mtmd Related to multimodal functionality (video/image/audio)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants