Skip to content

lfm2 : make LFM2.5-VL-3B work with a DSpark draft model - #1

Open
tugot17 wants to merge 1 commit into
masterfrom
piotr/lfm2-vl-dspark
Open

tugot17 wants to merge 1 commit into
masterfrom
piotr/lfm2-vl-dspark

Conversation

@tugot17

@tugot17 tugot17 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

This makes LFM2.5-VL-3B work with a DSpark draft model.

The lfm2 graph lost its t_layer_inp capture in the models/ split (still there at 38a5b42, and qwen3.cpp kept its copy). Without it the server asserts on the first request whenever a drafter is attached:

llama-graph.cpp:1352: GGML_ASSERT(t_layer_inp[il] != nullptr && "layer input tensor is null")

One line to put it back.

Testing:

hf download LiquidAI/LFM2.5-VL-3B-GGUF LFM2.5-VL-3B-F16.gguf mmproj-LFM2.5-VL-3B-F16.gguf --local-dir models
hf download tugot17/mango-gguf mango-draft-F16.gguf --local-dir models

./build/bin/llama-server -m models/LFM2.5-VL-3B-F16.gguf \
  --mmproj models/mmproj-LFM2.5-VL-3B-F16.gguf \
  -md models/mango-draft-F16.gguf \
  --spec-type draft-dspark --spec-draft-n-max 8 --spec-draft-n-min 0 \
  -fa on -ngl 99 -c 8192

Target and mmproj are public, the draft sidecar is gated on my HF, ping me for access.

Accept length on a 20-sample MMSpec probe: 3.86 (at 38a5b42), our internal fork gets ~4.1 on the same samples. Greedy output is identical with and without the drafter.

Things I hit on the way that this PR does not fix:

  • accept length dropped ~30% for every DSpark draft between 38a5b42 and current master (3.86 -> 2.85, same probe). needs a bisect, my guess is spec: support speculators-format checkpoints for DSpark ggml-org/llama.cpp#26275
  • the Lfm2DSparkDraftModel converter path is gone from master, I converted the sidecar with an older tree. btw the old converter applied the q/k rope permute twice for rope_is_neox_style=false checkpoints, which is why all our earlier VL sidecars measured ~1.7. it has to be applied once
  • convert_hf_to_gguf --target-model-dir doesn't accept a VL target dir, Lfm2VlForConditionalGeneration is only in the mmproj map

The models/ split dropped the loop-head capture (present at 38a5b42,
src/models/lfm2.cpp:252), so any decode with embeddings_layer_inp set -
i.e. serving an LFM2 target with a DFlash/DSpark drafter - aborts in
llm_graph_result::set_outputs (llama-graph.cpp:1352 'layer input tensor
is null') as soon as graph_reserve builds the embd-batch graph variant,
e.g. on the first mtmd image chunk of a vision request.

Same unconditional capture style as qwen3.cpp. Verified: LFM2.5-VL-3B +
DSpark sidecar no longer crashes on image requests; text-only LFM2.5-2.6B
+ its released sidecar unaffected (tau ~3.8 on a math probe).
@tugot17 tugot17 changed the title staging: LFM2.5-VL DSpark upstreaming (crash fix + parity investigation) add DSpark draft-model support for LFM2.5-VL-3B Sep 23, 2026
@tugot17 tugot17 changed the title add DSpark draft-model support for LFM2.5-VL-3B lfm2 : make LFM2.5-VL-3B work with a DSpark draft model Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant