Skip to content

Gemma 4 visual token budgets (image_min_tokens / image_max_tokens) — on last Go-runner base - #1

Closed
glennneuber wants to merge 10 commits into
base/gemma4-last-go-runnerfrom
feat/gemma4-visual-token-budgets-last-go-runner
Closed

glennneuber wants to merge 10 commits into
base/gemma4-last-go-runnerfrom
feat/gemma4-visual-token-budgets-last-go-runner

Conversation

@glennneuber

Copy link
Copy Markdown

Summary

Adds Gemma 4 visual token budget options image_min_tokens / image_max_tokens on api.Options, snapping to the ladder {70,140,280,560,1120} (defaults 70 / 560 when unset), with per-completion wiring through ollamarunner and a non-MLX scheduler reload when those options change.

Full feature design: docs/design/gemma4-vision-token-budgets.md.

Base branch — please read

This PR targets base/gemma4-last-go-runner (= upstream f63eea3d, 2026-05-24), not main, on purpose:

Rationale and forward-port analysis: docs/design/gemma4-vision-token-budgets-upstream-rebase.md.

Verification

  • 9 feature commits replayed onto f63eea3d with zero conflicts; git range-diff confirms every patch is byte-identical to the original branch.
  • Wired end-to-end and live for GGUF gemma4 (OllamaEngineRequired("gemma4") == true): api.Options → server → ollamarunner → EncodeMultimodalWithBudgets → ProcessImageWithBudgets; scheduler reloads on budget change.
  • go build ./... succeeds, go vet clean, feature unit tests pass, 30 scheduler tests pass (go test ./server/ -run 'Sched|Image|Token|Reload').

Known risk (pre-existing, not introduced here)

Raising the max budget to 560/1120 (vs the reference max ~280) exercises the vision position-embedding table (model/models/gemma4/model_vision.go:323), which is indexed without clamping. A large / extreme-aspect image at a high budget could index out of bounds. Recommend a smoke test with a real Gemma 4 vision model at image_max_tokens: 1120 on a wide image before relying on high budgets.

Testing

go test ./internal/gemma4vision/... ./model/models/gemma4/...
go test ./runner/ollamarunner/... ./server/...

🤖 Generated with Claude Code

Local Dev and others added 10 commits July 15, 2026 00:02
Add SnapGemma4VisualTokens and NormalizeGemma4ImageBudgets for ladder
{70,140,280,560,1120} with defaults 70/560, tie-break to lower rung, and
table-driven tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
JSON keys match Gemma visual token budget semantics (not bitmap pixel caps).

Co-authored-by: Cursor <cursoragent@cursor.com>
Optional interface for models that encode vision with per-call min/max
token budgets.

Co-authored-by: Cursor <cursoragent@cursor.com>
Implement EncodeMultimodalWithBudgets, ProcessImageWithBudgets, smartResize
min+max, default 70/560 via gemma4vision, and OLLAMA_DEBUG-friendly logs.

Co-authored-by: Cursor <cursoragent@cursor.com>
Normalize from req.Options, pass through NewSequence into inputs, use
MultimodalBudgetEncoder when present, and worst-case graph encode with
max ladder budget.

Co-authored-by: Cursor <cursoragent@cursor.com>
Extend needsReload for non-MLX alongside Runner DeepEqual. Add sched test
for ImageMaxTokens-only delta.

Co-authored-by: Cursor <cursoragent@cursor.com>
Vision is not implemented on MLX; log at Debug for visibility.

Co-authored-by: Cursor <cursoragent@cursor.com>
Preserves the Cursor plan in-repo for reviewers and future maintenance.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Records why the feature cannot be rebased onto current main (upstream
removed the Go inference runner in ollama#16031 and the Go model path in
ollama#17007), the base-selection analysis, the choice of f63eea3 as the last
fully-wired base, the verification results, and what a forward-port to
the llama-server/mtmd architecture would require.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

Forward-port tracking

GitHub Issues are disabled on this repo, so tracking the mtmd forward-port here on the PR.


Context

The Gemma 4 visual token budget feature (image_min_tokens / image_max_tokens) currently lives on feat/gemma4-visual-token-budgets-last-go-runner (#1), based on upstream f63eea3d (2026-05-24). It cannot be rebased onto current main because upstream removed the architecture it was built on:

On current main, Ollama ships raw image bytes to llama-server; all resize + tokenization happens inside llama.cpp mtmd/clip (C++), which is fetched at build (LLAMA_CPP_VERSION), not vendored.

Problem

The feature worked by having Go resize the image to hit a token budget before running vision inference itself. That Go-side hook no longer exists, and for Gemma specifically:

  • mtmd currently produces a fixed token count per image and ignores the min/max-token levers.
  • Those levers (--image-min/max-tokens) are per-process and only honored by dynamic-resolution projectors (Qwen-VL etc.), not Gemma.
  • Upstream's vision.min_image_tokens / vision.max_image_tokens GGUF KV is write-only / inert in Ollama today (only convert_lfm2_vl.go sets it, nothing reads it), and gemma4 doesn't emit it.

So this is a re-implementation against mtmd, not a rebase.

First step (decides the whole approach)

  • Confirm what the pinned llama.cpp mtmd actually does for Gemma today — is the token count truly fixed (~256), or does it already vary by resolution / pan-and-scan? Read the clip.cpp Gemma path at the pinned LLAMA_CPP_VERSION. If tokens already vary with resolution, the feature may be largely obsolete and this issue can close.

Candidate approaches (pick after the first step)

  • A — patch mtmd/clip (via llama/compat/) to give Gemma a dynamic-resolution path honoring a token budget, plumbed per-request through the /completion contract. Faithful, but high-maintenance (re-patch on every LLAMA_CPP_VERSION bump).
  • B — per-process flag — enable --image-min/max-tokens for Gemma at launch, distinct runners per budget (reuse the old sched.go reload idea). Still needs A's C++ enablement; coarse granularity.
  • C — upstream it — contribute Gemma dynamic-resolution + per-request budget to llama.cpp/mtmd and Ollama. Cleanest long-term; stops the divergence.
  • D — model-level default — bake budgets at convert time via vision.*_image_tokens KV (align with lfm2), if/when mtmd honors them for Gemma. Loses per-request control.

Also carry over from #1

  • Position-embedding table (model/models/gemma4/model_vision.go:323) is indexed without clamping; high budgets (560/1120 vs reference ~280) risk out-of-bounds on large / extreme-aspect images. Clamp and/or smoke-test with a real model at image_max_tokens: 1120.

References

@glennneuber

Copy link
Copy Markdown
Author

Closing as superseded by #2.

This PR rebased the original Go-runner-based Gemma 4 visual-token-budget feature onto f63eea3d (the last upstream commit before the Go inference runner was removed), on the premise that the feature was "stranded" and would need a C++/mtmd port to work on current main.

That premise turned out to be wrong. Verifying against the actually-pinned llama.cpp (b9888) showed upstream now supports Gemma 4 vision natively: the gemma4v projector uses the dynamic-resolution preprocessor with an image-token budget (set_limit_image_tokens(40, 280)), overridable via --image-min-tokens / --image-max-tokens. Ollama's single-file gemma4 GGUF is routed to that path via a compat shim (handle_gemma4_clip: … translating).

So the real fix for current main is a small Go-only change — no clip.cpp patch, no rebase onto old upstream:

➡️ #2 — tune the Gemma 4 image-token budget (min 40 / max 1120) via visionServerArgs. Empirically verified: a 2048² image goes from a fixed ~258 to ~1091 image tokens, generation not regressed.

The design notes on this branch (and the tracking comment above) that describe Gemma vision as "fixed / needs a C++ port" were based on a stale materialized llama/llama.cpp copy, not the pinned b9888, and are superseded by the verification in #2.

Leaving the branch in place for historical reference; no merge intended.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant