Repository navigation
Gemma 4 visual token budgets (image_min_tokens / image_max_tokens) — on last Go-runner base - #1
glennneuber wants to merge 10 commits into
Conversation
Add SnapGemma4VisualTokens and NormalizeGemma4ImageBudgets for ladder
{70,140,280,560,1120} with defaults 70/560, tie-break to lower rung, and
table-driven tests.
Co-authored-by: Cursor <cursoragent@cursor.com>
JSON keys match Gemma visual token budget semantics (not bitmap pixel caps). Co-authored-by: Cursor <cursoragent@cursor.com>
Optional interface for models that encode vision with per-call min/max token budgets. Co-authored-by: Cursor <cursoragent@cursor.com>
Implement EncodeMultimodalWithBudgets, ProcessImageWithBudgets, smartResize min+max, default 70/560 via gemma4vision, and OLLAMA_DEBUG-friendly logs. Co-authored-by: Cursor <cursoragent@cursor.com>
Normalize from req.Options, pass through NewSequence into inputs, use MultimodalBudgetEncoder when present, and worst-case graph encode with max ladder budget. Co-authored-by: Cursor <cursoragent@cursor.com>
Extend needsReload for non-MLX alongside Runner DeepEqual. Add sched test for ImageMaxTokens-only delta. Co-authored-by: Cursor <cursoragent@cursor.com>
Vision is not implemented on MLX; log at Debug for visibility. Co-authored-by: Cursor <cursoragent@cursor.com>
Preserves the Cursor plan in-repo for reviewers and future maintenance. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Records why the feature cannot be rebased onto current main (upstream removed the Go inference runner in ollama#16031 and the Go model path in ollama#17007), the base-selection analysis, the choice of f63eea3 as the last fully-wired base, the verification results, and what a forward-port to the llama-server/mtmd architecture would require. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Forward-port trackingGitHub Issues are disabled on this repo, so tracking the mtmd forward-port here on the PR. ContextThe Gemma 4 visual token budget feature (
On current ProblemThe feature worked by having Go resize the image to hit a token budget before running vision inference itself. That Go-side hook no longer exists, and for Gemma specifically:
So this is a re-implementation against First step (decides the whole approach)
Candidate approaches (pick after the first step)
Also carry over from #1
References
|
|
Closing as superseded by #2. This PR rebased the original Go-runner-based Gemma 4 visual-token-budget feature onto That premise turned out to be wrong. Verifying against the actually-pinned llama.cpp ( So the real fix for current ➡️ #2 — tune the Gemma 4 image-token budget (min 40 / max 1120) via The design notes on this branch (and the tracking comment above) that describe Gemma vision as "fixed / needs a C++ port" were based on a stale materialized Leaving the branch in place for historical reference; no merge intended. |
Summary
Adds Gemma 4 visual token budget options
image_min_tokens/image_max_tokensonapi.Options, snapping to the ladder{70,140,280,560,1120}(defaults 70 / 560 when unset), with per-completion wiring throughollamarunnerand a non-MLX scheduler reload when those options change.Full feature design: docs/design/gemma4-vision-token-budgets.md.
Base branch — please read
This PR targets
base/gemma4-last-go-runner(= upstreamf63eea3d, 2026-05-24), notmain, on purpose:runner/ollamarunner) and the Go model path (model/model.go,model/models/gemma4/).runner/ollamarunnerin runner: Remove CGO engines, use llama-server exclusively for GGML models ollama/ollama#16031 (9db4bdba) andmodel.go+model/models/gemma4/in llama: clean up dead code from llama-server work ollama/ollama#17007 (7b22ac96) — so the branch cannot be rebased onto currentmain; its core commits patch files that no longer exist.f63eea3dis the last upstream commit where the full Go stack is present and wired, so the feature actually runs. Targeting it keeps this PR to a clean, reviewable feature diff.Rationale and forward-port analysis: docs/design/gemma4-vision-token-budgets-upstream-rebase.md.
Verification
f63eea3dwith zero conflicts;git range-diffconfirms every patch is byte-identical to the original branch.OllamaEngineRequired("gemma4") == true):api.Options→ server →ollamarunner→EncodeMultimodalWithBudgets→ProcessImageWithBudgets; scheduler reloads on budget change.go build ./...succeeds,go vetclean, feature unit tests pass, 30 scheduler tests pass (go test ./server/ -run 'Sched|Image|Token|Reload').Known risk (pre-existing, not introduced here)
Raising the max budget to 560/1120 (vs the reference max ~280) exercises the vision position-embedding table (
model/models/gemma4/model_vision.go:323), which is indexed without clamping. A large / extreme-aspect image at a high budget could index out of bounds. Recommend a smoke test with a real Gemma 4 vision model atimage_max_tokens: 1120on a wide image before relying on high budgets.Testing
🤖 Generated with Claude Code