Repository navigation
feat(gemma4): tunable per-request vision image-token budget - #2
Merged
Merged
Conversation
Gemma 4 vision (gemma4v projector) defaults to a max of 280 image tokens in llama.cpp (set_limit_image_tokens(40, 280)). Pass --image-max-tokens 1120 for the gemma4 architecture so high-resolution images can use more visual tokens for detail. The min is left at the model default. Rename qwenVLServerArgs -> visionServerArgs since it now tunes image token budgets for more than one architecture. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pair the raised max ceiling with an explicit --image-min-tokens 40 for the gemma4 architecture. 40 is llama.cpp's documented gemma4v floor (set_limit_image_tokens(40, 280)); passing it explicitly makes the min a visible, tunable knob alongside the max without changing default behavior. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Expose the Gemma 4 vision image-token budget as api.Options fields (image_min_tokens / image_max_tokens on Runner) so it can be tuned per request (e.g. from Open WebUI), defaulting to 40 / 1120 via DefaultOptions. visionServerArgs now reads them, clamping min <= max since llama-server refuses to load when image_max_pixels < image_min_pixels. Because they are Runner options, the scheduler's existing reload-on-change (reflect.DeepEqual over Runner) reloads the runner when the budget changes, so no scheduler change is required. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Expose the Gemma 4 vision image-token budget as per-request options —
image_min_tokens/image_max_tokensonapi.Options(in theRunnergroup) — defaulting to 40 / 1120 viaDefaultOptions. This lets it be tuned per request (e.g. from Open WebUI's params) instead of being fixed.visionServerArgsreads the values and passes them to llama-server as--image-min-tokens/--image-max-tokens, which override llama.cpp'sgemma4vdefaults (set_limit_image_tokens(40, 280)) viacustom_image_min/max_tokens(clip-model.h). Min is clamped to max because llama-server refuses to load whenimage_max_pixels < image_min_pixels.Restores the Open WebUI workflow of the old (removed)
model.gofeature, but Go-only on currentmain— noclip.cpppatch.Semantics (important)
These are
Runner(load-time) options, so they behave exactly likenum_ctx: set them per request, and if the value differs from the loaded runner, the scheduler reloads the runner with the new flags. There is no per-image budget in the mtmd/completioncontract, so a true no-reload per-request budget would require patching llama.cpp — deliberately out of scope.Blast radius
api/types.go— 2 additiveRunnerfields +DefaultOptionsdefaults (parsed by the existing genericreflect.Intpath).llm/llama_server.go—visionServerArgsreads opts + clamps; 1 call site.server/sched.go— unchanged; reload-on-change is handled by the existingreflect.DeepEqualoverRunner.Empirical validation
Built against
b9888and tested withgemma4:e2b-it-q4_K_M(image tokens =prompt_eval_count− text baseline). Per-request options honored, with reload-on-change confirmed in the log (sched.go:263 msg=reloading runner; distinct launches40/1120,40/560,300/1120):image_max_tokens: 560image_min_tokens: 300Generation is not regressed
Real descriptions at the raised budget (~1091 tokens, 4× the 280 reference) are coherent and accurate, stopping naturally (
done=stop) — no garbage/crash (the position-embedding-OOB risk did not materialize). Text-only unaffected.Testing
Commits
--image-min-tokens.api.Optionsknob (this makes 1–2 the default case).🤖 Generated with Claude Code