Skip to content

feat(gemma4): tunable per-request vision image-token budget - #2

Merged
glennneuber merged 3 commits into
mainfrom
feat/gemma4-image-max-tokens
Jul 14, 2026
Merged

glennneuber merged 3 commits into
mainfrom
feat/gemma4-image-max-tokens

Conversation

@glennneuber

@glennneuber glennneuber commented Jul 14, 2026 •

Copy link
Copy Markdown

Summary

Expose the Gemma 4 vision image-token budget as per-request options — image_min_tokens / image_max_tokens on api.Options (in the Runner group) — defaulting to 40 / 1120 via DefaultOptions. This lets it be tuned per request (e.g. from Open WebUI's params) instead of being fixed.

visionServerArgs reads the values and passes them to llama-server as --image-min-tokens / --image-max-tokens, which override llama.cpp's gemma4v defaults (set_limit_image_tokens(40, 280)) via custom_image_min/max_tokens (clip-model.h). Min is clamped to max because llama-server refuses to load when image_max_pixels < image_min_pixels.

Restores the Open WebUI workflow of the old (removed) model.go feature, but Go-only on current main — no clip.cpp patch.

Semantics (important)

These are Runner (load-time) options, so they behave exactly like num_ctx: set them per request, and if the value differs from the loaded runner, the scheduler reloads the runner with the new flags. There is no per-image budget in the mtmd /completion contract, so a true no-reload per-request budget would require patching llama.cpp — deliberately out of scope.

Blast radius

  • api/types.go — 2 additive Runner fields + DefaultOptions defaults (parsed by the existing generic reflect.Int path).
  • llm/llama_server.go — visionServerArgs reads opts + clamps; 1 call site.
  • server/sched.go — unchanged; reload-on-change is handled by the existing reflect.DeepEqual over Runner.

Empirical validation

Built against b9888 and tested with gemma4:e2b-it-q4_K_M (image tokens = prompt_eval_count − text baseline). Per-request options honored, with reload-on-change confirmed in the log (sched.go:263 msg=reloading runner; distinct launches 40/1120, 40/560, 300/1120):

Request image tokens
2048² image, default ~1091
2048² image, image_max_tokens: 560 ~530 (max honored)
64² image, image_min_tokens: 300 ~333 (min floor honored, vs ~51 default)

Generation is not regressed

Real descriptions at the raised budget (~1091 tokens, 4× the 280 reference) are coherent and accurate, stopping naturally (done=stop) — no garbage/crash (the position-embedding-OOB risk did not materialize). Text-only unaffected.

Testing

go test ./api/ ./llm/   # incl. TestVisionServerArgs: defaults, custom budget, min>max clamp

Commits

  1. Raise the ceiling to 1120 (hardcoded launch flag).
  2. Add explicit --image-min-tokens.
  3. Make both a per-request api.Options knob (this makes 1–2 the default case).

🤖 Generated with Claude Code

glennneuber and others added 2 commits July 15, 2026 01:17
Gemma 4 vision (gemma4v projector) defaults to a max of 280 image tokens
in llama.cpp (set_limit_image_tokens(40, 280)). Pass --image-max-tokens
1120 for the gemma4 architecture so high-resolution images can use more
visual tokens for detail. The min is left at the model default.

Rename qwenVLServerArgs -> visionServerArgs since it now tunes image
token budgets for more than one architecture.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pair the raised max ceiling with an explicit --image-min-tokens 40 for the
gemma4 architecture. 40 is llama.cpp's documented gemma4v floor
(set_limit_image_tokens(40, 280)); passing it explicitly makes the min a
visible, tunable knob alongside the max without changing default behavior.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@glennneuber glennneuber changed the title feat(gemma4): raise vision image-token ceiling to 1120 feat(gemma4): tune vision image-token budget (min 40 / max 1120) Jul 14, 2026
Expose the Gemma 4 vision image-token budget as api.Options fields
(image_min_tokens / image_max_tokens on Runner) so it can be tuned per
request (e.g. from Open WebUI), defaulting to 40 / 1120 via DefaultOptions.
visionServerArgs now reads them, clamping min <= max since llama-server
refuses to load when image_max_pixels < image_min_pixels.

Because they are Runner options, the scheduler's existing reload-on-change
(reflect.DeepEqual over Runner) reloads the runner when the budget changes,
so no scheduler change is required.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@glennneuber glennneuber changed the title feat(gemma4): tune vision image-token budget (min 40 / max 1120) feat(gemma4): tunable per-request vision image-token budget Jul 14, 2026
@glennneuber
glennneuber merged commit 85ebcb7 into main Jul 14, 2026
9 checks passed
@glennneuber
glennneuber deleted the feat/gemma4-image-max-tokens branch July 29, 2026 05:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant