Skip to content

docs(design): salvage Gemma 4 vision-budget design notes before branch retirement - #5

Merged
glennneuber merged 1 commit into
mainfrom
docs/salvage-gemma4-design
Jul 29, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/salvage-gemma4-design

Conversation

@glennneuber

Copy link
Copy Markdown

Rescues two design records from feat/gemma4-visual-token-budgets-last-go-runner — the
only place they exist — so that branch can be retired without losing the reasoning behind
the gemma4 vision-budget work.

Why this is worth keeping

Both documents describe the original approach: implemented against Ollama's Go-native
inference runner, which upstream deleted in two stages (ollama#16031 removed
runner/ollamarunner/, ollama#17007 removed model/model.go and model/models/gemma4/). That
code was never merged and cannot be revived.

What shipped instead — PR #2 (d06138a9) — passes --image-min-tokens /
--image-max-tokens to llama-server, touching only api/types.go and
llm/llama_server.go. Different layer, different numbers:

the plan shipped
layer Go runner (ollamarunner) llama-server flags (C++ mtmd)
defaults 70 / 560 40 / 1120
ladder snap {70,140,280,560,1120} none — passed through

What survives the rewrite is the reasoning: the Google visual-token ladder, why a budget
change must force a scheduler reload, the option-naming rationale, and the base-selection
trade-off behind the -last-go-runner branch.

Not a verbatim copy

Dropped verbatim into main, a 422-line implementation plan with todos: status: pending
frontmatter reads as live guidance for work that can never be done. So:

  • HISTORICAL banner at the top of each file, with the plan-vs-shipped table above.
  • Dropped docs/design/PR_BODY.md — submission scaffolding for the closed upstream PR,
    and it restates the superseded 70/560 defaults as fact.
  • README.md rewritten as an archive index pointing at docs/maxusai/ for current
    behavior.

Two inline annotations correct claims that later proved wrong. I left the original text
unedited underneath both, so the record stays honest:

  1. "Forward-porting to current main" concluded that a forward-port would require
    patching C++ mtmd/clip, because Gemma supposedly ignores the image-token levers.
    It doesn't. PR feat(gemma4): tunable per-request vision image-token budget #2 passes them straight through with no C++ change, and
    prompt_eval_count goes 1,435 (~220 image tokens) → 2,233 (~1,015), matching the
    reference server. The section's closing instinct — "first confirm what the pinned
    mtmd actually does for Gemma today"
    — is exactly what resolved it.
  2. "Known risk" (unclamped vision position-embedding lookup at high budgets) lives in
    model/models/gemma4/model_vision.go, which no longer exists, so it's moot for the
    shipped path. Annotated as such — while noting the wide/extreme-aspect-ratio smoke test
    it recommends has still never been run against llama-server. The deployed 1120 config
    has been exercised on exactly one image.

⚠️ Merge after #3

The banners link to docs/maxusai/gemma4-budget-image.md, which arrives in #3. Merge that
one first or these two links dangle. No file overlap, so no conflict either way.

Once this lands, both feat/gemma4-visual-token-budgets branches can be deleted — the
second is a strict superset of the first, and this PR takes everything durable from it.

🤖 Generated with Claude Code

…h retirement

Rescues the two design records from feat/gemma4-visual-token-budgets-last-go-runner,
the only place they existed, so the branch can be deleted without losing the reasoning
behind the shipped feature.

Both describe the ORIGINAL approach — implemented against Ollama's Go-native inference
runner, which upstream deleted in ollama#16031 (runner/ollamarunner) and ollama#17007 (model/,
model/models/gemma4/). That code was never merged and cannot be revived. What shipped
instead (PR #2, d06138a) passes --image-min-tokens/--image-max-tokens to llama-server,
with different defaults: 40/1120, not the plan's 70/560.

Each file therefore gets a HISTORICAL banner up top so neither reads as live guidance.
Two inline annotations correct claims that later proved wrong:

- The rebase note's "Forward-porting" section concluded a forward-port would require
  patching C++ mtmd/clip because Gemma ignores the image-token levers. It does not —
  PR #2 passes them through with no C++ change and prompt_eval_count rises 1,435 ->
  2,233, matching the reference server.
- Its "Known risk" (unclamped vision position-embedding lookup) is in a file that no
  longer exists. Flagged moot for the shipped path, while noting the wide-image smoke
  test it recommends was never actually run against llama-server.

Dropped docs/design/PR_BODY.md: submission scaffolding for the closed upstream PR, and
it restates the superseded 70/560 defaults as fact.

Note: the banners link to docs/maxusai/gemma4-budget-image.md, which arrives in PR #3.
Merge that first or these two links dangle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants