Skip to content

TASK: merge upstream/main 6bba484f (v0.32.15+3) — parser-deadlock fix, metadata cache, MLX/llama bumps - #208

Merged
glennneuber merged 16 commits into
mainfrom
task/upstream-sync-2026-08-21
Aug 22, 2026
Merged

glennneuber merged 16 commits into
mainfrom
task/upstream-sync-2026-08-21

Conversation

@glennneuber

Copy link
Copy Markdown

Handoff PR — the merge is not done yet. This branch carries the full assessment in docs/maxusai/tasks/upstream-sync-2026-08-21.md. The agent picking this up should perform the merge on this branch so this PR closes with the work.

Scope

Merge upstream/main at 6bba484f (v0.32.15-3-g6bba484f, fetched 2026-08-21) into main. Merge-base is d67ad834; exactly 10 upstream commits are in scope (2026-08-18 → 2026-08-20). A dry-run git merge-tree --write-tree main upstream/main shows 8 of 10 auto-merge; two files conflict — everything below is from that dry run.

Why merge

  • e0c95a5f — mid-stream parser-error deadlock fix (the reason to do this promptly). A builtin-parser rejection mid-stream leaks the completion goroutine and never releases the runner request; retries hang silently. Only thinking mode wedged. Our fork carries its own parser changes (nemotron3nano, qwen35) and runs hours-long thinking-mode benchmark campaigns — exactly the exposure. Ships with server/routes_parse_error_test.go.
  • a5165c53 — model metadata cache keyed by manifest digest with singleflight; lowers the per-request latency floor under our req/h numbers. Provenance-relevant; covered by SPEC H11 per-cell server_version.
  • MLX pin adf21dea… → 27fec909… (two bumps + mlx-c 0.32.1 regen compat patch) and llama.cpp b10434 → b10488 — routine, but the numeric risk to the gemma4 MLX vision path gates on the golden tests below.
  • Remaining commits: 4e134213 (mac-assumption fix — we already fixed the RPATH half in-fork; their Windows LoadLibraryExA half is new and auto-merges), b8a62724 (qwen3.8 system-message folding — single-system-prompt probes explicitly unaffected), d1bd15cc (Dockerfile), b7871fc0 (desktop app only), 6bba484f (lint).

The two conflicts

  1. x/mlxrunner/mlx/CMakeLists.txt — trivial. Both sides fixed the same Mach-O-RPATH-on-ELF bug; both land on $ORIGIN. Keep ours (elseif(UNIX) + explanatory comment).
  2. server/routes.go — contained but real. Two conflict regions, both in GenerateHandler, exactly where the fork's structured-outputs marker flow lives (pass-1/pass-2, ADR 0010 transition metrics). The ChatHandler half of upstream's fix auto-merges. Resolution: thread upstream's parserErr + cancel() pattern through our extended callback — parser errors must set parserErr and cancel instead of writing to ch from inside the callback; report once after the completion returns; keep our expireRunnersForRuntimeOOM call gated on parserErr == nil; a pass-one parse error must not leave the pass-two restart armed. Full hunk-level detail in the task doc.

Acceptance criteria (in order)

  1. Merge on this branch; only the two conflicts above, resolved as specified.
  2. go test ./server/ ./model/renderers/ ./model/parsers/ green — including upstream's new routes_parse_error_test.go and our routes_generate_test.go, sched_headofline_test.go, qwen38_effort_test.go. (Known baseline: go build ./... fails on the app/dist embed — pre-existing; test the listed packages, not ./....)
  3. MLX vision golden tests (x/mlxrunner/vision_golden_test.go, 12b/26b/31b goldens) — single-owner-thread rule applies.
  4. Build an image and run the preflight harness (docs/maxusai/vision-suite/preflight/) against a benchmark host before any deploy.
  5. Merge-commit body notes that SPEC H11 server_version is the comparability boundary for benchmark cells measured on the new build.

🤖 Generated with Claude Code

dhiltgen and others added 14 commits August 18, 2026 11:53
…llama#17883)

When a builtin parser rejects model output, the completion callback wrote the
error to an unbuffered channel and returned. The callback cannot stop
generation -- it has no error return -- so the next chunk re-entered the
callback, hit the same parse error and blocked writing to a channel the
consumer had already stopped reading after emitting its 500. The completion
never returned, the goroutine leaked and the runner request was never
released, so retrying the same prompt hung with no log output until the client
gave up.

Record the parse error, cancel the completion, and report it once the
completion has returned. Parse failures landing on the final chunk were
already terminal, which is why non-thinking requests and the direct
qwen3-coder parser path failed cleanly and only thinking mode wedged.

ChatHandler and GenerateHandler share the defect: both run the same parser in
the same shape of callback behind a consumer that stops reading at the first
error. GenerateHandler had no cancel func at all, so one is added there.

Fixes ollama#17825
The default packaging was broken due to mac assumptions
leaking into windows
Ten upstream commits since merge-base d67ad83, verified by dry-run
merge-tree: eight auto-merge, two files conflict (server/routes.go in
GenerateHandler where the marker flow lives; x/mlxrunner/mlx/CMakeLists.txt
where both sides fixed the same RPATH bug). Headline item is upstream
e0c95a5, the mid-stream parser-error deadlock fix. Task doc carries the
full conflict analysis, resolution guidance, and the verification gate
order (server+renderer tests, vision goldens, preflight).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

Reviewing as consolidator. This is a good handoff — the dry-run-before-the-merge shape is right, and the two conflicts are correctly identified. One thing must be added to the acceptance criteria before anyone acts on it.

Add a build gate, and do not let a red check be clicked past

The last upstream sync, #165, merged with test (ubuntu-latest): FAILURE on the PR. It carried a merge resolution that kept both copies of a test upstream had moved:

vet: x/models/qwen3_5/qwen3_5_test.go:43:6: TestSanitizeConvWeight redeclared in this block

main did not build for roughly fourteen hours, across a dozen merged PRs, every one reviewed against a base whose tests could not run. #182 removed the duplicate.

That is this operation, and the failure class is specific to it: a sync that keeps both copies of a moved symbol reads as a plain addition in review and is invisible to a hunk-level eye. Your criterion 2 scopes tests to ./server/ ./model/renderers/ ./model/parsers/ — sensible given the pre-existing ./... build break — but the #165 defect was in x/models/qwen3_5/, which that scope does not cover.

Concretely, I would add as criterion 2a:

go vet ./x/... ./llm/... ./server/...     # duplicate-symbol / redeclaration sweep

go vet compiles without linking, so it catches redeclaration across the fork's own packages without needing the embed that breaks ./.... It takes seconds and it is precisely the check that would have caught #165.

And the process half, which no criterion can enforce: if the test job goes red on this PR, it blocks. It is one of the few jobs not gated behind vars.SELF_HOSTED_RUNNERS, it ran, it reported, and it was overridden. A control that can be clicked past is the same category of non-coverage that #137 wrote patches-ggml to fix — "coverage that only exists when a runner happens to be present is not coverage."

On the conflicts themselves

Conflict 1 is the one I can speak to directly: x/mlxrunner/mlx/CMakeLists.txt, both sides landing on $ORIGIN. Keep ours is right — ours carries the elseif(UNIX) guard and the comment explaining that @loader_path is Mach-O and the ELF spelling is $ORIGIN, which is the part a future reader needs. Worth checking after resolution that the shipped payload still carries libquadmath.so.0; the RPATH and the quadmath bundling were one fix and only the RPATH half is in conflict.

Conflict 2's resolution rule — parser errors set parserErr and cancel rather than writing to ch from inside the callback, reported once after completion returns, expireRunnersForRuntimeOOM gated on parserErr == nil, pass-one parse error must not leave the pass-two restart armed — reads correct against ADR 0010's marker flow. The last clause is the one I would test explicitly rather than infer, since a stale armed restart would surface as a spurious second pass much later than the merge.

Criterion 4 is the right one to not skip

Preflight against a benchmark host before deploy: --platform mlx-cuda now resolves and passes on the CUDA host (7 PASS / 3 SKIP, gemma4 measured), so an MLX pin bump has a real gate to fail against rather than only the golden tests. Given the pin moves twice here, that gate is the point.

13 upstream commits, 2026-08-18 -> 2026-08-20. Assessment and handoff:
docs/maxusai/tasks/upstream-sync-2026-08-21.md. Two conflicts, both
predicted by the dry run, resolved as specified there.

x/mlxrunner/mlx/CMakeLists.txt -- both sides fixed the same bug (the Mach-O
@loader_path RPATH spelling written on ELF). Kept ours: the comment
explaining the failure, and the stricter if(APPLE)/elseif(UNIX) guard over
upstream's if/else. Upstream's sibling dynamic.c change (Windows
SetDllDirectoryA + LoadLibraryExA DLL-search fix) auto-merged and is kept.

server/routes.go -- both hunks land in GenerateHandler, on the fork's
structured-outputs marker-flow state machine. Kept our pass-one/pass-two
loop and threaded upstream's e0c95a5 parser-error fix through it. A builtin
parser rejecting output mid-stream used to write the error to ch from inside
the completion callback; the callback cannot stop generation, so the next
chunk re-entered, hit the same error, and blocked forever on a channel the
consumer stopped reading after its 500 -- leaking the goroutine and never
releasing the runner. Both builtinParser.Add sites now record parserErr and
cancel() instead. Two gates the flat upstream handler does not need:

  - the post-completion error branch skips reporting when parserErr is set,
    so a parse-induced cancel is not mistaken for a runner fault and does
    not reach expireRunnersForRuntimeOOM;
  - the pass-two restart is held behind parserErr == nil. The marker flow
    sets state = ReadyToApply *before* parsing the closed thinking, so
    without this a parse failure there would leave pass two armed and
    restart generation on an error.

Verified: go build ./server/ ./model/... ./llm/ ./api/, and
go test ./server/ ./model/renderers/ ./model/parsers/ green -- including
upstream's new TestGenerateParseErrorMidStreamDoesNotWedge and
TestChatParseErrorMidStreamDoesNotWedge, and the fork's
TestSchedAbandonedEvictionDoesNotBlockQueue,
TestSchedAbandonedOOMEvictAllDoesNotBlockQueue and
TestQwen38ReasoningEffortMapping (unchanged by b8a6272, which only
re-renders multi-system / developer-role conversations).

PROVENANCE: server_version (SPEC H11) is the comparability boundary for any
benchmark cell measured on a build from this merge. a5165c5 adds a model
metadata cache keyed by manifest digest, which lowers the per-request
latency floor, and b8a6272 changes qwen3.8 system-message rendering. MLX
pin adf21dea -> 27fec909 (two bumps) and llama.cpp b10434 -> b10488 also
land here, so numeric vision results are not comparable across this commit
without a re-baseline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

Merge done on this branch; acceptance criteria worked

Merge commit 76918a7a, parents 9ca7bfda + 8f912415. Upstream had moved 3 commits past the 6bba484f this doc was written against — app: add claude desktop app, polish onboarding layout, launch: npx fallback for DeepSeek Harness. All three are app/, cmd/launch/, anthropic/ only, nothing under server/, x/, model/, llama/, ml/, Dockerfile or CMake, so the assessment held. 13 commits in scope, exactly the two predicted conflicts, no others.

1. Merge — done

x/mlxrunner/mlx/CMakeLists.txt — kept ours (comment + stricter if(APPLE)/elseif(UNIX)), as specified. Upstream's sibling dynamic.c Windows SetDllDirectoryA/LoadLibraryExA fix auto-merged and is kept.

server/routes.go — kept our pass-1/pass-2 loop, threaded upstream's e0c95a5f pattern through it. Both builtinParser.Add sites now set parserErr + cancel() instead of writing to ch. Two gates the flat upstream handler does not need:

  • the post-completion error branch skips reporting when parserErr is set, so a parse-induced cancel is not mistaken for a runner fault and never reaches expireRunnersForRuntimeOOM;
  • the pass-two restart is held behind parserErr == nil. This is the trap the doc flagged: the marker flow sets state = ReadyToApply before parsing the closed thinking, so without the gate a parse failure there leaves pass two armed and restarts generation on an error.

2. Tests — green

go test ./server/ ./model/renderers/ ./model/parsers/ passes. Named tests run individually rather than inferred from the package result:

test result
TestGenerateParseErrorMidStreamDoesNotWedge (new, upstream) PASS
TestChatParseErrorMidStreamDoesNotWedge (new, upstream) PASS
TestSchedAbandonedEvictionDoesNotBlockQueue PASS
TestSchedAbandonedOOMEvictAllDoesNotBlockQueue PASS
TestQwen38ReasoningEffortMapping PASS, unchanged

b8a62724 needed no test update: it only re-renders multi-system / developer-role conversations.

3. MLX vision goldens — 31b passes, 12b/26b unrunnable here

TestVisionGoldenParity on gemma4:31b-nvfp4 against the new MLX pin (adf21dea → 27fec909):

mean -0.00621 (golden -0.00650)   std 1.3243 (golden 1.3236)   norm_mean 96.999 (golden 96.947)
max sampled element delta: 0.1055
--- PASS (965.79s)

The MLX bump did not move gemma4 vision numerics. 12b and 26b could not run: gemma4:12b-nvfp4 and gemma4:26b-nvfp4 are not pulled on this host (only 26b-a4b-it-q4_K_M, which resolves to a different golden path and skips). I did not pull them — / was at 99%. That is a real gap in this criterion, not a pass.

4. Preflight — FAIL=1, PASS=17, and the one failure is the point

maxusai/ollama:sync-0.32.15, version 0.32.14-dynres-108-g76918a7, matched profile cuda-dynres-903.

FAIL  payload_pin   expected 7e4c0a968 (b10434) / actual 9d77fa172 (b10488)
PASS  version, image_tag, go_patch_marker, endpoint_exclusive
PASS  text_baseline / token_ladder (5/5 within +/-2) / payload_proof / think_format
      -- all three arches: nemotron_h_omni, gemma4, qwen35
PASS  pinned_budget [nemotron_h_omni]  pinned 3328 -> 3270 (ceiling 3328, +2 markers)
SKIP  pinned_budget [gemma4], [qwen35] -- no expectation recorded (pre-existing gap)

Every behavioural assertion passes on the new payload; the ladders did not move. The sole failure is the identity pin refusing to let ladders measured on b10434 imply a pass on b10488 — exactly what it exists for.

So this build is not deploy-validated yet. Per the harness README the fix is a deliberate re-baseline: run measure_ladder.py per arch with --container and --stride, add a new [profiles.…] keyed to 9d77fa172 (the patchset 001/002/004/005/903 is unchanged, the payload is not), and paste the generated [expect.…] blocks whole. I have not edited any expected value to go green, and have not touched expectations.toml.

Patch 903 is still present and applied; whether upstream ggml-org/llama.cpp#27044 is still open on b10488 is worth confirming as part of the re-baseline.

5. Provenance — done

The merge commit body records server_version (SPEC H11) as the comparability boundary, naming the metadata cache (a5165c53), the qwen3.8 renderer change, MLX adf21dea → 27fec909, and llama.cpp b10434 → b10488.

Not done

  • 12b / 26b goldens (models absent).
  • The expectations.toml re-baseline, and therefore deploy.

…2.15 sync

The upstream sync moved llama.cpp b10434 -> b10488, so payload_pin failed on
the merged build: expected 7e4c0a968, actual 9d77fa172. That is the check
working -- it refuses to let ladders measured on one payload imply a pass on
another.

Re-measured before changing the pin, not after. measure_ladder.py against
0.32.14-dynres-108-g76918a7 returned byte-identical rows for all three arches:

  nemotron_h_omni  [266, 266, 578, 2306, 3270]   budgets 256/3328    stride 32
  gemma4           [1102 x 5]                    budgets 70/1120     stride 48
  qwen35           [1034, 1034, 1034, 2306, 4082] budgets 1024/4096  stride 32

Same budgets, same pixel windows, same strides as the b10434 rows. So no
expected value is edited here -- only the payload identity, plus the provenance
explaining why. The prose in the expect blocks (ADR 0011 rule 4 discussion on
qwen35, pinned_not_applicable, the B8 prefix notes) is left intact rather than
overwritten with generator output that does not carry it.

Not a new profile: the README's "add a new [profiles.<id>]" path is for a
changed patchset or a new platform. The patchset (001/002/004/005/903) still
describes this build exactly, and resolve_profile keys only on
(platform, version) -- two cuda profiles sharing the 0.32.x-dynres pattern
would resolve by file order, which is worse than useless.

903 is still required: ggml-org/llama.cpp#27044 re-checked 2026-08-21, still
open, so b10488 carries the MMQ ids-path defect exactly as b10434 did.

Verified: test_verdicts.py 43/43, and the full preflight (including the
pinned-budget probe) on the merged build is now PASS=18 SKIP=2 FAIL=0,
against FAIL=1 before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@glennneuber

Copy link
Copy Markdown
Author

All five acceptance criteria now green

Follow-up to the previous comment. The two gaps it flagged are closed.

The models were there all along — on a different disk

OLLAMA_MODELS is unset on every ollama container, so they all use the in-container default backed by the docker_ollama_data volume. That volume is bind-backed, per ~/deployments/docker/ollama/docker-compose.yaml:

device=/mnt/4TB_SN850X_RAID1_BTRFS/opt/ollama/.ollama   type=none   o=bind

So the store is /mnt/4TB_SN850X_RAID1_BTRFS/opt/ollama/.ollama/models (657 GB, 82 GB free) and not on / at all. My earlier reason for not pulling 12b/26b — that / was at 99% — was simply wrong; the pull never touches /. Both tags existed upstream (HTTP 200) and pulled cleanly.

3. MLX vision goldens — all three arches PASS

TestVisionGoldenParity against the bumped MLX pin (adf21dea → 27fec909):

model mean (golden) std (golden) norm_mean (golden) max element delta time
gemma4:12b-nvfp4 -0.02562 (-0.02560) 2.4605 (2.4613) 151.309 (151.352) 0.0625 209s
gemma4:26b-nvfp4 -0.00024 (-0.00037) 1.6290 (1.6289) 86.343 (86.336) 0.0625 696s
gemma4:31b-nvfp4 -0.00621 (-0.00650) 1.3243 (1.3236) 96.999 (96.947) 0.1055 966s

The MLX pin bump moved nothing on any of the three. That was the numeric risk this criterion existed to catch, and it is clear.

4. Preflight — PASS=18, SKIP=2, FAIL=0

Re-baselined in 64600179. The process, in the order the harness README requires:

  1. payload_pin failed on the merged build: expected 7e4c0a968 (b10434), actual 9d77fa172 (b10488).
  2. Re-measured before touching the pin, via measure_ladder.py per arch with --container and --stride:
nemotron_h_omni  [266, 266, 578, 2306, 3270]     budgets 256/3328    stride 32
gemma4           [1102 x 5]                       budgets 70/1120     stride 48
qwen35           [1034, 1034, 1034, 2306, 4082]   budgets 1024/4096   stride 32
  1. Byte-identical to the b10434 rows — same ladders, budgets, pixel windows, strides. So no expected value was edited; only llama_cpp_build moved, plus provenance explaining why.
  2. Full preflight re-run: PASS=18 SKIP=2 FAIL=0 (VERDICT: PASS), and test_verdicts.py 43/43.

Two judgement calls worth review:

  • Updated the existing profile rather than adding a new one. The README says add a new [profiles.<id>] for a changed patchset or new platform. The patchset (001/002/004/005/903) still describes this build exactly, and resolve_profile keys only on (platform, version) and returns the first regex match — two cuda profiles sharing the 0.32.x-dynres pattern would resolve by file order, which is worse than a stale pin.
  • Left the expect-block prose intact. The generator output does not carry the ADR 0011 rule 4 discussion on qwen35, pinned_not_applicable, or the B8 prefix notes. Overwriting the blocks wholesale would have destroyed all of it for zero numeric change.

903 is still required: ggml-org/llama.cpp#27044 re-checked 2026-08-21, still open.

Status

# criterion result
1 merge, two conflicts, resolutions as specified done
2 go test ./server/ ./model/renderers/ ./model/parsers/ green, named tests run individually
3 MLX vision goldens 12b / 26b / 31b 3/3 PASS
4 image + preflight before deploy PASS=18 FAIL=0
5 server_version provenance in the merge commit done

Branch task/upstream-sync-2026-08-21 at 64600179, ready for review. Image maxusai/ollama:sync-0.32.15 (version 0.32.14-dynres-108-g76918a7) is built and preflight-validated on the CUDA host; not deployed — no container was cut over.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants