Repository navigation
TASK: merge upstream v0.34.1 into main (llama.cpp b10864, MLX d9add9d1 + MLX-C carry patch, scoped lifetimes) - #302
Merged
Merged
Conversation
Loading GGUF metadata is an expensive operation. Two caches had evolved to mitigate this, and the two capability implementations produced inconsistent results for some models. This PR now extracts the metadata once per blob into a file at <OLLAMA_MODELS>/metadata/sha256-<hex>.json. Only arrays over 4096 elements and non-finite floats are left out. Direct Capabilities() discovery now costs us instead of ms. /api/tags can build directly from manifests and the extracted metadata.
KV snapshots must cover a node's edge exactly. Recurrent and sliding-window state is only useful at a node's end, and a node may have none: a request resuming there lands on the previous checkpoint and begin schedules a capture at the match. The header claimed every node carries its snapshots from creation, which a node split out of an existing edge at close cannot. That hid a gap: when a response is a prefix of a stored one, close lands on the split-off head with the caches resting at its end, and pageOut skipped the capture because the node already had a KV snapshot. Restate the header as the rules that hold, and make pageOut capture whatever layers a node is missing. The scheduling comment also said eviction preserves user nodes; it only resists compaction.
The tokens of a non-causal media item attend to each other in both directions, so the item has to be evaluated in one forward. Prefill honors that when it picks chunk boundaries, but the prefix cache did not: a snapshot could be taken partway through an item, and a request that resumed there would evaluate the rest of the item alone and compute different attention for it. Snapshots scheduled inside a non-causal item now land at its end, and a match that ends inside one resumes at its start.
When a request resumes partway through a cached edge, the node holding that edge was dropped from the active path, because the path has to end at the live offset for close and the prefill captures to extend the trie from its last node. Off the path, the node was an ordinary leaf with a stale last-used time, so eviction removed it first. The captures taken during that request start at the resume offset, but attach rebuilds the missing node from the path's last node, so the new node's edge begins earlier than its KV snapshot. A later request resuming there was refused by the KV cache and re-prefilled from scratch. If eviction first merged the node into its parent, the two snapshots were concatenated as if adjacent, and the restore reported a hit while the buffer held tokens from other positions. Split the node at the live offset instead. The head stays on the path and gets the last-used update. Only the unused tail can be evicted, and losing it costs nothing. The split only happens when every layer can rewind into the edge, so it never involves a recurrent layer, and the head gets the same KV-only snapshots a close-time split already produces. When the request follows the edge, compaction merges the halves back.
…budget
Eviction skipped every node on the active path, so a conversation's own
turn checkpoints were never reclaimed no matter how far over budget the
trie was. On models with sliding-window or recurrent layers each turn's
checkpoint is a full copy of that state, 800 MiB per turn on
gemma4:31b-mlx, and a long chat grows without bound. The scheduler then
counts that memory as in use and evicts the model to load anything
else.
Only the frontier and branch points are protected now. Any other node,
active or not, is evicted least recently used first. On the active path
that merges a turn into the next one: the merged node keeps the newer
whole-state, and the KV snapshots there are lazy views of the live
buffer, so nothing is copied. Rewinding to an evicted turn resumes at
the newest surviving checkpoint before it.
qwen3.8:27b-mlx on an M5 Max, the same short question every turn with
24 tokens generated per reply, 8 GiB budget, 17.2 GiB of weights:
turn | before: paged out nodes reported | after: paged out nodes reported
11 | 4.61 GiB 33 21.6 GiB | 4.61 GiB 33 21.6 GiB
21 | 7.91 GiB 56 24.9 GiB | 7.92 GiB 56 24.9 GiB
31 | 8.46 GiB 60 25.5 GiB | 7.94 GiB 56 25.0 GiB
41 | 9.90 GiB 70 26.9 GiB | 7.96 GiB 56 25.0 GiB
50 | 11.19 GiB 79 28.2 GiB | 7.98 GiB 56 25.0 GiB
Fixes ollama#17783
…re loaded On Apple silicon the scheduler's free-memory figure for the GPU is the Metal working set minus what Ollama's own runners report. It does not see memory held by other applications, so a second MLX model can pass the fit check on a machine that is already short of memory, and the load pushes the system into swap and compression. While other models are loaded, the MLX fit check now also bounds the available memory by the system's free memory on shared-memory GPUs, the same rule llama-server loads already apply. A miss evicts an idle model and retries instead of starting the load. First loads are unchanged: with nothing else loaded, the model loads against the working-set figure alone, as both engines do today. The check also does not cover memory that grows after load, such as KV caches and prefix-cache snapshots.
…s the next model The scheduler starts the next load as soon as Close returns. The MLX client sent SIGINT, gave the process five seconds, then sent SIGKILL and returned without waiting, so a runner that could not take the signal was still exiting, with its memory still held, when the next load began. The runner has no signal handler, so SIGINT was already a kill. Load also started the process and recorded it without the client's mutex, so a Close racing with a load at server shutdown could find nothing to stop and leave the runner it missed running. Close now kills the process and waits for it to be reaped, as the llama-server client does. Load starts and records the process under the mutex and refuses to start once Close has run.
The bindings freed arrays by sweeping everything not pinned, so freeing anything required knowing what every other caller still held, and code that never swept accumulated until memory ran out. The prefix cache's eviction of a long stored path did exactly that: each merge copied the KV snapshots and nothing freed the consumed copies until the request ended, which drove a second long request past physical memory. Every array now belongs to a scope. A function scope, entered with Scoped or one of the ScopedEval forms, frees what was created in it when the function returns; results leave only by being returned. A held scope is closed by its holder and frees what was attached to it. A graph is built in a function scope and evaluated after it, so the eval frees each intermediate as it consumes it. Pin, Unpin, Sweep, and the array list's mutex are gone. On an M5 Max with qwen3.8:27b-mlx, the second 84k-token request after a stored one peaks at 35 GB instead of 57 GB; the cold path is unchanged. The copies themselves are untouched, so restoring an owned path can still exceed memory.
Models and layers checked optional weights for nil and also for a handle that no longer refers to an array, and evaluation and weight collection skipped such handles. No path produces one: a missing tensor is nil, and a handle only loses its array when its scope frees it, after which using it is a bug. The nil checks stay; the validity check is internal to the bindings now.
Loading a model can transform tensors after reading them: qwen3.5 models pack their linear-attention projections into one layout, and MoE models fuse the gate and up expert stacks. The buffers those transforms consume go back to MLX's allocator pool rather than to the system, and nothing releases the pool until the first request finishes. On qwen3.8:27b-mlx that is 2.15 GiB held idle on top of 16.9 GiB of weights, counted in the runner's reported memory the whole time. Clear the pool once the weights are evaluated. Models whose tensors load unchanged, such as gemma4, leave nothing in the pool and are unaffected.
Gemma3n's MobileNetV5 projector silently produces corrupted image embeddings on the CPU backend - no error, the model just describes the wrong image (reproduced on llama.cpp b10760; gemma4's encoder is fine on CPU). Without this guard the existing partial-offload, limited-VRAM, and OOM-retry fallbacks would pick the CPU projector on exactly the small GPUs where gemma3n lands.
* MLX: version bump * mlx: support ModelOpt global scales in MoE models * address comments * address comments
…14969) * create: add server-side MLX imports and drop GGUF conversion Support safetensors imports through the MLX create pipeline both locally and on the server, including remote upload/staging, draft layer handling, cancellation propagation, transfer limits, and shared manifest/blob writing. Limit GGUF create to wrapping existing GGUF inputs into Ollama manifests. Remove the in-tree safetensors-to-GGUF converter, server quantization path, and converter-only dependencies so GGUF conversion and quantization stay in llama.cpp tooling. Keep the MLX path focused on supported safetensors model creation with validation before MLX work, and expose that flow without the --experimental CLI gate. * address comments * add client side gguf create fast path * address comments * rebase adjustments
Includes quantized matmul corruption fix, which impacts nvfp4 multimodal models (gemma4 vision towers). Deferring wiring up the new MLX-C thread-local stream/sync APIs for now.
Drop support for creating new models with typical_p parameters, while retaining support for existing GGUF models with the setting.
Upstream v0.34.1 (38fdb5d): llama.cpp b10760 -> b10864, MLX ce916dbb -> d9add9d1 (which carries our idle-core pin, ollama#4452, and reaches past the ollama#4458 gather_qmm wall through upstream's own MLX-C carry patch in mlx/compat/), MLX-C c74db530 -> ebc88f10. Ten conflicts. MLX_VERSION: upstream's. convert/ (deleted upstream, GGUF conversion is llama.cpp tooling now) and server/model_list_cache.go (deleted; our Nemotron-text-only and gemma4-safetensors capability rules already live in images.go's filterUnsupportedCapabilities): deletions accepted. llm/llama_server.go: upstream's repeat-limit error in our (result, error) arity. x/mlxrunner/client.go: our admission signature and fields with upstream's closed flag, CheckRuntime and the locked closed-check before Start; upstream's integrated-GPU clamp folded into admit(), which now takes systemInfo. x/mlxrunner/runner.go: model setup lives in upstream's loadModel now; only configureCacheLimit() kept, after the weights are resident. x/mlxrunner/pipeline.go: our guardClose teardown on upstream's scoped shape, logging memory peak/held; our stop-sequence handling ported into upstream's scoped stream loop, which already records before streaming. The two server test files: our fixtures on upstream's gguftest types (ggml.KV is gone). Upstream replaced Pin/Unpin/Sweep/LogArrays with scoped lifetimes. Three files of ours used them and are ported: the qqmm bench (function scopes per shape, per input and per timed call; QuantizedMatmul gained a global-scale argument with the carry patch) and the two vision tests (the worker's stop hook no longer sweeps). A prefix-cache test fixture gains the PreparedItem upstream's non-causal boundary rule now reads on every item. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… preflight pins move, task doc runnerlog.py takes the runner's new per-request `memory peak=… held=…` line as a request boundary and exposes `held`; summarize_retained_memory.py prints held and its step between requests on such logs, since the per-array listing and `active − tracked` no longer exist past v0.34.0. Two tests; README entry. cuda-dynres-903 pins move to the fold's payload: llama_cpp_build 5d806aa25 (b10864) and mlx_build d9add9d11f31… in full, as the profile tests require. The task doc records the assessment, the conflict table, the ports and the decisions; gates are pending the image. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Between b10760 and b10864 upstream changed the gemma4 projector case's set_limit_image_tokens(40, 280) to (70, 1120) and dropped the comment above it. Both were context lines of 004's clip.cpp hunk, so the patch stopped applying and the first image build died in the vulkan stage at 3.5 minutes. The insertion point had not moved. Re-cut on a real b10864 checkout: the added and removed lines are identical to before (+67/-8); only context and hunk offsets differ. The full six-patch series applies clean from a clean checkout in build order with plain git apply, which is what apply-git-patches.cmake and CI's patches jobs do. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…age-token limits Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…34.0 fold's identity The header still said the current fold was v0.33.3-dynres and the deployed build v0.33.2-dynres.1; v0.34.0-dynres has been the fold and the deployed build since 2026-09-14. The "what tested green" matrix is regenerated by release_matrix.py from that build's full preflight run (21 pass, 7 skip, measured on the container that serves). "What differs from upstream, concretely" is rewritten one row per capability against upstream v0.34.1 at llama.cpp b10864, grouped by what the row is for: vision correctness on the llama.cpp path, structured output and generation control, serving and scheduling, and measurement. Two rows changed since it was written: upstream adopted the fork's gemma4 default limits (70/1120) in b10864, so only the budget fill is still ours; and upstream added its own integrated-GPU admission bound, which now sits inside admit(). Both upstream tickets the table cites are still open. Verified before writing: nemotron's canvas is still fixed at b10864, the MLX runner upstream has no stop-sequence handling, and its KV cache type is one global env. The task doc's directory-level table is replaced by a pointer to this one plus the size of the divergence, which is what the merge needs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The README row for think + format said upstream constrains the whole generation. It does not: since v0.34.1 upstream's routes layer defers the grammar until the thinking→content transition and folds pass-one metrics into the final response. The fork's flow is a superset — the marker-stop pass one, the metrics reconstruction when a runner does not report, the second pass pinned to pass one's truncation window — and the row now says so. That makes it a retirement candidate, which is what the new register is for: every item the fork carries, what retires it, and the test that gates its deletion, reviewed at each fold. x/structured heads it with Glenn's decision: carried until parity against upstream's engine is tested, deleted if there is no regression. The README's "fixes carried until upstream takes them" bullet links the register. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
x/structured (ADR 0009/0013) has had no importers since ADR 0033 adopted upstream's engine. Glenn's condition for deleting it is a test against that engine with no regression. This is the test: both engines driven live over a byte-level vocabulary (one piece per byte plus EOS, so xgrammar's token matcher and x/structured's byte machine see the same input), every verdict computed on both sides and diffed, differences classified rather than transcribed. It skips wherever libollama_xgrammar.so is not present and accepts OLLAMA_XGRAMMAR_LIBDIR for a host without a payload. Measured on xgrammar v0.2.5: 108 agreements, 0 regressions. Whitespace policy differs where llama.cpp's grammar never allowed space before a colon; -0 is refused by xgrammar's integer grammar and accepted by its number grammar; prefixItems is upstream-only; and allOf with more than one branch degrades to a permissive object on xgrammar - required properties, per-branch types and property order are not enforced, per the converter's own warning - which the test reports as a known limitation rather than a failure, so the decision stays a decision. ADR 0013's property holds without its guard: the five unbounded schemas compile in 3-30 ms within 1 MiB. The retirement register carries the result and the recommendation. Delete this file with the package. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ty gate found x/structured (ADR 0009, 0013) has had no importers since ADR 0033 adopted upstream's xgrammar engine, and ADR 0033 named its deletion as the follow-up. Glenn's condition was a test against that engine with no regression. The parity gate ran both engines live over a byte-level vocabulary on xgrammar v0.2.5: 108 agreements, 0 regressions — no schema and no output the package accepted is refused by xgrammar. Deleting it changes no request path; it removes 3,829 lines of a second grammar engine every fold merged around. The gate imported the package and goes with it. What it found is kept as engine_behaviour_test.go, xgrammar-only characterisation tests that pass while the measured behaviour holds and fail, naming the record to update, when an xgrammar bump changes it: allOf with more than one branch is permissive on xgrammar (required, per-branch types and order unenforced — its converter says allOf support is still ongoing), the integer grammar refuses -0 where the number grammar takes it, and whitespace before a colon is admitted where llama.cpp's grammar never allowed it. ADR 0013's property stays as a budget test: the five schemas the deleted guard refused compile in 3-30 ms within 1 MiB, because xgrammar compiles repetition lazily. ADR 0033 carries the amendment; the retirement register moves the package to retired; the task doc records it as decision D6. go mod tidy is unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ve carried 070580c deleted x/structured and the parity gate but staged nothing else: the first pathspec of the git add named the directory that had just been removed, and the failure stopped the rest from being staged. This is the rest. engine_behaviour_test.go keeps what the gate found, xgrammar-only: allOf with more than one branch is permissive, the integer grammar refuses -0 where the number grammar takes it, whitespace before a colon is admitted, and ADR 0013's bound holds as a budget. Each pin fails, naming the record to update, when an xgrammar bump changes the behaviour. ADR 0033 carries the amendment, the retirement register moves the package to retired, the task doc records D6. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Preflight PASS 21/7 from the fold worktree. GGUF think-off against the 0.34.0 fold: no quality cell regressed, e2b and e4b recovered, one e2b cell tripped upstream's new token-repeat error, two n=1 movers under repeat. Qwen2.5-VL: identical cell for cell. MLX: admission refused three models under 50 GB of foreign GPU0 occupancy, so that leg is invalid and re-runs alone once the GPU frees. Throughput columns not read: shared GPU. The gemma4 comment in llama_server.go named llama.cpp's old (40, 280) defaults; b10864 made them (70, 1120). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
MLX think-off re-run alone once GPU0 had 85 GB free: five models, no errors, every scored cell equal to the deployed 0.34.0 build, one 35b-a3b name_bbox cell moved at n=1 (repeats running). The run's `memory` lines are the first real ones: slog writes msg=memory unquoted, runnerlog.py accepted only the quoted form and read nothing — widened, with the real line pinned as a test. held is bounded on gemma4 12b/26b/31b and released on qwen3.8; on qwen3.6:35b-a3b it grows 0.60 GiB per request for all 28 requests with no eviction, which this log cannot split between trie fill and a leak; a trace-level probe is queued and the task doc says so. README: current fold v0.34.1-dynres, deployed stays v0.34.0-dynres, the matrix regenerated from the candidate's full preflight run (preflight-runs/full-0341.json on the array; runs/ is gitignored). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`held` alone cannot tell a prefix trie filling towards its cap from a leak. The 0.34.1 runner still logs the trie's accounting line at trace level, so the parser keeps its active_size and the retained-memory summariser prints held − trie and its step: the weights plus whatever nothing tracks, with the trie's growth removed. One test, 134 pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…X regression Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…dence Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…b removes it Trace probe: held − trie grows +0.23 GiB per request on 35b-a3b and up to 1.1 GiB per request on 27b (released later in 1 GiB steps). Control with OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0: flat for all 28 requests. The summariser's trie figure is active + paged out. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ar than the deployed build; the knob takes it to flat Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…(upstream ec3cc23) Upstream v0.34.2's fix, re-rooted under x/: the decode loop released MLX's pool of freed buffers only when the generated count landed exactly on a multiple of 256, which a speculative round steps over, so the pool was never released and the runner's footprint climbed over a generation. Measured on this fold with the trace suite (qwen3.6:35b-a3b-nvfp4, 28 requests, drafting under a grammar on): resident − trie grew +5.3 GiB without it and +2.4 GiB with it, against +1.8 GiB on the deployed 0.34.0 build and +0.1 GiB with OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0. The hunk removes the fold's memory regression against the deployed build; the drafting retention that both builds share stays behind the knob (ADR 0033, D5). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…d carried; memory re-measure closed Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
glennneuber
marked this pull request as ready for review
September 18, 2026 00:10
This was referenced Sep 18, 2026
glennneuber
added a commit
that referenced
this pull request
Sep 18, 2026
Bisected on the encoder rather than the tier: the encoder measurement is deterministic and answers in seconds, while the tier needs a scored benchmark on a quantity that drifts hundreds of tokens between runs. Midpoint 907deff (v0.34.0 fold + #300 + #301, pins b10760 / MLX ce916dbb) is bit-identical to 0.33.2 on the encoder for both affected models, scores 4 on the think-on 9px tier, and has a greedy eval range of one token. So the image embeddings, the OCR tier and the run-to-run spread all flip at the same commit boundary, and all of them are in #302. That eliminates the whole v0.34.0 upstream fold, #300 and #301. It does not prove one cause; it rules out the findings being scattered and needing separate hunts. Also records a separation route that does NOT work: pairing 8a7ba94's Go with ce916dbb's MLX payload segfaults at package init, because the new Go calls mlx_install_capture_handler and the old MLX-C lacks the symbol. The halves are hard-coupled and cannot be split by swapping payloads. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fold of upstream v0.34.1 (tag
38fdb5dd5, 2026-09-14; 23 first-parent commits; 346 files, +17k/−57k) on top ofv0.34.0-dynres. Record:docs/maxusai/tasks/upstream-sync-0.34.1.md. Gates (2026-09-18): imagesync-0.34.1(0.34.0-dynres-6-gfb18f5c) built; preflight PASS 21/7; GGUF think-off no scored-cell regression (e2b/e4b recovered, two deterministic movers pinned at n = 4); Qwen2.5-VL identical cell for cell; MLX think-off every scored cell equal, one 35b-a3b cell at n = 1 under repeat. Memory (closed 2026-09-18): under drafting + grammar the fold retained ~3× the deployed build on the recurrent-state models (35b-a3b: +5.3 vs +1.8 GiB outside the trie over 28 requests); upstream'sec3cc2307(v0.34.2) brings it to +2.4 and is cherry-picked here, measured on a Go-only swap;OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0takes it to +0.1 and stays the mitigation for what both builds share. The one moved MLX cell is bimodal across cold loads (n = 4), not a regression. Ready for review; merge on the maintainer's word.Pins
b10760(0f3a71be1)b10864(5d806aa25)ce916dbbd9add9d1— carries our idle-core pin (ollama#4452), 24 commits onc74db530ebc88f10The
#4458gather_qmm global_scalewall that capped our MLX pin is gone: upstream hit it too and shipsmlx/compat/0001-mlx-c-qmm-global-scale.patch, applied through thecmake/apply-git-patches.cmakemachinery fromcmake/local.cmake(which we never changed) and copied into the mlx stage by the Dockerfile. Nothing extra is carried. The swap-validity input list gainsmlx/compat/.The one structural change: scoped array lifetimes
Upstream replaced
Pin/Unpin/Sweep/LogArrayswith function and held scopes (x/mlxrunner/mlx/scope.go). Three files of ours used the old API and are ported — theqqmmbench (per-shape, per-input and per-timed-call scopes;QuantizedMatmulalso gained a global-scale argument) and the two vision tests. Thetensors total:trace line is gone with it; the runner now logsmemory peak=… held=…per request.runnerlog.pyreads that as a request boundary andsummarize_retained_memory.pyprintsheldand its step on such logs (two tests). The v0.34.0 drafting-leak finding was of the old model and is re-measured on this build, not carried.Ten conflicts
MLX_VERSIONconvert/convert_nemotron_h.go+ testserver/model_list_cache.goimages.go'sfilterUnsupportedCapabilitiesllm/llama_server.go(result, error)arityx/mlxrunner/client.goclosedflag,CheckRuntime, locked closed-check beforeStart; upstream's integrated-GPU clamp folded intoadmit(), which now takessystemInfo(10 test callers)x/mlxrunner/runner.goloadModel; onlyconfigureCacheLimit()kept, after the weights are residentx/mlxrunner/pipeline.goguardCloseteardown on upstream's scoped shape; our stop-sequence handling ported into upstream's scoped stream loop, which already records before streamingserver/routes_debug_test.go,routes_generate_test.gogguftesttypes (ggml.KVis gone)Auto-merged, proven by tests rather than read:
speculate.go(theOLLAMA_MLX_DRAFT_UNDER_GRAMMARknob is intact; upstream still drafts under a grammar, ADR 0033 unchanged),prefix_cache.go,sched.go,Dockerfile,test.yaml,media.go,images.go. One prefix-cache test fixture gained thePreparedItemthat upstream's non-causal boundary rule now reads on every item.Gate 2
go build ./...,go vet,go mod tidyclean;go testgreen on all 13x/mlxrunnerpackages,llmandserver; gofumpt applied. Preflight: 106 tests;cuda-dynres-903pins moved to5d806aa25/d9add9d1…(full commit, as the tests require). Vision-suite: 132 tests.Not in this fold
The Metal half (held).
v0.34.2-rc1moves the MLX engine out ofx/— a separate, structural fold. PR #301 (whitespace bound, ADR 0035) touchesclient.go; merging it first means a small merge here, after means a rebase — either is fine.🤖 Generated with Claude Code