Skip to content

TASK: merge upstream v0.34.1 into main (llama.cpp b10864, MLX d9add9d1 + MLX-C carry patch, scoped lifetimes) - #302

Merged
glennneuber merged 41 commits into
mainfrom
task/upstream-sync-0.34.1
Sep 18, 2026
Merged

glennneuber merged 41 commits into
mainfrom
task/upstream-sync-0.34.1

Conversation

@glennneuber

@glennneuber glennneuber commented Sep 17, 2026 •

Copy link
Copy Markdown

Fold of upstream v0.34.1 (tag 38fdb5dd5, 2026-09-14; 23 first-parent commits; 346 files, +17k/−57k) on top of v0.34.0-dynres. Record: docs/maxusai/tasks/upstream-sync-0.34.1.md. Gates (2026-09-18): image sync-0.34.1 (0.34.0-dynres-6-gfb18f5c) built; preflight PASS 21/7; GGUF think-off no scored-cell regression (e2b/e4b recovered, two deterministic movers pinned at n = 4); Qwen2.5-VL identical cell for cell; MLX think-off every scored cell equal, one 35b-a3b cell at n = 1 under repeat. Memory (closed 2026-09-18): under drafting + grammar the fold retained ~3× the deployed build on the recurrent-state models (35b-a3b: +5.3 vs +1.8 GiB outside the trie over 28 requests); upstream's ec3cc2307 (v0.34.2) brings it to +2.4 and is cherry-picked here, measured on a Go-only swap; OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0 takes it to +0.1 and stays the mitigation for what both builds share. The one moved MLX cell is bimodal across cold loads (n = 4), not a regression. Ready for review; merge on the maintainer's word.

Pins

v0.34.0-dynres this branch
llama.cpp b10760 (0f3a71be1) b10864 (5d806aa25)
MLX ce916dbb d9add9d1 — carries our idle-core pin (ollama#4452), 24 commits on
MLX-C c74db530 ebc88f10

The #4458 gather_qmm global_scale wall that capped our MLX pin is gone: upstream hit it too and ships mlx/compat/0001-mlx-c-qmm-global-scale.patch, applied through the cmake/apply-git-patches.cmake machinery from cmake/local.cmake (which we never changed) and copied into the mlx stage by the Dockerfile. Nothing extra is carried. The swap-validity input list gains mlx/compat/.

The one structural change: scoped array lifetimes

Upstream replaced Pin/Unpin/Sweep/LogArrays with function and held scopes (x/mlxrunner/mlx/scope.go). Three files of ours used the old API and are ported — the qqmm bench (per-shape, per-input and per-timed-call scopes; QuantizedMatmul also gained a global-scale argument) and the two vision tests. The tensors total: trace line is gone with it; the runner now logs memory peak=… held=… per request. runnerlog.py reads that as a request boundary and summarize_retained_memory.py prints held and its step on such logs (two tests). The v0.34.0 drafting-leak finding was of the old model and is re-measured on this build, not carried.

Ten conflicts

file resolution
MLX_VERSION upstream's
convert/convert_nemotron_h.go + test deletion accepted — upstream dropped GGUF conversion (llama.cpp tooling now); our nemotron3 Omni converter fix goes with it, nothing at runtime depends on it
server/model_list_cache.go deletion accepted — our Nemotron-text-only / gemma4-safetensors capability rules already live in images.go's filterUnsupportedCapabilities
llm/llama_server.go upstream's repeat-limit error in our (result, error) arity
x/mlxrunner/client.go our admission signature and fields + upstream's closed flag, CheckRuntime, locked closed-check before Start; upstream's integrated-GPU clamp folded into admit(), which now takes systemInfo (10 test callers)
x/mlxrunner/runner.go model setup lives in upstream's loadModel; only configureCacheLimit() kept, after the weights are resident
x/mlxrunner/pipeline.go our guardClose teardown on upstream's scoped shape; our stop-sequence handling ported into upstream's scoped stream loop, which already records before streaming
server/routes_debug_test.go, routes_generate_test.go our fixtures on upstream's gguftest types (ggml.KV is gone)

Auto-merged, proven by tests rather than read: speculate.go (the OLLAMA_MLX_DRAFT_UNDER_GRAMMAR knob is intact; upstream still drafts under a grammar, ADR 0033 unchanged), prefix_cache.go, sched.go, Dockerfile, test.yaml, media.go, images.go. One prefix-cache test fixture gained the PreparedItem that upstream's non-causal boundary rule now reads on every item.

Gate 2

go build ./..., go vet, go mod tidy clean; go test green on all 13 x/mlxrunner packages, llm and server; gofumpt applied. Preflight: 106 tests; cuda-dynres-903 pins moved to 5d806aa25 / d9add9d1… (full commit, as the tests require). Vision-suite: 132 tests.

Not in this fold

The Metal half (held). v0.34.2-rc1 moves the MLX engine out of x/ — a separate, structural fold. PR #301 (whitespace bound, ADR 0035) touches client.go; merging it first means a small merge here, after means a rebase — either is fine.

🤖 Generated with Claude Code

hoyyeva and others added 30 commits September 9, 2026 15:02
Loading GGUF metadata is an expensive operation. Two caches had evolved to
mitigate this, and the two capability implementations produced inconsistent
results for some models.

This PR now extracts the metadata once per blob into a file at
<OLLAMA_MODELS>/metadata/sha256-<hex>.json. Only arrays over 4096 elements and
non-finite floats are left out.

Direct Capabilities() discovery now costs us instead of ms.  /api/tags can
build directly from manifests and the extracted metadata.
KV snapshots must cover a node's edge exactly. Recurrent and
sliding-window state is only useful at a node's end, and a node may
have none: a request resuming there lands on the previous checkpoint
and begin schedules a capture at the match.

The header claimed every node carries its snapshots from creation,
which a node split out of an existing edge at close cannot. That hid a
gap: when a response is a prefix of a stored one, close lands on the
split-off head with the caches resting at its end, and pageOut skipped
the capture because the node already had a KV snapshot.

Restate the header as the rules that hold, and make pageOut capture
whatever layers a node is missing. The scheduling comment also said
eviction preserves user nodes; it only resists compaction.
The tokens of a non-causal media item attend to each other in both
directions, so the item has to be evaluated in one forward. Prefill
honors that when it picks chunk boundaries, but the prefix cache did
not: a snapshot could be taken partway through an item, and a request
that resumed there would evaluate the rest of the item alone and
compute different attention for it.

Snapshots scheduled inside a non-causal item now land at its end, and
a match that ends inside one resumes at its start.
When a request resumes partway through a cached edge, the node holding
that edge was dropped from the active path, because the path has to end
at the live offset for close and the prefill captures to extend the trie
from its last node. Off the path, the node was an ordinary leaf with a
stale last-used time, so eviction removed it first.

The captures taken during that request start at the resume offset, but
attach rebuilds the missing node from the path's last node, so the new
node's edge begins earlier than its KV snapshot. A later request
resuming there was refused by the KV cache and re-prefilled from
scratch. If eviction first merged the node into its parent, the two
snapshots were concatenated as if adjacent, and the restore reported a
hit while the buffer held tokens from other positions.

Split the node at the live offset instead. The head stays on the path
and gets the last-used update. Only the unused tail can be evicted, and
losing it costs nothing. The split only happens when every layer can
rewind into the edge, so it never involves a recurrent layer, and the
head gets the same KV-only snapshots a close-time split already
produces. When the request follows the edge, compaction merges the
halves back.
…budget

Eviction skipped every node on the active path, so a conversation's own
turn checkpoints were never reclaimed no matter how far over budget the
trie was. On models with sliding-window or recurrent layers each turn's
checkpoint is a full copy of that state, 800 MiB per turn on
gemma4:31b-mlx, and a long chat grows without bound. The scheduler then
counts that memory as in use and evicts the model to load anything
else.

Only the frontier and branch points are protected now. Any other node,
active or not, is evicted least recently used first. On the active path
that merges a turn into the next one: the merged node keeps the newer
whole-state, and the KV snapshots there are lazy views of the live
buffer, so nothing is copied. Rewinding to an evicted turn resumes at
the newest surviving checkpoint before it.

qwen3.8:27b-mlx on an M5 Max, the same short question every turn with
24 tokens generated per reply, 8 GiB budget, 17.2 GiB of weights:

  turn | before: paged out  nodes  reported | after: paged out  nodes  reported
    11 |          4.61 GiB     33  21.6 GiB |         4.61 GiB     33  21.6 GiB
    21 |          7.91 GiB     56  24.9 GiB |         7.92 GiB     56  24.9 GiB
    31 |          8.46 GiB     60  25.5 GiB |         7.94 GiB     56  25.0 GiB
    41 |          9.90 GiB     70  26.9 GiB |         7.96 GiB     56  25.0 GiB
    50 |         11.19 GiB     79  28.2 GiB |         7.98 GiB     56  25.0 GiB

Fixes ollama#17783
…re loaded

On Apple silicon the scheduler's free-memory figure for the GPU is the
Metal working set minus what Ollama's own runners report. It does not see
memory held by other applications, so a second MLX model can pass the fit
check on a machine that is already short of memory, and the load pushes
the system into swap and compression.

While other models are loaded, the MLX fit check now also bounds the
available memory by the system's free memory on shared-memory GPUs, the
same rule llama-server loads already apply. A miss evicts an idle model
and retries instead of starting the load. First loads are unchanged: with
nothing else loaded, the model loads against the working-set figure alone,
as both engines do today. The check also does not cover memory that grows
after load, such as KV caches and prefix-cache snapshots.
…s the next model

The scheduler starts the next load as soon as Close returns. The MLX client
sent SIGINT, gave the process five seconds, then sent SIGKILL and returned
without waiting, so a runner that could not take the signal was still
exiting, with its memory still held, when the next load began. The runner
has no signal handler, so SIGINT was already a kill.

Load also started the process and recorded it without the client's mutex,
so a Close racing with a load at server shutdown could find nothing to stop
and leave the runner it missed running.

Close now kills the process and waits for it to be reaped, as the
llama-server client does. Load starts and records the process under the
mutex and refuses to start once Close has run.
The bindings freed arrays by sweeping everything not pinned, so freeing
anything required knowing what every other caller still held, and code
that never swept accumulated until memory ran out. The prefix cache's
eviction of a long stored path did exactly that: each merge copied the
KV snapshots and nothing freed the consumed copies until the request
ended, which drove a second long request past physical memory.

Every array now belongs to a scope. A function scope, entered with Scoped
or one of the ScopedEval forms, frees what was created in it when the
function returns; results leave only by being returned. A held scope is
closed by its holder and frees what was attached to it. A graph is
built in a function scope and evaluated after it, so the eval frees each
intermediate as it consumes it. Pin, Unpin, Sweep, and the array list's
mutex are gone.

On an M5 Max with qwen3.8:27b-mlx, the second 84k-token request after a
stored one peaks at 35 GB instead of 57 GB; the cold path is unchanged.
The copies themselves are untouched, so restoring an owned path can still
exceed memory.
Models and layers checked optional weights for nil and also for a handle
that no longer refers to an array, and evaluation and weight collection
skipped such handles. No path produces one: a missing tensor is nil, and a
handle only loses its array when its scope frees it, after which using it
is a bug. The nil checks stay; the validity check is internal to the
bindings now.
Loading a model can transform tensors after reading them: qwen3.5 models
pack their linear-attention projections into one layout, and MoE models
fuse the gate and up expert stacks. The buffers those transforms consume
go back to MLX's allocator pool rather than to the system, and nothing
releases the pool until the first request finishes. On qwen3.8:27b-mlx
that is 2.15 GiB held idle on top of 16.9 GiB of weights, counted in the
runner's reported memory the whole time.

Clear the pool once the weights are evaluated. Models whose tensors load
unchanged, such as gemma4, leave nothing in the pool and are unaffected.
Gemma3n's MobileNetV5 projector silently produces corrupted image
embeddings on the CPU backend - no error, the model just describes the
wrong image (reproduced on llama.cpp b10760; gemma4's encoder is fine on
CPU). Without this guard the existing partial-offload, limited-VRAM, and
OOM-retry fallbacks would pick the CPU projector on exactly the small
GPUs where gemma3n lands.
* MLX: version bump

* mlx: support ModelOpt global scales in MoE models

* address comments

* address comments
…14969)

* create: add server-side MLX imports and drop GGUF conversion

Support safetensors imports through the MLX create pipeline both locally and on the server, including remote upload/staging, draft layer handling, cancellation propagation, transfer limits, and shared manifest/blob writing.

Limit GGUF create to wrapping existing GGUF inputs into Ollama manifests. Remove the in-tree safetensors-to-GGUF converter, server quantization path, and converter-only dependencies so GGUF conversion and quantization stay in llama.cpp tooling.

Keep the MLX path focused on supported safetensors model creation with validation before MLX work, and expose that flow without the --experimental CLI gate.

* address comments

* add client side gguf create fast path

* address comments

* rebase adjustments
Includes quantized matmul corruption fix, which impacts nvfp4 multimodal models (gemma4 vision towers). Deferring wiring up the new MLX-C thread-local stream/sync APIs for now.
Drop support for creating new models with typical_p parameters, while
retaining support for existing GGUF models with the setting.
Upstream v0.34.1 (38fdb5d): llama.cpp b10760 -> b10864, MLX ce916dbb ->
d9add9d1 (which carries our idle-core pin, ollama#4452, and reaches past the ollama#4458
gather_qmm wall through upstream's own MLX-C carry patch in mlx/compat/),
MLX-C c74db530 -> ebc88f10.

Ten conflicts. MLX_VERSION: upstream's. convert/ (deleted upstream, GGUF
conversion is llama.cpp tooling now) and server/model_list_cache.go (deleted;
our Nemotron-text-only and gemma4-safetensors capability rules already live in
images.go's filterUnsupportedCapabilities): deletions accepted.
llm/llama_server.go: upstream's repeat-limit error in our (result, error)
arity. x/mlxrunner/client.go: our admission signature and fields with
upstream's closed flag, CheckRuntime and the locked closed-check before Start;
upstream's integrated-GPU clamp folded into admit(), which now takes
systemInfo. x/mlxrunner/runner.go: model setup lives in upstream's loadModel
now; only configureCacheLimit() kept, after the weights are resident.
x/mlxrunner/pipeline.go: our guardClose teardown on upstream's scoped shape,
logging memory peak/held; our stop-sequence handling ported into upstream's
scoped stream loop, which already records before streaming. The two server
test files: our fixtures on upstream's gguftest types (ggml.KV is gone).

Upstream replaced Pin/Unpin/Sweep/LogArrays with scoped lifetimes. Three files
of ours used them and are ported: the qqmm bench (function scopes per shape,
per input and per timed call; QuantizedMatmul gained a global-scale argument
with the carry patch) and the two vision tests (the worker's stop hook no
longer sweeps). A prefix-cache test fixture gains the PreparedItem upstream's
non-causal boundary rule now reads on every item.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… preflight pins move, task doc

runnerlog.py takes the runner's new per-request `memory peak=… held=…` line as
a request boundary and exposes `held`; summarize_retained_memory.py prints
held and its step between requests on such logs, since the per-array listing
and `active − tracked` no longer exist past v0.34.0. Two tests; README entry.

cuda-dynres-903 pins move to the fold's payload: llama_cpp_build 5d806aa25
(b10864) and mlx_build d9add9d11f31… in full, as the profile tests require.

The task doc records the assessment, the conflict table, the ports and the
decisions; gates are pending the image.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Between b10760 and b10864 upstream changed the gemma4 projector case's
set_limit_image_tokens(40, 280) to (70, 1120) and dropped the comment above
it. Both were context lines of 004's clip.cpp hunk, so the patch stopped
applying and the first image build died in the vulkan stage at 3.5 minutes.
The insertion point had not moved.

Re-cut on a real b10864 checkout: the added and removed lines are identical
to before (+67/-8); only context and hunk offsets differ. The full six-patch
series applies clean from a clean checkout in build order with plain
git apply, which is what apply-git-patches.cmake and CI's patches jobs do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…age-token limits

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…34.0 fold's identity

The header still said the current fold was v0.33.3-dynres and the deployed build
v0.33.2-dynres.1; v0.34.0-dynres has been the fold and the deployed build since
2026-09-14. The "what tested green" matrix is regenerated by release_matrix.py
from that build's full preflight run (21 pass, 7 skip, measured on the container
that serves).

"What differs from upstream, concretely" is rewritten one row per capability
against upstream v0.34.1 at llama.cpp b10864, grouped by what the row is for:
vision correctness on the llama.cpp path, structured output and generation
control, serving and scheduling, and measurement. Two rows changed since it was
written: upstream adopted the fork's gemma4 default limits (70/1120) in b10864,
so only the budget fill is still ours; and upstream added its own integrated-GPU
admission bound, which now sits inside admit(). Both upstream tickets the table
cites are still open. Verified before writing: nemotron's canvas is still fixed
at b10864, the MLX runner upstream has no stop-sequence handling, and its KV
cache type is one global env.

The task doc's directory-level table is replaced by a pointer to this one plus
the size of the divergence, which is what the merge needs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The README row for think + format said upstream constrains the whole
generation. It does not: since v0.34.1 upstream's routes layer defers the
grammar until the thinking→content transition and folds pass-one metrics into
the final response. The fork's flow is a superset — the marker-stop pass one,
the metrics reconstruction when a runner does not report, the second pass
pinned to pass one's truncation window — and the row now says so.

That makes it a retirement candidate, which is what the new register is for:
every item the fork carries, what retires it, and the test that gates its
deletion, reviewed at each fold. x/structured heads it with Glenn's decision:
carried until parity against upstream's engine is tested, deleted if there is
no regression. The README's "fixes carried until upstream takes them" bullet
links the register.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
x/structured (ADR 0009/0013) has had no importers since ADR 0033 adopted
upstream's engine. Glenn's condition for deleting it is a test against that
engine with no regression. This is the test: both engines driven live over a
byte-level vocabulary (one piece per byte plus EOS, so xgrammar's token
matcher and x/structured's byte machine see the same input), every verdict
computed on both sides and diffed, differences classified rather than
transcribed. It skips wherever libollama_xgrammar.so is not present and
accepts OLLAMA_XGRAMMAR_LIBDIR for a host without a payload.

Measured on xgrammar v0.2.5: 108 agreements, 0 regressions. Whitespace policy
differs where llama.cpp's grammar never allowed space before a colon; -0 is
refused by xgrammar's integer grammar and accepted by its number grammar;
prefixItems is upstream-only; and allOf with more than one branch degrades to
a permissive object on xgrammar - required properties, per-branch types and
property order are not enforced, per the converter's own warning - which the
test reports as a known limitation rather than a failure, so the decision
stays a decision. ADR 0013's property holds without its guard: the five
unbounded schemas compile in 3-30 ms within 1 MiB.

The retirement register carries the result and the recommendation. Delete
this file with the package.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
glennneuber and others added 11 commits September 17, 2026 12:27
…ty gate found

x/structured (ADR 0009, 0013) has had no importers since ADR 0033 adopted
upstream's xgrammar engine, and ADR 0033 named its deletion as the follow-up.
Glenn's condition was a test against that engine with no regression. The
parity gate ran both engines live over a byte-level vocabulary on xgrammar
v0.2.5: 108 agreements, 0 regressions — no schema and no output the package
accepted is refused by xgrammar. Deleting it changes no request path; it
removes 3,829 lines of a second grammar engine every fold merged around.

The gate imported the package and goes with it. What it found is kept as
engine_behaviour_test.go, xgrammar-only characterisation tests that pass while
the measured behaviour holds and fail, naming the record to update, when an
xgrammar bump changes it: allOf with more than one branch is permissive on
xgrammar (required, per-branch types and order unenforced — its converter says
allOf support is still ongoing), the integer grammar refuses -0 where the
number grammar takes it, and whitespace before a colon is admitted where
llama.cpp's grammar never allowed it. ADR 0013's property stays as a budget
test: the five schemas the deleted guard refused compile in 3-30 ms within
1 MiB, because xgrammar compiles repetition lazily.

ADR 0033 carries the amendment; the retirement register moves the package to
retired; the task doc records it as decision D6. go mod tidy is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ve carried

070580c deleted x/structured and the parity gate but staged nothing else:
the first pathspec of the git add named the directory that had just been
removed, and the failure stopped the rest from being staged. This is the rest.

engine_behaviour_test.go keeps what the gate found, xgrammar-only: allOf with
more than one branch is permissive, the integer grammar refuses -0 where the
number grammar takes it, whitespace before a colon is admitted, and ADR 0013's
bound holds as a budget. Each pin fails, naming the record to update, when an
xgrammar bump changes the behaviour. ADR 0033 carries the amendment, the
retirement register moves the package to retired, the task doc records D6.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Preflight PASS 21/7 from the fold worktree. GGUF think-off against the
0.34.0 fold: no quality cell regressed, e2b and e4b recovered, one e2b
cell tripped upstream's new token-repeat error, two n=1 movers under
repeat. Qwen2.5-VL: identical cell for cell. MLX: admission refused three
models under 50 GB of foreign GPU0 occupancy, so that leg is invalid and
re-runs alone once the GPU frees. Throughput columns not read: shared GPU.

The gemma4 comment in llama_server.go named llama.cpp's old (40, 280)
defaults; b10864 made them (70, 1120).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
MLX think-off re-run alone once GPU0 had 85 GB free: five models, no
errors, every scored cell equal to the deployed 0.34.0 build, one 35b-a3b
name_bbox cell moved at n=1 (repeats running). The run's `memory` lines
are the first real ones: slog writes msg=memory unquoted, runnerlog.py
accepted only the quoted form and read nothing — widened, with the real
line pinned as a test. held is bounded on gemma4 12b/26b/31b and released
on qwen3.8; on qwen3.6:35b-a3b it grows 0.60 GiB per request for all 28
requests with no eviction, which this log cannot split between trie fill
and a leak; a trace-level probe is queued and the task doc says so.

README: current fold v0.34.1-dynres, deployed stays v0.34.0-dynres, the
matrix regenerated from the candidate's full preflight run
(preflight-runs/full-0341.json on the array; runs/ is gitignored).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`held` alone cannot tell a prefix trie filling towards its cap from a
leak. The 0.34.1 runner still logs the trie's accounting line at trace
level, so the parser keeps its active_size and the retained-memory
summariser prints held − trie and its step: the weights plus whatever
nothing tracks, with the trie's growth removed. One test, 134 pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…X regression

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…dence

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…b removes it

Trace probe: held − trie grows +0.23 GiB per request on 35b-a3b and up
to 1.1 GiB per request on 27b (released later in 1 GiB steps). Control
with OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0: flat for all 28 requests. The
summariser's trie figure is active + paged out.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ar than the deployed build; the knob takes it to flat

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…(upstream ec3cc23)

Upstream v0.34.2's fix, re-rooted under x/: the decode loop released MLX's
pool of freed buffers only when the generated count landed exactly on a
multiple of 256, which a speculative round steps over, so the pool was
never released and the runner's footprint climbed over a generation.

Measured on this fold with the trace suite (qwen3.6:35b-a3b-nvfp4, 28
requests, drafting under a grammar on): resident − trie grew +5.3 GiB
without it and +2.4 GiB with it, against +1.8 GiB on the deployed 0.34.0
build and +0.1 GiB with OLLAMA_MLX_DRAFT_UNDER_GRAMMAR=0. The hunk removes
the fold's memory regression against the deployed build; the drafting
retention that both builds share stays behind the knob (ADR 0033, D5).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…d carried; memory re-measure closed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@glennneuber
glennneuber marked this pull request as ready for review September 18, 2026 00:10
@glennneuber
glennneuber merged commit 8a7ba94 into main Sep 18, 2026
34 of 35 checks passed
glennneuber added a commit that referenced this pull request Sep 18, 2026
Bisected on the encoder rather than the tier: the encoder measurement is
deterministic and answers in seconds, while the tier needs a scored
benchmark on a quantity that drifts hundreds of tokens between runs.

Midpoint 907deff (v0.34.0 fold + #300 + #301, pins b10760 / MLX ce916dbb)
is bit-identical to 0.33.2 on the encoder for both affected models, scores 4
on the think-on 9px tier, and has a greedy eval range of one token. So the
image embeddings, the OCR tier and the run-to-run spread all flip at the
same commit boundary, and all of them are in #302.

That eliminates the whole v0.34.0 upstream fold, #300 and #301. It does not
prove one cause; it rules out the findings being scattered and needing
separate hunts.

Also records a separation route that does NOT work: pairing 8a7ba94's Go
with ce916dbb's MLX payload segfaults at package init, because the new Go
calls mlx_install_capture_handler and the old MLX-C lacks the symbol. The
halves are hard-coupled and cannot be split by swapping payloads.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants