Repository navigation
preflight: pin the payload on native hosts, not just containerised ones - #363
Conversation
check_payload_pin returned SKIP whenever there was no container, and the
native Metal path has none. That is why not one of the seven metal/mlx-metal
profiles carries a llama_cpp_build: the field would have asserted nothing, so
nobody added it.
Found promoting 0.34.2 on Metal. Its LLAMA_CPP_VERSION moved b10864 -> b10969
-- a Metal fusion subsystem that did not exist before and is default-on, plus
a restructured tensor-path matmul inside kernel_mul_mm_id -- while every Metal
expectation was inherited across the move untouched. Nothing in the harness
said so; the move was found by diffing the file by hand. That is the same
accident payload_pin was written for (b10091/b10353 under one version string),
on the platform where it could not run.
The resolution already existed for check_metal_tensor_payload:
local_listener_exe() -> lib_ollama_llama_server(), which is ollama's own exeDir
rule. This binds it to payload_pin too, and to the run artefact's
meta.llama_cpp_build, so a native run names its payload like a containerised
one does.
Three specifics worth keeping:
- remote hosts still SKIP. `llama-server --version` runs on the HARNESS
host, so a remote server's path either does not exist here or belongs to
something else; naming the wrong binary is worse than declining.
- a resolved path that does not exist SKIPs rather than failing. ollama takes
the first lib/ollama DIRECTORY that exists and looks no further, so the
binary may genuinely be absent; that is cannot-answer, not wrong-payload.
- llama_cpp_build with neither exec_cmd nor container used to build
["docker", "exec", None, ...] and die on the None in argv. It now runs the
payload directly, with no shell, so a path with spaces or braces needs no
quoting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Independent corroboration from the ROCm side, for the "payload move is not cosmetic" argument. I spent today measuring ROCm 10.0.0 on gfx1151 (#362) and ran into the same
I eliminated the fork's own causes one at a time: not direct-I/O (on and off both give 22/5), not Two things that may be useful here:
No conflict with #362, which touches only |
What
check_payload_pinreturned SKIP whenever there was no container, and the nativeMetal path has none. This adds the second route — the payload the listening
executable resolves to, by ollama's own exeDir rule — so the pin works on native
hosts as well as containerised ones.
Why this was found, and why it matters
Not one of the seven
metal/mlx-metalprofiles carries allama_cpp_build. Thatis not an oversight: the field would have asserted nothing, because the only check
that consumes it could not run there.
The cost showed up promoting 0.34.2 on Metal.
LLAMA_CPP_VERSIONmovedb10864 → b10969 while
MLX_VERSIONstayed put, and every Metal expectation wasinherited across that move untouched. The payload move is not cosmetic —
5d806aa25..391fac164is +1214/−339 across 17 Metal backend files, including afusion subsystem that did not exist before and is default-on, and a restructured
tensor-path matmul inside
kernel_mul_mm_id.Nothing in the harness said so. I found it by diffing the file by hand.
That is precisely the accident
payload_pinwas written for — b10091/b10353 passingunder one version string — occurring on the platform where the check could not run.
How
The resolution already existed for
check_metal_tensor_payload:local_listener_exe()→lib_ollama_llama_server(). This binds it topayload_pintoo, and to the run artefact's
meta.llama_cpp_build, so a native run names itspayload the way a containerised one does.
Three deliberate behaviours, each with a test:
llama-server --versionruns on the harnesshost, so a remote server's path either does not exist here or belongs to
something else. Naming the wrong binary is worse than declining to answer — the
same rule
check_metal_tensor_payloadalready follows.first
lib/ollamadirectory that exists and looks no further, so the binary maygenuinely be absent. That is cannot-answer, not wrong-payload.
llama_cpp_build(None, path=...)now works. With neitherexec_cmdnorcontainerit used to build["docker", "exec", None, ...]and die on theNonein argv — the same shape as the
container_logsbug fixed in preflight: gate the M5 Neural Accelerators in all three places they can be lost #341. It now runsthe payload directly with no shell, so a path containing spaces or braces needs
no quoting and cannot be re-interpreted.
Verification
python3 test_verdicts.py— 186 tests, OK (skipped=6); 9 of them new.Live on this host, against the 0.34.2 candidate on
:11437with no container:which is b10969, matching
llama-server --versionread directly from the install.Not in this PR
llama_cpp_buildto the Metal profiles. That is a profile change andbelongs with the
mlx-metal-0-34-2seeding, where the value can be measuredrather than asserted.
see a GGUF payload move at all. That is on Metal: what promoting 0.34.2 costs, and why the MLX half is probably inert #361.
Measured on
macbook-pro-m5-max-128GB/mlx-metal.🤖 Generated with Claude Code