Repository navigation
task: Gate 5 completed on CUDA — think-on half and the same-platform 0.33.2 baseline - #278
Merged
Merged
Conversation
…0.33.2 baseline Criterion 7 compared the 0.33.3 cells against the sync15 (0.32.14) CUDA build because no 0.33.2 vision campaign had ever run on this host; the 0.33.2 numbers were Metal measurements and never mix with CUDA ones. 7b records the missing half: the think-on campaign (every cell equal to its baseline count), the 0.33.2 think-off baseline run today on the same host (quality at parity, 0 OOMs under 16 GiB of headroom), and the per-request peak-memory comparison that answers the OOM question -- like for like the builds are within ~1 GiB at every quantile, 0.33.2 the higher, so 0.33.3 does not allocate more and the Sep-4 OOMs were zero-headroom contention. Also records what the think-on run found: at a hard 40 GiB cap the MLX runner stalls rather than OOMs (thrash check off by default), two hours lost to one stall and the three requests queued behind it; the serving notes in 7 and 9 now point at OLLAMA_GPU_OVERHEAD (headroom the runner subtracts from free VRAM) instead of a hard OLLAMA_MLX_MEMORY_LIMIT. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Docs only: the fold task doc gets criterion 7b with the two halves of Gate 5 that were still owed on the CUDA host, plus a correction to the serving notes.
mlx0333cu_1_*_thinkonvssync15nt_1_): 26b 25/27 = 25/27, 31b 27/27 = 27/27, qwen3.8 27/27 = 27/27, qwen3.6 24/27 = 24/27 with two of the three standing arms identical. No regression.mlx0332cu_1_*_thinkfalse, run today on the CUDA host withOLLAMA_GPU_OVERHEAD16 GiB, 0 OOMs): quality at parity with the 0.33.3 cells (qwen3.8 27/27 byte-identical; every arm 0.33.3 lacks is one of the four OOM'd arms), and per-request peak memory within ~1 GiB at every quantile like for like, 0.33.2 the higher — so 0.33.3 does not allocate more for the same work and the Sep-4 OOMs were zero-headroom contention. Criterion 7's comparator was the sync15 (0.32.14) build; the 0.33.2 vision campaigns were Metal-host measurements and never mix with CUDA ones.OLLAMA_MLX_MEMORY_LIMITthe runner stalled rather than OOMing (thrash check off by default, mlx: run with MLX's graph-cache thrashing check off; keep the first panic #212); one stall plus three queued timeouts cost 2 h. Serving notes in 7 and 9 now say headroom viaOLLAMA_GPU_OVERHEAD, not a hard cap. Companion changes: mlx: price the context rung at admission, headroom calibrated on GPU (#275) #276 (admission prices the rung; GPU calibration pending), vsuite: evict on the way out; headroom, not a cap, on the shared host #277 (driver evicts on exit; README advice corrected).🤖 Generated with Claude Code