Skip to content

task: Gate 5 completed on CUDA — think-on half and the same-platform 0.33.2 baseline - #278

Merged
glennneuber merged 1 commit into
mainfrom
docs/gate5-same-platform-results
Sep 6, 2026
Merged

glennneuber merged 1 commit into
mainfrom
docs/gate5-same-platform-results

Conversation

@glennneuber

Copy link
Copy Markdown

Docs only: the fold task doc gets criterion 7b with the two halves of Gate 5 that were still owed on the CUDA host, plus a correction to the serving notes.

  • Think-on (mlx0333cu_1_*_thinkon vs sync15nt_1_): 26b 25/27 = 25/27, 31b 27/27 = 27/27, qwen3.8 27/27 = 27/27, qwen3.6 24/27 = 24/27 with two of the three standing arms identical. No regression.
  • Same-platform 0.33.2 baseline (mlx0332cu_1_*_thinkfalse, run today on the CUDA host with OLLAMA_GPU_OVERHEAD 16 GiB, 0 OOMs): quality at parity with the 0.33.3 cells (qwen3.8 27/27 byte-identical; every arm 0.33.3 lacks is one of the four OOM'd arms), and per-request peak memory within ~1 GiB at every quantile like for like, 0.33.2 the higher — so 0.33.3 does not allocate more for the same work and the Sep-4 OOMs were zero-headroom contention. Criterion 7's comparator was the sync15 (0.32.14) build; the 0.33.2 vision campaigns were Metal-host measurements and never mix with CUDA ones.
  • What the think-on run found: at a hard 40 GiB OLLAMA_MLX_MEMORY_LIMIT the runner stalled rather than OOMing (thrash check off by default, mlx: run with MLX's graph-cache thrashing check off; keep the first panic #212); one stall plus three queued timeouts cost 2 h. Serving notes in 7 and 9 now say headroom via OLLAMA_GPU_OVERHEAD, not a hard cap. Companion changes: mlx: price the context rung at admission, headroom calibrated on GPU (#275) #276 (admission prices the rung; GPU calibration pending), vsuite: evict on the way out; headroom, not a cap, on the shared host #277 (driver evicts on exit; README advice corrected).

🤖 Generated with Claude Code

…0.33.2 baseline

Criterion 7 compared the 0.33.3 cells against the sync15 (0.32.14) CUDA build
because no 0.33.2 vision campaign had ever run on this host; the 0.33.2 numbers
were Metal measurements and never mix with CUDA ones. 7b records the missing
half: the think-on campaign (every cell equal to its baseline count), the 0.33.2
think-off baseline run today on the same host (quality at parity, 0 OOMs under
16 GiB of headroom), and the per-request peak-memory comparison that answers
the OOM question -- like for like the builds are within ~1 GiB at every
quantile, 0.33.2 the higher, so 0.33.3 does not allocate more and the Sep-4
OOMs were zero-headroom contention.

Also records what the think-on run found: at a hard 40 GiB cap the MLX runner
stalls rather than OOMs (thrash check off by default), two hours lost to one
stall and the three requests queued behind it; the serving notes in 7 and 9 now
point at OLLAMA_GPU_OVERHEAD (headroom the runner subtracts from free VRAM)
instead of a hard OLLAMA_MLX_MEMORY_LIMIT.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@glennneuber
glennneuber merged commit 5777f3f into main Sep 6, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant