Skip to content

[FEAT] 4 unimplemented optimizations from PR391 plus 2 new ones - six serving optimizations for Qwen3.8 Flash-Next on MTPLX 2.11.2 (four PR-391 remainder optimizations + two exact decode optimizations) - #475

Open
davidtai wants to merge 11 commits into
youssofal:mainfrom
davidtai:perf/qwen38-aux-lanes

Conversation

@davidtai

@davidtai davidtai commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Six serving optimizations for Qwen3.8 Flash-Next: the PR 391 remainder and two exact decode optimizations

Interim update: the context ladders, the 16,384-token interleave, and the charts are final. Quality columns and the typical-threshold sweep fill in as those runs complete; cells that read "measuring" are still pending.

This pull request adds six serving optimizations to the Qwen3.8 Flash-Next path. The base is upstream MTPLX 2.11.2. Four are the PR 391 optimizations that 2.11.2 did not re-land: a hyper-connection read kernel, a prefill causal-mask fuse, a QSA prefill query tile, and a QSA split-K sparse decode kernel. Two are exact decode optimizations: cached async PLE rows and pooled QSA row selection. The four remainder optimizations are rounding-class, so the output differs from the stock path only by floating-point rounding. The two aux optimizations are exact, so their output is byte-for-byte the base. The base commits also re-engage four upstream verify optimizations that 2.11.2 shipped but left silently off when served (Section 3).

The comparison has two arms: release 2.11.2 alone (A) and release plus this pull request's six optimizations (C, this pull request). The headline is C against A.

Headline figures at 16,384 tokens of context, 3 seeds, cold prefill:

  • Decode: 71.97 tok/s on release 2.11.2, 82.84 tok/s with all six optimizations, +15.1%.
  • Prefill: 1381.5 tok/s on release, 1397.9 tok/s with the optimizations, +1.2%.
  • Long context: both release and this pull request exceed the 100 GiB memory knob at 261,120 tokens on this machine; the prefill completes and the overflow is the speculative verify's KV-cache write, which [BUG FIX] OOM - Fit 261,120-token decode: write the verify KV in place and keep the verify attention fused #482 fixes so the cell decodes at a 100.82 GB peak (Section 1.5).
  • Quality at the benchmark's own sampler (temperature 1, xhigh, one seed): HumanEval strict pass@1 0.9695 on release against 0.9634 with the six optimizations, completed-task 1.0000 against 0.9814; MBPP at that sampler ran for one arm only (Section 4).

Terms used in this document:

  • decode: the token-by-token generation phase, after the prompt is read.
  • prefill: reading the prompt into the KV cache before the first output token.
  • optimization: one change an operator can switch off on its own.
  • PLE: a per-layer embedding, a small extra lookup the model adds inside each layer.
  • QSA: the sparse-attention path, which selects the earlier tokens each layer attends to.
  • HC: the hyper-connection read inside the verify step.
  • M4: the fixed four-row speculative verify width.
  • rounding-class: the output differs from the stock path only by floating-point rounding.
  • tok/s: tokens per second.

1. Benchmark results

Two arms ran in one benchmark run on one machine. Each arm received the same request body per cell. Section 2 gives the settings.

1.1 Summary

summary
Summary grid: decode, prefill, time-to-first-token and peak memory by context size, one line per arm, each point the fastest of its seeds. Bars show the slowest-to-fastest spread.

1.2 Decode tok/s

decode_tok_s

Context A release C (this PR) C vs A
1,024 87.32 (73.66-87.32) n=3 88.62 (85.88-88.62) n=3 +1.5%
8,192 72.29 (62.44-72.29) n=3 85.95 (75.93-85.95) n=3 +18.9%
16,384 71.97 n=9 82.84 n=9 +15.1%
32,768 76.47 (69.65-76.47) n=3 81.30 (71.72-81.30) n=3 +6.3%
65,536 69.26 (65.68-69.26) n=3 84.43 (70.82-84.43) n=3 +21.9%
131,072 64.26 (59.70-64.26) n=3 68.62 (65.67-68.62) n=3 +6.8%

The 261,120-token cell is in Section 1.5; both arms exceed the memory knob there.

1.3 Prefill tok/s

prefill_tok_s

The two exact decode optimizations do not change prefill; the candidate column is arm C prefill.

Context A release C (this PR) Delta
1,024 916.1 n=3 912.2 n=3 -0.4%
8,192 1359.3 n=3 1361.0 n=3 +0.1%
16,384 1381.5 n=9 1397.9 n=9 +1.2%
32,768 1212.5 n=3 1260.3 n=3 +3.9%
65,536 1150.7 n=3 1177.2 n=3 +2.3%
131,072 1116.7 n=3 1131.4 n=3 +1.3%

1.4 Time-to-first-token and peak memory

ttft_s
Time-to-first-token by context size.

peak_memory_gb
Peak memory by context size. The 100 GiB cap holds on every arm that reaches the cell.

The per-context tables for time-to-first-token and peak memory are in the perf report under docs/perf/.

1.5 261,120-token context

Both release 2.11.2 and this pull request exceed the 100 GiB memory knob at 261,120 tokens on this machine under the served launch. Metal returns insufficient_memory at the same point on both, at a sampled peak of 100.82 GB on release and 100.79 GB on this pull request.

The prefill is not what overflows. Every arm's receipt shows the full prompt read and tokens already emitted before the failure: new_prefill_tokens 261120, time to first token 243.5 s on release and 241.9 s on this pull request, then completion_tokens 12 on release and 3 on this pull request, with finish_reason error and [METAL] Command buffer execution failed: Insufficient Memory. The overflow is in the first decode step.

That decode step is the fixed four-row speculative verify, and the allocation that overflows is its KV-cache write. The verify holds its KV in TensorOffsetKVCache, whose update_and_fetch writes the new rows with the functional mx.slice_update. That op returns a new array, so it reallocates the whole [1, 2, capacity, 256] bf16 buffer for keys and again for values, about 0.27 GB each at the 262K capacity, in every one of the 12 full-attention layers: about 6.4 GB inside one verify command buffer, on a steady state already near 100.8 GB against a 107.374 GB limit. The stock KVCache used by plain autoregressive decode writes in place and allocates nothing, which is why the same prompt fits with native MTP off. A direct probe of that cache at this geometry measures 6.42 GB resident and, once the write is in place, a 0.00 GB update transient.

A second, smaller term sits in the same step: with a query length of 3 to 6 rows and 12 query heads per key/value head, query length times 12 is past the 32-row limit of MLX's fused vector attention kernel, so MLX declines the fused call and the unfused route materializes a [24, S, 261120] score tensor per layer. Removing only that term does not make the cell fit. A head-chunked build engaged and still failed at the same 100.82 GB peak. The prefill mask-fuse optimization declines in exactly this band, query length 3 to 8, so it does not help either; the wide 256-row prefill chunks still fuse. The analysis trail is in over100-reports/oom255k/report.md.

The earlier #391 tree only looked like it fit. That fit was one cold seed; its other two seeds also ran out of memory at the same cell.

This pull request does not change 261,120-token behaviour; the fix is #482, which writes the verify KV in place and keeps the verify attention on the fused kernel. Two windows confirm the cause and the fix:

Window Configuration Result at 261,120 Peak (GB)
W5 native MTP off (plain autoregressive decode, stock in-place KVCache) fits, 547 tokens, normal stop 94.19
F2 #482, MTP on, verify running fits on all 3 cold seeds, 461 / 1024 / 845 tokens 100.82

With the fix the cell decodes at 62.84 tok/s fastest (56.82-62.84), prefill 1082.9 tok/s, and sits about 6.5 GB under the knob, while 16,384-token output stays byte-identical to release.

1.6 Interleaved 16,384-token protocol

The arms ran interleaved at the 16,384 / 1,024 cell in the order A C D E, three windows per arm (D and E are the #478 typical-acceptance arms). The paired delta compares each arm-C window against its neighbouring arm-A windows, which cancels the slow warm-up drift across the run.

decode_16k_windows

Window Arm Decode tok/s (fastest seed)
a1 A (release) 71.31
c1 C (this PR) 76.94
a2 A (release) 71.86
c2 C (this PR) 77.11
a3 A (release) 71.57
c3 C (this PR) 76.66

Release 71.97 tok/s and candidate 82.84 tok/s are each the fastest window. The fastest-window gap is +15.1%. The paired per-seed 95% interval on the mean delta is [+1.76, +8.89] tok/s. Drift bracket (slowest-to-fastest spread between the release windows) +0.26 tok/s.

1.7 Supplementary measurements

These two sets are supplementary, not headline. Both use the fastest-of-seeds rule.

Greedy code-eval on the release arm, kept for reference only. It is not the quality gate: greedy is temperature 0, which is not the setting anything here was benchmarked at.

Suite (release, greedy) Strict pass@1 Completed-task pass@1 Truncation (rate; mean/max tokens)
HumanEval / HumanEval+ / MBPP 0.9756 / 0.9512 / 0.9040 0.9756 / 0.9512 / 0.9082 0.0% / 0.0% / 0.5%

Three per-optimization knock-outs at 16,384 tokens, on this pull request's tree, each one window of three seeds. The plan was ten leave-one-out windows; the run was cancelled after three, which had already finished and passed their gates. The other seven were not measured, and each of these three is a single window, so its per-seed range is wide: read the change as indicative, not a paired result. This pull request's fastest ABAB decode is 82.84 tok/s.

Optimization off Decode with it off (fastest seed, range) Change vs this PR (82.84 fastest) Note
HC mixer read 79.45 (71.90-79.45) -3.39 tok/s (-4.1%) removed lane read off (armed:false)
QSA sparse decode 82.00 (70.46-82.00) -0.84 tok/s (-1.0%) removed lane dropped from the install report
Prefill mask fuse 81.94 (72.15-81.94) -0.90 tok/s (-1.1%) removed lane logged no engaged line

The remainder-only arm: release plus the four ported optimizations and the four re-engaged keys, without the two exact decode optimizations. The battery calls it arm B. It was stopped after 32K, so it is not a headline. Decode is the fastest of its seeds with the slowest-to-fastest range.

Context Remainder-only decode (fastest, range) vs release (A) vs this PR (C)
1K 87.86 (83.75-87.86) n=3 +0.6% -0.9%
8K 84.71 (75.03-84.71) n=3 +17.2% -1.4%
32K 79.92 (70.29-79.92) n=3 +4.5% -1.7%

2. Benchmark method

Setting Value
Model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed, revision 29ba90f82124961d0d902a9ea9bbb1034972af2f
Arms release MTPLX 2.11.2 (A); release plus this pull request's six optimizations (C, this pull request)
Engine build MTPLX 2.11.2 on MLX 0.32.2
Profile cli-resolved (Turbo), the profile mtplx serve selects for this pack; a fresh SSD session cache directory per window
Prompt sizes 1,024 / 8,192 / 16,384 / 32,768 / 65,536 / 131,072 / 261,120 tokens
Output length 1,024 tokens maximum per request
Sampler temperature 1, top-p 0.95, top-k 20
Reasoning effort xhigh
MTP depth native, depth 3
Seeds 20260829, 20260830, 20260831 (3 seeds per load up to 32,768 tokens, 1 seed per load from 65,536 tokens)
Quality seeds one seed, 20260829, at the same sampler as the speed cells (temperature 1, top-p 0.95, top-k 20, xhigh)
16,384 protocol interleaved A C D E, 3 windows per arm, 3 seeds each; each arm's 16K value is its fastest window; paired per-seed deltas of C against A with a 95% interval on the mean delta
Cell statistic throughput cells report the fastest of the seeds with the slowest-to-fastest range in parentheses; TTFT and wall report the fastest (lowest); peak memory reports the highest
261,120 note both release and this pull request exceed the memory knob at this cell (Section 1.5)
n-gram table pre-warmed before each cell
Memory cap 100 GiB on every arm
Prefill state cold; cross-request prefix restore off, so each seed prefills fresh; new prefill tokens equal prompt tokens; every cell cold by receipt
Thermal state fans at maximum, a 40 degree Celsius gate before every cell

Per-optimization attribution was not measured on this base; the numbers of record are the whole-PR deltas from the ladder (Section 1) and the 16K interleave (Section 1.6). A partial remainder-only arm (release plus the four remainder optimizations, at 1K, 8K and 32K) is kept as supplementary data in the docs appendix under docs/perf/; it is not one of the two arms in these tables.


3. What each optimization does

Optimization Phase Exactness Key
Hyper-connection read at M4 decode rounding MTPLX_QWEN4_HC_M4
QSA split-K sparse decode decode rounding MTPLX_QSA_SPARSE_DECODE
Cached async PLE rows decode exact MTPLX_QWEN4_PLE_CACHED_AUX
Pooled QSA row selection decode exact MTPLX_QSA_POOLED_ROWSEL
Prefill causal-mask fuse prefill rounding MTPLX_QWEN4_PREFILL_MASK_FUSE
QSA prefill query tile prefill rounding MTPLX_QSA_PREFILL_QUERY_TILE

Decode

Hyper-connection read at M4 (MTPLX_QWEN4_HC_M4)

  • Problem. The verify step reads the hyper-connection for a small block of rows as a chain of normalize, down-project and up-project steps.
  • Change. It runs the whole read as one GPU kernel over a compact 8-bit copy of the weights. No native extension.
  • Effect. Rounding-class output; decode gets faster.
  • Exactness. Rounding: differs from stock only by floating-point rounding; covered by the quality gate.
  • Files. mtplx/kernels/qwen4_m4_hyper_read.py, mtplx/models/qwen4_exp.py.
  • Switch. Default on for a served Flash-Next pack. MTPLX_QWEN4_HC_M4=0 turns it off.

QSA split-K sparse decode (MTPLX_QSA_SPARSE_DECODE)

  • Problem. Each verify cycle the sparse attention gathers the selected KV rows into a large tensor per layer, then reads it.
  • Change. A native split-K kernel reads the selected rows once in place and skips the gather and the extra copy. It needs the native extension mtplx_native_qsa; if the extension is not built the optimization declines to stock, prints the reason, and still serves.
  • Effect. Rounding-class output; decode gets faster.
  • Exactness. Rounding: differs from stock only by floating-point rounding; covered by the quality gate. A load-time parity probe on stock MLX gives worst relative L2 3.14e-05 against an fp32 reference with top-1 identical (1.0000), and relative L2 4.60e-03 against the stock-gather path with top-1 0.9896.
  • Files. mtplx/kernels/qsa_sparse_decode.py, native_extensions/qsa_sparse_gqa/ (mtplx_native_qsa).
  • Switch. Default on for a served Flash-Next pack. MTPLX_QSA_SPARSE_DECODE=0 turns it off.

Cached async PLE rows (MTPLX_QWEN4_PLE_CACHED_AUX)

  • Problem. The model builds the extra per-layer embedding lookup inside the timed compute on every decode step, so the step waits for it.
  • Change. A CPU-side native provider computes the fixed 64 row IDs and reads the cold rows, reusing the stock row cache, and the auxiliary embedding plane is built beside the compiled verify step with mx.async_eval so it overlaps the compiled replay instead of sitting on the critical path. It needs the native extension mtplx_native_ple_cpu_rows; if the extension is not built the optimization declines to stock and still serves, still exact.
  • Effect. Byte-identical output; decode gets faster.
  • Exactness. Exact: byte-identical output.
  • Files. mtplx/ple_cached_aux.py, mtplx/ple_cached_row_handoff.py, native_extensions/ple_cpu_rows/ (mtplx_native_ple_cpu_rows).
  • Switch. Default on for a served Flash-Next pack. MTPLX_QWEN4_PLE_CACHED_AUX=0 turns it off.

Pooled QSA row selection (MTPLX_QSA_POOLED_ROWSEL)

  • Problem. The twelve QSA indexers each rebuild the same pool-kernel setup on every decode step, and each holds its own copy of the RoPE rotation table.
  • Change. It binds the existing qsa_indexer_pool_keys_metal kernel metadata once per indexer at construction and shares one inv_freq object across all twelve indexers, so the repeated per-step setup leaves the decode step. Stock MLX, no native extension. The install validates the 48-layer QSA layout, the per-indexer geometry, the RMS-norm epsilon, the RoPE scale and the shared inv_freq identity, and a contract failure fails the model load.
  • Effect. Byte-identical output; decode gets faster.
  • Exactness. Exact: byte-identical output.
  • Files. mtplx/qsa_pooled_rowsel.py.
  • Switch. Default on for a served Flash-Next pack. MTPLX_QSA_POOLED_ROWSEL=0 turns it off.

Prefill

Prefill causal-mask fuse (MTPLX_QWEN4_PREFILL_MASK_FUSE)

  • Problem. During prefill the dense attention builds a full score tensor and applies the causal mask by hand, because at head_dim 256 MLX declines its fused kernel and takes this unfused route.
  • Change. It sends the same attention through MLX's fused SDPA with a causal flag, so it never builds the full score tensor. Stock MLX, no native extension. It declines only for the verify-width dead band (query length 3 to 8), falls back to stock there, and logs the refusal.
  • Effect. Rounding-class output; prefill reads and writes less memory. It does not by itself make 261,120 tokens fit: the binding allocation there is the verify's KV-cache write, not attention, and this optimization declines in the verify's own shape band anyway (Section 1.5).
  • Exactness. Rounding: differs from stock only by floating-point rounding; covered by the quality gate.
  • Files. mtplx/models/qwen4_exp.py.
  • Switch. Default on for a served Flash-Next pack. MTPLX_QWEN4_PREFILL_MASK_FUSE=0 turns it off.

QSA prefill query tile (MTPLX_QSA_PREFILL_QUERY_TILE)

  • Problem. The dense QSA attention in a wide prefill chunk builds a large score block.
  • Change. It splits the query rows into tiles and runs attention one tile at a time, so a wide chunk keeps a narrow chunk's peak at the same total work. Stock MLX, no native extension. The value is rows per tile; the default 2,048 is inert at the production chunk width.
  • Effect. Rounding-class output; prefill keeps a lower memory peak at the same work.
  • Exactness. Rounding: differs from stock only by floating-point rounding; covered by the quality gate.
  • Files. mtplx/qwen4_prefill_chunk.py, mtplx/models/qwen4_exp.py.
  • Switch. Default on for a served Flash-Next pack. MTPLX_QSA_PREFILL_QUERY_TILE=0 turns it off.

Every old MTPLX_FABLE_* name still works as an alias when the new key is unset.

Four upstream optimizations that were silently off

These are not new optimizations. They are upstream 2.11.2's own re-landed verify optimizations, now actually engaged; an audit of the ~31 auto-armed keys found four whose readers were dead as served.

Key Reader What it does in the verify step
MTPLX_QWEN4_DRAFT_K20_PRESCATTER qwen4_draft_k20_prescatter.py:200 (re-cached generation.py:323) pre-scattered K20 draft read on the device draft chain
MTPLX_QWEN4_BLOCK_VERIFY qwen4_block_verify.py:132 (re-cached generation.py:337) block speculative verification in the accept loop
MTPLX_QWEN4_OPDIET runtime_options.py:61 exact-preserving op diet in the compiled fixed-M4 verify graph
MTPLX_QWEN4_VERIFY_GLUE runtime_options.py:159 fused QSA-rope glue inside the fixed-M4 verify body

Cause: mtplx serve imports the generation and runtime modules before it stamps the auto-armed keys, so these four readers froze their default (off) at import and the auto-arm was a silent no-op; the arm A (release) numbers above were measured with the four off, which is the release as a user runs it. Fix: this pull request's base commits change the four readers to resolve the environment at use, change no defaults, and add tests/test_qwen4_remainder_arming.py, which replays the served import order. Arm C carries the fix, so it includes the four re-engaged optimizations.

What is not ported

The fifth PR 391 optimization, verify graph-build overlap (MTPLX_QWEN4_GRAPH_BUILD_OVERLAP), is not in this pull request. Upstream 2.11.2 re-landed the fixed-M4 verify as a single compiled graph. The optimization rode a prefix and suffix split of that graph, and that split does not exist upstream. Porting it is a redesign of the compiled verify, and its whole claim is a submission-timing reorder that cannot be proven on the CPU path.

Feasibility, from the port report: on the PR 391 campaign the optimization targeted 1.934 ms per cycle of host-late GPU idle at 16,384 tokens; at the default one-layer split the binding term is about 0.53 ms per cycle, about +1.4 tok/s (about 1.4%); at a three-to-four-layer split it peaks near 1.7 to 1.8 ms per cycle, about +3.6 to +4.7 tok/s (about 4.5%). The optimization is exact by construction, so it is a speed-only optimization, not a quality risk. Recommendation: defer it to a dedicated split-compilation task with a bit-exact A/B gate. The one-page feasibility is in over100-reports/remainder-port-report.md.


4. Quality

The four remainder optimizations are rounding-class, so this pull request ships on the quality gate, not on bit-identity. The two aux optimizations are exact and do not affect it. The gate is the repo's own code-eval, HumanEval and MBPP, served, paired release against this pull request. Every cell runs at the benchmark's own sampler: temperature 1, top-p 0.95, top-k 20, reasoning effort xhigh, one seed (20260829). That is the setting the speed numbers were measured at, so the quality cells answer the question the speed cells raise.

Each cell reports three numbers per arm: strict pass@1 counts a task cut off at the output cap as a failure; completed-task pass@1 excludes the cut-off tasks; and the truncation rate carries the mean and max completion tokens. With reasoning xhigh at temperature 1 the release cut off about 10% of HumanEval tasks at an 8,192-token cap while still thinking, so the strict number measured verbosity, not correctness; the cap is now 32,768 tokens. Truncation is reported in its own column and never re-run, so a non-zero rate is data rather than a failed cell; the strict and completed-task columns bracket what it costs.

Suite Arm Strict pass@1 Completed-task pass@1 Truncation (rate; mean/max tokens)
HumanEval release (A) 0.9695 1.0000 3.0% (mean 3058 tok)
HumanEval this PR (C) 0.9634 0.9814 1.8% (mean 2734 tok)
MBPP release (A) not run not run not run
MBPP this PR (C) 0.9461 0.9830 3.7% (mean 3367 tok)

Provenance of the candidate rows: this pull request's own tree was not re-run for quality. The candidate cells are the typical-acceptance sweep's exact cell, served from 264e0835 with the typical rule switched off. That tree is this pull request's code plus the typical rule, so with the rule off it executes this pull request's optimizations and nothing else. Release MBPP at this sampler was never run, and the programme was cut after the exact arm, so that cell reads "not run"; the greedy release MBPP figure is in the supplementary row of Section 1.7.

Scoring: prompt + solution + tests, solution = last fenced block defining the entry point (evalplus sanitize). Re-scored offline from saved completions after the live gate was found to drop the prompt's helper definitions (HumanEval/38, /50).

Each optimization carries a =0 opt-out, the kill switch if an optimization moves quality.


5. Changes that did not work

These candidates were measured and rejected on the earlier #391 stack. Their code is not in this pull request. The numbers are from that stack, at the 17,408-token shape.

Candidate Measured effect Reason
Native sidecar sync raw PLE about -8.5% The synchronous read sits on the critical path; the cached optimization supersedes it.
Native sidecar async raw PLE about -1.9% The uncached predecessor of the cached optimization.
GPU-stream PLE transport -1.88% The GPU-stream factory did not beat the CPU queue transport.
Queued sampled-D3 selector -4.49% The extra device-to-host boundary adds draft cost.
MTP depth 4 about -30% The deeper draft costs more than it accepts.
Draft temperature 0.85 about -2.5% The benchmark contract is temperature 1.

The full rejected inventory with per-row numbers is in the perf report under docs/perf/.


6. How to run and how to disable

Build the venv and the native extensions with the setup script, then serve the pack. The server arms all six optimizations by default:

scripts/fable/setup_over100_venv.sh
mtplx serve \
  --model ~/.mtplx/models/Youssofal--Qwen3.8-Flash-Next-MTPLX-Optimized-Speed \
  --model-id mtplx-flash-next-optimized-speed

Read GET /health and confirm the install verdicts name the optimizations. For the QSA sparse decode optimization, confirm the [mtplx] qsa_sparse_decode armed: line; a serve that printed declined to stock ran the stock path.

Turn one optimization off with its key set to 0:

MTPLX_QWEN4_HC_M4=0 mtplx serve --model ...
MTPLX_QSA_SPARSE_DECODE=0 mtplx serve --model ...
MTPLX_QWEN4_PLE_CACHED_AUX=0 mtplx serve --model ...
MTPLX_QSA_POOLED_ROWSEL=0 mtplx serve --model ...
MTPLX_QWEN4_PREFILL_MASK_FUSE=0 mtplx serve --model ...
MTPLX_QSA_PREFILL_QUERY_TILE=0 mtplx serve --model ...

The two native optimizations need their extensions built in the serve environment: mtplx_native_qsa for QSA sparse decode and ple_cpu_rows for cached PLE. The setup script builds both. Without an extension its optimization declines to the stock path and still serves.


7. File map and provenance

Area Files
Runtime mtplx/models/qwen4_exp.py, mtplx/runtime_options.py, mtplx/qwen4_prefill_chunk.py, mtplx/runtime.py, mtplx/ple_cached_aux.py, mtplx/ple_cached_row_handoff.py, mtplx/qsa_pooled_rowsel.py, mtplx/qwen4_aux_lanes.py
Kernels mtplx/kernels/qwen4_m4_hyper_read.py, mtplx/kernels/qsa_sparse_decode.py, mtplx/native/__init__.py
Native extensions native_extensions/qsa_sparse_gqa/ (mtplx_native_qsa), native_extensions/ple_cpu_rows/ (mtplx_native_ple_cpu_rows)
Server mtplx/server/openai.py, mtplx/profiles.py
Wheel signer scripts/bundle_native_runtime_wheel.py
Tests tests/test_qwen4_hc_m4.py, tests/test_qwen4_prefill_mask_fuse.py, tests/test_qsa_sparse_decode.py, tests/test_qsa_sparse_decode_wiring.py, tests/test_qsa_sparse_gqa_native.py, tests/test_pr391_ple_cached_aux_cpu.py, tests/test_pr391_ple_cached_row_handoff_cpu.py, tests/test_pr391_fixed_m4_pool_install_cpu.py, tests/test_qwen4_aux_lanes.py, tests/test_bundle_native_runtime_wheel.py
Documentation the perf report and charts under docs/perf/
  • Base: upstream main 21be78b3 (MTPLX v2.11.2).
  • Branch: perf/qwen38-aux-lanes-main @ 72f58d41 = upstream main, then four remainder commits (hyper-connection read, prefill mask fuse, QSA prefill query tile, QSA split-K sparse decode), then one aux commit (cached async PLE rows, pooled QSA row selection).
  • Optimizations source, read-only: PR 391 head a5e38bb7.

History. On the PR 391 campaign at 16,384 tokens the release default reached 71.17 tok/s, the full PR 391 stack reached 80.92 tok/s, and upstream 2.10.2 reached 57.65 tok/s on the same instrument. The nine keys 2.11.2 re-landed recovered about 60% of the full-stack gain; these six optimizations carry the rest and the two exact decode optimizations. On the earlier #391 stack the two aux optimizations together gave +1.59% decode over the same build with both off, byte-identical, in a same-build pair at 16,384 tokens. Those numbers are superseded by the two-arm battery above, which is measured on 2.11.2.

… 2.11.2

MTPLX 2.11.2 re-landed nine of PR 391's twelve Flash-Next decode keys; this
ports three of the unlanded five onto upstream's re-landed structure, wired
into the fixed-M4 auto-arm block with the MTPLX_QWEN4_*/MTPLX_QSA_* namespace
and a per-key =0 opt-out (the old MTPLX_FABLE_* names kept as aliases).

- HC_M4 (MTPLX_QWEN4_HC_M4): the verify-width (2..8 row) hyper-connection read
  run as one multi-threadgroup GEMV (kernels/qwen4_m4_hyper_read). Reader in
  runtime_options read once at import; GatedResidual gains the geometry
  eligibility check, pack validation and the fused read; install validation
  runs after the M4-stage3 install and reports at /health
  qwen4_install_reports.hc_m4. Rounding-class.
- prefill causal-mask fuse (MTPLX_QWEN4_PREFILL_MASK_FUSE): the dense QSA
  prefill chunk goes through MLX's fused SDPA instead of a materialized score
  tensor MLX's head-dim-256 heuristic declines; a per-shape-class capability
  cache keeps a verify step MLX refuses from disarming a wide chunk.
  Rounding-class (exact visible set).
- QSA prefill query tile (MTPLX_QSA_PREFILL_QUERY_TILE): tiles only the dense
  QSA attention query rows so a wider prefill chunk keeps the narrow chunk's
  attention peak and cost. Value companion, default 2048 (inert at the
  production 2,048 chunk width). Rounding-class (exact visible set).

Each is default-on for a served fixed-M4 Flash-Next pack (server auto-arm
lane_defaults, gated on the fixed-M4 config predicate) with a per-key kill
switch through the existing pop loop, and registered in the boot-time
runtime-env validator. All three are rounding-class, so quality-gated on
HumanEval.

Two of the five remain and are documented in docs/perf/qwen38-391-remainder.md:
the QSA sparse split-K decode (a native kernel whose build and parity probe
need the GPU) and the graph-build overlap (its prefix/suffix split of the
fixed-M4 verify has no substrate on 2.11.2's single-graph verify).

CPU tests (venv mlx 0.32.2, no GPU): tests/test_qwen4_hc_m4.py 53 passed,
tests/test_qwen4_prefill_mask_fuse.py 40 passed.
The fourth of PR 391's five unlanded lanes: MTPLX_QSA_SPARSE_DECODE, the native
split-K sparse-GQA attention for the M=4 fixed verify. It reads the selected KV
rows of the fixed QSA cache once per verify cycle instead of materializing a
gathered [1,2,4,2052,256] K/V pair per layer, which is where the shipped lane's
bytes are. Rounding class: fp32 online softmax over the exact visible set.

- kernels/qsa_sparse_decode.py + native_extensions/qsa_sparse_gqa (package
  mtplx_native_qsa: the split-K Metal kernel, steel headers and a nanobind
  binding); mtplx/native loads it, runtime_options reads MTPLX_QSA_SPARSE_DECODE
  (+_TILE 128:32, +_SPLITS 17); the old MTPLX_FABLE_* names are honoured as
  aliases when the new key is unset.
- graphbank.TensorOffsetQSACache validates the lane ONCE at cache install (a real
  parity probe, outside any mx.compile trace); the twin re-promotion sites and
  the compiled verify_step carry it, and the verify body asserts the lane is in
  the traced graph. models/qwen4_exp routes the fixed-capacity verify width to
  the kernel (QSAIndexer._sparse_decode_route) or declines to stock for a request
  shape it cannot serve.
- Server auto-arm: default ON for the fixed-M4 pack ONLY when the native
  extension is built; a wheel without mtplx_native_qsa declines to stock with a
  logged verdict and still serves. An explicit MTPLX_QSA_SPARSE_DECODE=1 reaches
  the fail-closed install (armed and unbuilt raises). Registered in the boot-time
  runtime-env validator; kill switch through the existing pop loop.
- scripts/bundle_native_runtime_wheel.py signs and packages mtplx_native_qsa
  alongside mtplx_qsa_kernels (Developer ID, hardened runtime, secure timestamp),
  with tests.
- The mask-fuse refusal test now accepts either MLX build's native wording: the
  lane logs a version-independent per-class line and never raises under default
  arming (it falls to the stock dense SDPA).

Load-time parity on stock mlx 0.32.2 with the native kernel built: vs the fp32
reference worst rel_l2 3.1e-05 with the top-1 token identical, vs the stock
gather path rel_l2 4.6e-03 (rounding class), across the 4093 and 2052 probe
cells that stand in for the 16K and 261,120 serving regimes.

CPU tests (venv mlx 0.32.2): tests/test_qsa_sparse_decode.py,
tests/test_qsa_sparse_decode_wiring.py, tests/test_qsa_sparse_gqa_native.py and
tests/test_bundle_native_runtime_wheel.py all green.
Served via `mtplx serve` (cli-resolved Turbo, no lane flags), the QSA split-K
decode lane did not engage: qsa_sparse_decode_enabled() read the environment at
IMPORT and cached the default (False), but the fixed-M4 auto-arm stamps
MTPLX_QSA_SPARSE_DECODE (native-gated) into the environment AFTER
runtime_options is imported, so the cache froze the default before the stamp
landed -- the lane was absent from /health with neither an "armed:" nor a
"declined to stock" line. hc_m4 escaped only because its reader is read on a
path where the module was imported after the stamp.

Resolve the flag lazily on the FIRST read (which is the graphbank cache
install, after the overrides are applied), then cache; the _QSA_SPARSE_DECODE
module global stays (tests force it to a bool) and the native-gated default in
the server auto-arm is unchanged. The env is frozen once serving starts, so a
lazy first read is still a single cached bool on the hot path.

Regression tests, the shape that would have caught this:
- the reader picks up a stamp applied AFTER import (an import-frozen reader
  fails it),
- the fixed-M4 auto-arm block stamps the lane when the native extension is
  built, or prints the declined-to-stock verdict and leaves it unstamped when
  it is not.
Second arming failure (battery, 2026-09-07): served as `mtplx serve` launches
it, hc_m4 was OFF for the same reason the QSA decode lane was in commit 3 --
its reader froze the environment at import (default off) while the fixed-M4
auto-arm stamps the lane keys into the environment AFTER runtime_options is
imported. The earlier claim that hc_m4 read on a post-stamp path did not hold
for the served path.

Resolve every remainder-lane flag at USE (the install / route path, which runs
after the overrides are applied), never at import:
- runtime_options: qwen4_hc_m4_enabled, qsa_sparse_decode_tile and
  qsa_sparse_decode_splits (qsa_sparse_decode_enabled was fixed in commit 3);
  each keeps its module global as a test override (None = read env).
- models/qwen4_exp: _prefill_mask_fuse_enabled drops @lru_cache (its body
  already reads os.environ), so a stamp landing after import is seen.
- qwen4_prefill_chunk.resolve_query_tile_rows already read at use.
The native-gated default and the MTPLX_FABLE_* aliases are unchanged. Upstream's
own MTPLX_QWEN4_OPDIET / MTPLX_QWEN4_VERIFY_GLUE readers are left as-is (not part
of this remainder set).

Test reproducing the served order (tests/test_qwen4_remainder_arming.py):
import mtplx.runtime + mtplx.server.openai FIRST, assert all four readers off,
run _server_runtime_env_overrides for the fixed-M4 pack, apply it to
os.environ, then assert all four arm -- the decode lane armed when the native
extension is built, else an explicit declined-to-stock verdict, never silent
absence. The hc_m4 read-once test is rewritten to assert read-at-use, and the
mask-fuse test drops its now-defunct cache_clear() calls.
@davidtai
davidtai force-pushed the perf/qwen38-aux-lanes branch from 8c53d81 to 72f58d4 Compare September 7, 2026 14:51
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 7, 2026
… 391 remainder port

Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391
remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the
four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K
sparse-GQA decode extension. The two lanes here add to the same fixed-M4
lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and
wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa),
all cleanly additive).
PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the
turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never
merged, but the maintainer independently re-landed the Flash-Next stack under
the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are
present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane
(mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel
(mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is
re-siting the arming and native loading off the absent full_stack_env onto
upstream's own machinery.

Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out):
- ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram
  rows and the auxiliary embedding plane is produced with mx.async_eval outside
  the compiled verifier. The stock owner-side row cache is preserved; declines
  to stock with a printed reason when the ple_cpu_rows extension is not built.
- qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the
  pool kernel metadata once per indexer and shares one inv_freq object.

Rebase changes vs the closed-PR commit:
- mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are
  MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_*
  names kept as aliases (primary wins when both set).
- mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and
  _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an
  alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary.
- mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent
  fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET).
- mtplx/runtime.py: the two installs run after the fixed-M4 verify install,
  logging instead of the removed _print_install_receipt.
- mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so
  normalize_runtime_env_overrides accepts the server-stamped values.
- mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension
  / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not
  reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py).
- scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new
  mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized
  release wheel carries a signed ple_cpu_rows Mach-O.
- scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows.

772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default
off) touches the same _SidecarGather rows but at the disk-layout layer; it is
orthogonal to this runtime-scheduling lane and does not subsume it.

CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files
plus the wheel-bundler test.
@davidtai davidtai changed the title perf(qwen4): two exact decode lanes for Qwen3.8 Flash-Next (cached async PLE, pooled-key rowsel) perf(qwen4): six serving lanes for Qwen3.8 Flash-Next on MTPLX 2.11.2 (four PR-391 remainder lanes + two exact decode lanes) Sep 7, 2026
@davidtai davidtai changed the title perf(qwen4): six serving lanes for Qwen3.8 Flash-Next on MTPLX 2.11.2 (four PR-391 remainder lanes + two exact decode lanes) perf(qwen4): six serving optimizations for Qwen3.8 Flash-Next on MTPLX 2.11.2 (four PR-391 remainder optimizations + two exact decode optimizations) Sep 7, 2026
Arming audit (battery, 2026-09-07): `mtplx serve` imports generation / runtime /
model modules before parse_args stamps the auto-arm env, so any flag whose
reader resolves at module import freezes its default before the stamp lands and
the auto-arm's setdefault is a silent no-op as launched. A column-0 scan of
every module holding a stamped key's reader found exactly four such readers
among the ~31 auto-armed keys; every other stamped key reads the environment at
use or is consumed from config.json at model load. All four are decode-verify
lanes, so the release control (71.17 tok/s at 16K) ran without them -- a
plausible slice of the 71->81 gap.

Resolve all four at use (read the environment each call; the module global stays
a test/force override; the env is frozen once serving starts, so two traces of
one graph still read the same value):
- MTPLX_QWEN4_DRAFT_K20_PRESCATTER: qwen4_draft_k20_prescatter._ENABLED, and
  generation.py's cached _QWEN4_DRAFT_K20_PRESCATTER (removed; the one draft
  consult site calls the reader).
- MTPLX_QWEN4_BLOCK_VERIFY: qwen4_block_verify._ENABLED, and generation.py's
  cached _QWEN4_BLOCK_VERIFY (removed; the accept-loop consult calls the reader).
- MTPLX_QWEN4_OPDIET (+ _ITEMS): runtime_options.
- MTPLX_QWEN4_VERIFY_GLUE (+ _ITEMS): runtime_options (reset hook kept, now
  forcing the globals).
No default value changed; keys that already read at use are untouched. Upstream's
STRICT_CLAIMS and BATCH_PAGED_OFFSETS are also import-frozen but are not
auto-armed (operator sets them pre-launch), so they are left as-is.

Tests: tests/test_qwen4_remainder_arming.py extended to assert all four arm in
the served order (import first, stamp, read) and that the fixed-M4 auto-arm
stamps OPDIET / BLOCK_VERIFY / VERIFY_GLUE and their readers then arm. The
block-verify and draft-k20 source-inspection tests and the opdiet read-once test
are rewritten to assert read-at-use.
The arming audit's lesson: gate on the install verdict, not the env. Three
decode-verify lanes had no per-window observable in
/health qwen4_install_reports -- draft_k20_prescatter, block_verify, opdiet --
so a served window could not confirm they engaged. Add read-only reports (no
behaviour change, no defaults touched):
- draft_k20_prescatter: {armed (read at use), engaged (first-use latch set when
  claim_draft_route installs the route), receipt (the last install receipt)}.
- block_verify: {armed, engaged (latched when a block verifier is built for the
  accept loop)} plus a one-shot "[mtplx] MTPLX_QWEN4_BLOCK_VERIFY armed:" log.
- opdiet: {armed, items (configured selection), applied (first-use latch of the
  items that actually ran at a gated site)}.
Each appears only when ARMED (read at use, gate-able without a request), so an
unarmed lane stays absent (== off) like the other lanes; the engaged/applied
latch rides inside the armed report.

CPU test: tests/test_qwen4_remainder_arming.py asserts the three reports are
absent when off / =0 and present with armed True under a served-order stamp.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 7, 2026
… 391 remainder port

Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391
remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the
four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K
sparse-GQA decode extension. The two lanes here add to the same fixed-M4
lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and
wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa),
all cleanly additive).
PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the
turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never
merged, but the maintainer independently re-landed the Flash-Next stack under
the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are
present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane
(mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel
(mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is
re-siting the arming and native loading off the absent full_stack_env onto
upstream's own machinery.

Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out):
- ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram
  rows and the auxiliary embedding plane is produced with mx.async_eval outside
  the compiled verifier. The stock owner-side row cache is preserved; declines
  to stock with a printed reason when the ple_cpu_rows extension is not built.
- qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the
  pool kernel metadata once per indexer and shares one inv_freq object.

Rebase changes vs the closed-PR commit:
- mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are
  MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_*
  names kept as aliases (primary wins when both set).
- mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and
  _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an
  alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary.
- mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent
  fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET).
- mtplx/runtime.py: the two installs run after the fixed-M4 verify install,
  logging instead of the removed _print_install_receipt.
- mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so
  normalize_runtime_env_overrides accepts the server-stamped values.
- mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension
  / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not
  reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py).
- scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new
  mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized
  release wheel carries a signed ple_cpu_rows Mach-O.
- scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows.

772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default
off) touches the same _SidecarGather rows but at the disk-layout layer; it is
orthogonal to this runtime-scheduling lane and does not subsume it.

CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files
plus the wheel-bundler test.
@davidtai
davidtai force-pushed the perf/qwen38-aux-lanes branch from 72f58d4 to 8f0256d Compare September 7, 2026 17:00
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 7, 2026
… 391 remainder port

Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391
remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the
four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K
sparse-GQA decode extension. The two lanes here add to the same fixed-M4
lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and
wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa),
all cleanly additive).
PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the
turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never
merged, but the maintainer independently re-landed the Flash-Next stack under
the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are
present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane
(mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel
(mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is
re-siting the arming and native loading off the absent full_stack_env onto
upstream's own machinery.

Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out):
- ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram
  rows and the auxiliary embedding plane is produced with mx.async_eval outside
  the compiled verifier. The stock owner-side row cache is preserved; declines
  to stock with a printed reason when the ple_cpu_rows extension is not built.
- qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the
  pool kernel metadata once per indexer and shares one inv_freq object.

Rebase changes vs the closed-PR commit:
- mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are
  MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_*
  names kept as aliases (primary wins when both set).
- mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and
  _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an
  alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary.
- mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent
  fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET).
- mtplx/runtime.py: the two installs run after the fixed-M4 verify install,
  logging instead of the removed _print_install_receipt.
- mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so
  normalize_runtime_env_overrides accepts the server-stamped values.
- mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension
  / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not
  reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py).
- scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new
  mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized
  release wheel carries a signed ple_cpu_rows Mach-O.
- scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows.

772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default
off) touches the same _SidecarGather rows but at the disk-layout layer; it is
orthogonal to this runtime-scheduling lane and does not subsume it.

CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files
plus the wheel-bundler test.
@davidtai
davidtai force-pushed the perf/qwen38-aux-lanes branch from 8f0256d to cb32386 Compare September 7, 2026 18:10
… 391 remainder port

Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391
remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the
four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K
sparse-GQA decode extension. The two lanes here add to the same fixed-M4
lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and
wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa),
all cleanly additive).
PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the
turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never
merged, but the maintainer independently re-landed the Flash-Next stack under
the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are
present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane
(mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel
(mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is
re-siting the arming and native loading off the absent full_stack_env onto
upstream's own machinery.

Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out):
- ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram
  rows and the auxiliary embedding plane is produced with mx.async_eval outside
  the compiled verifier. The stock owner-side row cache is preserved; declines
  to stock with a printed reason when the ple_cpu_rows extension is not built.
- qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the
  pool kernel metadata once per indexer and shares one inv_freq object.

Rebase changes vs the closed-PR commit:
- mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are
  MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_*
  names kept as aliases (primary wins when both set).
- mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and
  _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an
  alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary.
- mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent
  fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET).
- mtplx/runtime.py: the two installs run after the fixed-M4 verify install,
  logging instead of the removed _print_install_receipt.
- mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so
  normalize_runtime_env_overrides accepts the server-stamped values.
- mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension
  / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not
  reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py).
- scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new
  mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized
  release wheel carries a signed ple_cpu_rows Mach-O.
- scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows.

772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default
off) touches the same _SidecarGather rows but at the disk-layout layer; it is
orthogonal to this runtime-scheduling lane and does not subsume it.

CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files
plus the wheel-bundler test.
@davidtai
davidtai force-pushed the perf/qwen38-aux-lanes branch from cb32386 to 27d5ff6 Compare September 7, 2026 18:30
@youssofal

Copy link
Copy Markdown
Owner

David, thank you for this. I read the full diff, ran your test files (562 passed, 23 skipped on CPU), built both native extensions in a scratch directory, and checked the C++ n-gram row hash against the numpy one on 3,690 cases (0 mismatches). Here is where I landed.

Things I want to take:

  • The cached async PLE rows path. I verified it produces the same bytes as the current path at every stage except the GPU op itself (same rows, same dequantize call, same LRU order), so I consider it exact.
  • Pooled QSA row selection, after one change (below). It reuses the existing fused pool kernel, which the current tests already pin as bit-exact against the eager chain.
  • The query-tile plumbing, which is inert at the shipped 2048-token prefill chunk.

Things that need fixing before they can merge:

  1. Block verification. Your change reads MTPLX_QWEN4_BLOCK_VERIFY at use, and the server already stamps it on for every served Flash-Next pack, so merging as-is would switch the accept loop to block verification in production. The law that was in the tree was not exact at depth 3 (I have an enumeration oracle showing the sampled distribution drifts from the target). I have rewritten it on our side so it is provably exact, with the oracle as a test, and I am measuring its speed tonight. Please drop the block-verify part of the last two commits and I will handle that lane separately.
  2. The prescatter re-engagement does not work. claim_draft_route still checks the raw module global, which is now None, so it returns early and the lane never runs while /health reports it as armed. One-line fix: check is_enabled() there. This also means your arm C ran without prescatter.
  3. Setting MTPLX_QWEN4_OPDIET=0 now makes the server fail to boot, because pooled row selection is stamped on unconditionally and its installer requires op-diet. It should only be stamped when op-diet resolves on, or the installer should decline instead of raising.
  4. The native wheels cannot be bundled as written. Neither setup.py declares an exact mlx==0.32.2 requirement, your own bundler rejects a wheel without one, and the sparse-GQA pyproject.toml pins nanobind 2.12.0, which its own CMake guard refuses against mlx 0.32.2. The existing qsa_kernels extension shows the convention. I would also rather see the decode kernel added to qsa_kernels than a second package that duplicates the prefill kernel.

Two smaller notes: the HC-M4 kernel reads bf16 weights, not an 8-bit copy as the description says, and the 4096-token prefill docstring is stale.

On the speed claim: the +15% at 16K is the whole stack including the frozen-key re-engagement, and your own knock-out table attributes roughly 6 to 8 points to the six lanes themselves. I have already landed a general fix for the frozen keys on our side (one refresh after the profile environment is applied), so that part of the gain will ship regardless. For the four rounding-class lanes (HC read, mask fuse, sparse decode, and the query tile when it is active) I want the HumanEval/MBPP cells in the table filled in before they go on by default, and I will run an A/B on this machine with each lane switched individually.

If you can rebase with the four fixes above and split the docs artifacts into their own commit, I can take the exact lanes right away and evaluate the rest with numbers.

@davidtai davidtai changed the title perf(qwen4): six serving optimizations for Qwen3.8 Flash-Next on MTPLX 2.11.2 (four PR-391 remainder optimizations + two exact decode optimizations) [FEAT] 4 unimplemented optimizations from PR391 plus 2 new ones - six serving optimizations for Qwen3.8 Flash-Next on MTPLX 2.11.2 (four PR-391 remainder optimizations + two exact decode optimizations) Sep 8, 2026
@youssofal

Copy link
Copy Markdown
Owner

A short update on the exact lanes. The pooled row selection lane is ported on our side with your authorship, gated so that a daemon with op-diet off still boots on the stock path, with your two test files and our environment and health tests green. On this M5 Max with the Flash-Next Speed pack, an 8.8k-token prompt decoded greedily to 1,024 tokens is byte-identical with the lane on and off, with the same acceptance ladder and 71.8 against 72.5 tok/s, so at that length it changes nothing either way. It needs the 16K and 100K cells before it goes on by default. The cached PLE rows lane waits on the native pins from the review.

davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
Add the cascade context-ladder decode chart (cascade_decode_by_context.svg):
exact pairing and cascade alpha 0.0/0.5/1.0/2.0 across 1K-128K, fastest of seeds
with min-max variance bands; 16,384 merged from the arm-G sweep; 261,120 absent
(every arm OOMs on the youssofal#475 base without youssofal#482). Update the charts manifest.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
cascade_decode_by_context now carries eight arms: the exact pairing, the OPT rule
at alpha 0.0/0.25/0.5/0.75/1.0/2.0, and the TokenV3 rule at alpha 0.95 across
1K-128K (fastest of three seeds, min-max bands). 261,120 is absent because every
arm exceeds the memory knob on the youssofal#475 base without youssofal#482.

Chart and manifest only; nothing under mtplx/.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
Two charts comparing this pull request's operating points against youssofal#475 and
youssofal#478, both re-derived from the receipt json rather than from any table:

  cascade_vs_475_478_16k.svg     decode at 16,384 tokens for release 2.11.2
                                 (exact), youssofal#475 (exact), youssofal#478 (typical 0.09),
                                 OPT alpha 0.25, and TokenV3 alpha 0.75 and
                                 0.95, each bar annotated with its own
                                 HumanEval strict pass@1
  cascade_vs_475_478_ladder.svg  decode against context size, 1,024 to
                                 131,072, for the same four arms that have a
                                 full ladder

Bar height and line point are the fastest seed, every band is min-max, and
acceptance mode is in each label: a tok/s figure is not readable without it.
The alphas of the two rules are different quantities and the captions say so.
No arm carries a 261,120 point, because on the youssofal#475 base without youssofal#482 every
arm on this pack exceeds the memory knob.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
Two charts comparing this pull request against youssofal#475 and youssofal#478 on BOTH model
packs, so the pack and the optimization set are separated rather than
confounded. Both are re-derived from the receipt json:

  bare_vs_475_478_16k.svg     decode at 16,384 tokens, the Optimized-Speed
                              pack (release exact, youssofal#475 exact, youssofal#478 typical
                              0.09) beside the Bare-Speed pack (the same
                              three, then this pull request at the exact
                              law, at typical 0.09, at cascade OPT alpha 0.5
                              and at cascade TokenV3 alpha 0.95)
  bare_vs_475_478_ladder.svg  decode against context size, 1,024 to 261,120,
                              for youssofal#475 and youssofal#478 on both packs and this pull
                              request on Bare-Speed

The two colour families are the two packs, not the two pull requests. Bar
height and line point are the fastest seed, every band is min-max, and the
acceptance mode is in each label: a tok/s figure is not readable without it.
At 261,120 the Bare-Speed arms carry a point and the Optimized-Speed arms do
not, because the latter exceed the memory knob on the youssofal#475 base without youssofal#482.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
… as the reference

David's rulings applied to the youssofal#485 chart set:

RULING 1 ("youssofal#475 is NOT exact"): no chart, legend, reference line, caption or
manifest labels youssofal#475 or release 2.11.2 as "exact". The acceptance-off state is
now named "acceptance mode off" / "cascade off"; the youssofal#475 arm is "youssofal#475 (base)".
  * cascade_decode_vs_alpha: reference lines relabeled "cascade off, youssofal#475 base
    (82.80)" and (see below) "typical 0.2 (99.16)".
  * cascade_decode_by_context: arm-C-caspair legend "exact (no cascade)" ->
    "cascade off (youssofal#475 base)".
  * cascade_vs_475_478_16k / _ladder: "youssofal#475 (base)", "acceptance mode off" /
    "cascade off"; the exact-acceptance disclaimer now names the ordinary
    speculative-decoding acceptance law, not any arm.

RULING 2 (typical 0.2, not 0.09, is the youssofal#478 reference): every youssofal#485-vs-youssofal#478
comparison now references youssofal#478 at typical threshold 0.2 (pooled 16,384 window,
99.16 tok/s fastest of n=9, HumanEval strict pass@1 0.9695), replacing typical
0.09. Applied to cascade_decode_vs_alpha (dotted reference line), the 16K bars
and the context ladder. The §2.4 ABAB (typical 0.09 vs TokenV3 0.95) is a
separate measurement and is unchanged.

Charts re-rendered from receipts (fastest-of-seeds, min-max band); manifest
bytes/sha256 refreshed. Docs only; no receipts touched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
…typical 0.2

David's rulings applied to the youssofal#488 chart set:

RULING 1 ("youssofal#475 is NOT exact"): the acceptance-off state is named
"acceptance mode off" (bars/ladders/legends), the youssofal#475 arm is "youssofal#475 (base)",
and release 2.11.2 is never labeled "exact".
  * bare_16k_bars, bare_decode_16k_windows: arm A/B legends "..., exact" ->
    "..., acceptance mode off".
  * decode_tok_s, prefill_tok_s, ttft_s, peak_memory_gb, summary (context
    ladder): A/B legends "..., exact (FR-Spec dark)" -> "..., acceptance mode
    off (FR-Spec dark)".
  * bare_vs_475_478_16k / _ladder: release/youssofal#475 relabeled; the disclaimer now
    names the ordinary speculative-decoding acceptance law, not any arm.
  * pr391-charts/pr391-typical-acceptance-decode: "exact (lane off)" ->
    "typical off"; footnote "+/-1 task of exact" -> "of typical off".

RULING 2 (typical 0.2 is the youssofal#478 reference): the Optimized-Speed youssofal#478
comparison arm in bare_vs_475_478_16k and _ladder now references typical 0.2
(16,384 fastest 99.16 tok/s, n=9), replacing typical 0.09. The Bare-Speed pack
has no typical-0.2 measurement, so its youssofal#478 arm stays typical 0.09, explicitly
labeled (data gap, not a substitution).

Layout: the longer mode labels needed room. bare_vs_475_478_16k now scales its
width with the bar count and drops the pack group-labels below their bracket
lines clear of the caption; bare_vs_475_478_ladder widens the right gutter so
the legend is not clipped; the summary grid falls back to a two-column legend
when labels are long so nothing is cut off.

Charts re-rendered from receipts (fastest-of-seeds, min-max band). Docs only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
…al off"

David's RULING 1 ("youssofal#475 is NOT exact"): the typical sweep's acceptance-off arm
must not be labeled "exact". It is the ordinary acceptance law with typical
acceptance turned off, so it is now named "typical off".
  * acc_vs_length, acc_vs_speed, acc_vs_length_mbpp, acc_vs_speed_mbpp: series
    label "exact (off)" -> "typical off"; the MBPP single-series title/caption
    "exact arm" -> "typical-off arm".
  * tok_s_vs_passk: the base point "exact (off)" -> "typical off".
  * pr391-charts/pr391-typical-acceptance-decode: legend "exact (lane off)" ->
    "typical off"; footnote "+/-1 task of exact (153/164)" -> "of typical off
    (153/164)".
  * manifest.json alt/caption "exact arm" -> "typical-off arm"; bytes refreshed.

RULING 2 does not change this PR's charts: the typical sweep shows every
threshold (off / 0.09 / 0.2 / 0.4) as its own series, not a youssofal#485-vs-youssofal#478
comparison, so there is no single youssofal#478 reference to switch here.

Charts re-rendered from the sweep receipts. Docs only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
…typical 0.2

David's rulings applied to the youssofal#488 chart set:

RULING 1 ("youssofal#475 is NOT exact"): the acceptance-off state is named
"acceptance mode off" (bars/ladders/legends), the youssofal#475 arm is "youssofal#475 (base)",
and release 2.11.2 is never labeled "exact".
  * bare_16k_bars, bare_decode_16k_windows: arm A/B legends "..., exact" ->
    "..., acceptance mode off".
  * decode_tok_s, prefill_tok_s, ttft_s, peak_memory_gb, summary (context
    ladder): A/B legends "..., exact (FR-Spec dark)" -> "..., acceptance mode
    off (FR-Spec dark)".
  * bare_vs_475_478_16k / _ladder: release/youssofal#475 relabeled; the disclaimer now
    names the ordinary speculative-decoding acceptance law, not any arm.
  * pr391-charts/pr391-typical-acceptance-decode: "exact (lane off)" ->
    "typical off"; footnote "+/-1 task of exact" -> "of typical off".

RULING 2 (typical 0.2 is the youssofal#478 reference): the Optimized-Speed youssofal#478
comparison arm in bare_vs_475_478_16k and _ladder now references typical 0.2
(16,384 fastest 99.16 tok/s, n=9), replacing typical 0.09. The Bare-Speed pack
has no typical-0.2 measurement, so its youssofal#478 arm stays typical 0.09, explicitly
labeled (data gap, not a substitution).

Layout: the longer mode labels needed room. bare_vs_475_478_16k now scales its
width with the bar count and drops the pack group-labels below their bracket
lines clear of the caption; bare_vs_475_478_ladder widens the right gutter so
the legend is not clipped; the summary grid falls back to a two-column legend
when labels are long so nothing is cut off.

Charts re-rendered from receipts (fastest-of-seeds, min-max band). Docs only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
… as the reference

David's rulings applied to the youssofal#485 chart set:

RULING 1 ("youssofal#475 is NOT exact"): no chart, legend, reference line, caption or
manifest labels youssofal#475 or release 2.11.2 as "exact". The acceptance-off state is
now named "acceptance mode off" / "cascade off"; the youssofal#475 arm is "youssofal#475 (base)".
  * cascade_decode_vs_alpha: reference lines relabeled "cascade off, youssofal#475 base
    (82.80)" and (see below) "typical 0.2 (99.16)".
  * cascade_decode_by_context: arm-C-caspair legend "exact (no cascade)" ->
    "cascade off (youssofal#475 base)".
  * cascade_vs_475_478_16k / _ladder: "youssofal#475 (base)", "acceptance mode off" /
    "cascade off"; the exact-acceptance disclaimer now names the ordinary
    speculative-decoding acceptance law, not any arm.

RULING 2 (typical 0.2, not 0.09, is the youssofal#478 reference): every youssofal#485-vs-youssofal#478
comparison now references youssofal#478 at typical threshold 0.2 (pooled 16,384 window,
99.16 tok/s fastest of n=9, HumanEval strict pass@1 0.9695), replacing typical
0.09. Applied to cascade_decode_vs_alpha (dotted reference line), the 16K bars
and the context ladder. The §2.4 ABAB (typical 0.09 vs TokenV3 0.95) is a
separate measurement and is unchanged.

Charts re-rendered from receipts (fastest-of-seeds, min-max band); manifest
bytes/sha256 refreshed. Docs only; no receipts touched.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
…al off"

David's RULING 1 ("youssofal#475 is NOT exact"): the typical sweep's acceptance-off arm
must not be labeled "exact". It is the ordinary acceptance law with typical
acceptance turned off, so it is now named "typical off".
  * acc_vs_length, acc_vs_speed, acc_vs_length_mbpp, acc_vs_speed_mbpp: series
    label "exact (off)" -> "typical off"; the MBPP single-series title/caption
    "exact arm" -> "typical-off arm".
  * tok_s_vs_passk: the base point "exact (off)" -> "typical off".
  * pr391-charts/pr391-typical-acceptance-decode: legend "exact (lane off)" ->
    "typical off"; footnote "+/-1 task of exact (153/164)" -> "of typical off
    (153/164)".
  * manifest.json alt/caption "exact arm" -> "typical-off arm"; bytes refreshed.

RULING 2 does not change this PR's charts: the typical sweep shows every
threshold (off / 0.09 / 0.2 / 0.4) as its own series, not a youssofal#485-vs-youssofal#478
comparison, so there is no single youssofal#478 reference to switch here.

Charts re-rendered from the sweep receipts. Docs only.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
…typical 0.2

David's rulings applied to the youssofal#488 chart set:

RULING 1 ("youssofal#475 is NOT exact"): the acceptance-off state is named
"acceptance mode off" (bars/ladders/legends), the youssofal#475 arm is "youssofal#475 (base)",
and release 2.11.2 is never labeled "exact".
  * bare_16k_bars, bare_decode_16k_windows: arm A/B legends "..., exact" ->
    "..., acceptance mode off".
  * decode_tok_s, prefill_tok_s, ttft_s, peak_memory_gb, summary (context
    ladder): A/B legends "..., exact (FR-Spec dark)" -> "..., acceptance mode
    off (FR-Spec dark)".
  * bare_vs_475_478_16k / _ladder: release/youssofal#475 relabeled; the disclaimer now
    names the ordinary speculative-decoding acceptance law, not any arm.
  * pr391-charts/pr391-typical-acceptance-decode: "exact (lane off)" ->
    "typical off"; footnote "+/-1 task of exact" -> "of typical off".

RULING 2 (typical 0.2 is the youssofal#478 reference): the Optimized-Speed youssofal#478
comparison arm in bare_vs_475_478_16k and _ladder now references typical 0.2
(16,384 fastest 99.16 tok/s, n=9), replacing typical 0.09. The Bare-Speed pack
has no typical-0.2 measurement, so its youssofal#478 arm stays typical 0.09, explicitly
labeled (data gap, not a substitution).

Layout: the longer mode labels needed room. bare_vs_475_478_16k now scales its
width with the bar count and drops the pack group-labels below their bracket
lines clear of the caption; bare_vs_475_478_ladder widens the right gutter so
the legend is not clipped; the summary grid falls back to a two-column legend
when labels are long so nothing is cut off.

Charts re-rendered from receipts (fastest-of-seeds, min-max band). Docs only.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
…n from the README

The README carried the same vocabulary defects the pull request body did, and
was fixed the same way, against pr-bodies/GLOSSARY.md.

"Exact" describes the acceptance LAW, never an arm. This pack's arms A and B
are release 2.11.2 and youssofal#475, which are rounding-class builds that happen to run
exact acceptance, so calling them "(A, exact)" reads as a claim about their
kernels. Every arm label now states which lossy mode is off: the five four-arm
tables, the all-kernels table, the HumanEval rows, and the two prose sites that
said "youssofal#475 exact code" and "16K exact".

Section 6 claimed the four-arm sweep was still being measured while Sections
2.1 to 2.5 published it in full, and listed the queue order and the GPU-lock
protocol to get there. It now names only the three real gaps. The Section 2
banner made the same stale claim and is gone, the pack footnote no longer
names an internal role, and the method row states what the GPU lock guarantees
rather than how it is taken.

Checked with pr-bodies/check_glossary.py, which now gates files by path.
…senses of "exact"

Against pr-bodies/GLOSSARY.md, the same pass the pull request bodies had.

The arm table gave an acceptance mode for D and E only, so A and C read as
though they had none. Both run with typical off, and their rows now say so.

"Exact" appears once here, in the optimization-class sense: an exact
optimization is byte-for-byte the stock path, against a rounding-class one that
differs only by floating-point rounding. Nothing in the file bound that, so it
was indistinguishable from the exact acceptance LAW that youssofal#478 and youssofal#485 use the
same word for. The paragraph now binds it and states that it never labels an
arm, since arms A and C are rounding-class builds overall.

The prose mixed short context labels with exact counts; prose now gives the
counts and the short form is left to table row labels.
…verflow

The long-context verdict table blamed a QSA-indexer prefill transient for the
261,120-token OOM on all four arms. The prefill completes on every arm; the
overflow is the first decode step's speculative-verify KV-cache write, where
TensorOffsetKVCache.update_and_fetch used the functional mx.slice_update and
reallocated the full per-layer KV buffers (about 6.4 GB in one command
buffer). The probes W1 and W2 targeted prefill and so could not have fit. This
matches the youssofal#475 body's Section 1.5 and the youssofal#482 fix arm, which fits 261,120
on all three cold seeds at 100.82 GB.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 9, 2026
cascade_humaneval_by_alpha.svg: strict and completed-task pass@1 as grouped
bars for cascade off (youssofal#475 base), the OPT rule at alpha 0.0/0.25/0.5/0.75 and
the TokenV3 rule at 0.75/0.95, with the cascade-off strict level as a
reference line, so the quality cost of the OPT rule is visible next to
TokenV3. Data: evalsweep478/armG_summary.json and the youssofal#478 sweep's typical-off
cell; nothing re-measured.
davidtai added a commit to davidtai/MTPLX-STREAMING that referenced this pull request Sep 10, 2026
youssofal added a commit that referenced this pull request Sep 17, 2026
Four hot-path gates are read once at import so decode never touches
os.environ and two traces of one compiled graph cannot disagree:
MTPLX_QWEN4_OPDIET, MTPLX_QWEN4_VERIFY_GLUE, MTPLX_QWEN4_DRAFT_K20_PRESCATTER
and MTPLX_QWEN4_BLOCK_VERIFY (generation.py keeps copies of the last two).
`mtplx serve` imports those modules before ServerState stamps the model
family's runtime env, and the Flash-Next lane defaults arm all four, so every
served daemon since 2.11.1 reported the keys as configured on /health while
running with all of them off: the PR #391 ports measured on 2026-09-03 never
reached a user, and the K20 prescatter and op diet receipts were harness-only.
davidtai's PR #475 found the same four frozen readers.

apply_profile_env now re-reads the gates in their owning modules right after
it writes os.environ (runtime_options.refresh_env_flags, which also re-copies
generation's constants) and returns the live values as a receipt. Every call
site of apply_profile_env runs before a model load, so nothing has been
compiled when the values change and the traces-agree invariant holds. A test
mapping passed as `environ` leaves module state alone. Block verification is
re-engaged only now that its law is exact (previous commit).
@youssofal

Copy link
Copy Markdown
Owner

2.11.3 ships the fix for the frozen setting readers this PR found, with credit in the release notes. The pooled row-selection kernel measured a tie on the 9k-token prompt (79.4 to 79.6 tok/s with it, 79.6 to 79.7 without), so it is not in this release and the PR stays open. Release: https://github.com/youssofal/MTPLX/releases/tag/v2.11.3

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants