[FEAT] 4 unimplemented optimizations from PR391 plus 2 new ones - six serving optimizations for Qwen3.8 Flash-Next on MTPLX 2.11.2 (four PR-391 remainder optimizations + two exact decode optimizations) - #475
Conversation
… 2.11.2 MTPLX 2.11.2 re-landed nine of PR 391's twelve Flash-Next decode keys; this ports three of the unlanded five onto upstream's re-landed structure, wired into the fixed-M4 auto-arm block with the MTPLX_QWEN4_*/MTPLX_QSA_* namespace and a per-key =0 opt-out (the old MTPLX_FABLE_* names kept as aliases). - HC_M4 (MTPLX_QWEN4_HC_M4): the verify-width (2..8 row) hyper-connection read run as one multi-threadgroup GEMV (kernels/qwen4_m4_hyper_read). Reader in runtime_options read once at import; GatedResidual gains the geometry eligibility check, pack validation and the fused read; install validation runs after the M4-stage3 install and reports at /health qwen4_install_reports.hc_m4. Rounding-class. - prefill causal-mask fuse (MTPLX_QWEN4_PREFILL_MASK_FUSE): the dense QSA prefill chunk goes through MLX's fused SDPA instead of a materialized score tensor MLX's head-dim-256 heuristic declines; a per-shape-class capability cache keeps a verify step MLX refuses from disarming a wide chunk. Rounding-class (exact visible set). - QSA prefill query tile (MTPLX_QSA_PREFILL_QUERY_TILE): tiles only the dense QSA attention query rows so a wider prefill chunk keeps the narrow chunk's attention peak and cost. Value companion, default 2048 (inert at the production 2,048 chunk width). Rounding-class (exact visible set). Each is default-on for a served fixed-M4 Flash-Next pack (server auto-arm lane_defaults, gated on the fixed-M4 config predicate) with a per-key kill switch through the existing pop loop, and registered in the boot-time runtime-env validator. All three are rounding-class, so quality-gated on HumanEval. Two of the five remain and are documented in docs/perf/qwen38-391-remainder.md: the QSA sparse split-K decode (a native kernel whose build and parity probe need the GPU) and the graph-build overlap (its prefix/suffix split of the fixed-M4 verify has no substrate on 2.11.2's single-graph verify). CPU tests (venv mlx 0.32.2, no GPU): tests/test_qwen4_hc_m4.py 53 passed, tests/test_qwen4_prefill_mask_fuse.py 40 passed.
The fourth of PR 391's five unlanded lanes: MTPLX_QSA_SPARSE_DECODE, the native split-K sparse-GQA attention for the M=4 fixed verify. It reads the selected KV rows of the fixed QSA cache once per verify cycle instead of materializing a gathered [1,2,4,2052,256] K/V pair per layer, which is where the shipped lane's bytes are. Rounding class: fp32 online softmax over the exact visible set. - kernels/qsa_sparse_decode.py + native_extensions/qsa_sparse_gqa (package mtplx_native_qsa: the split-K Metal kernel, steel headers and a nanobind binding); mtplx/native loads it, runtime_options reads MTPLX_QSA_SPARSE_DECODE (+_TILE 128:32, +_SPLITS 17); the old MTPLX_FABLE_* names are honoured as aliases when the new key is unset. - graphbank.TensorOffsetQSACache validates the lane ONCE at cache install (a real parity probe, outside any mx.compile trace); the twin re-promotion sites and the compiled verify_step carry it, and the verify body asserts the lane is in the traced graph. models/qwen4_exp routes the fixed-capacity verify width to the kernel (QSAIndexer._sparse_decode_route) or declines to stock for a request shape it cannot serve. - Server auto-arm: default ON for the fixed-M4 pack ONLY when the native extension is built; a wheel without mtplx_native_qsa declines to stock with a logged verdict and still serves. An explicit MTPLX_QSA_SPARSE_DECODE=1 reaches the fail-closed install (armed and unbuilt raises). Registered in the boot-time runtime-env validator; kill switch through the existing pop loop. - scripts/bundle_native_runtime_wheel.py signs and packages mtplx_native_qsa alongside mtplx_qsa_kernels (Developer ID, hardened runtime, secure timestamp), with tests. - The mask-fuse refusal test now accepts either MLX build's native wording: the lane logs a version-independent per-class line and never raises under default arming (it falls to the stock dense SDPA). Load-time parity on stock mlx 0.32.2 with the native kernel built: vs the fp32 reference worst rel_l2 3.1e-05 with the top-1 token identical, vs the stock gather path rel_l2 4.6e-03 (rounding class), across the 4093 and 2052 probe cells that stand in for the 16K and 261,120 serving regimes. CPU tests (venv mlx 0.32.2): tests/test_qsa_sparse_decode.py, tests/test_qsa_sparse_decode_wiring.py, tests/test_qsa_sparse_gqa_native.py and tests/test_bundle_native_runtime_wheel.py all green.
Served via `mtplx serve` (cli-resolved Turbo, no lane flags), the QSA split-K decode lane did not engage: qsa_sparse_decode_enabled() read the environment at IMPORT and cached the default (False), but the fixed-M4 auto-arm stamps MTPLX_QSA_SPARSE_DECODE (native-gated) into the environment AFTER runtime_options is imported, so the cache froze the default before the stamp landed -- the lane was absent from /health with neither an "armed:" nor a "declined to stock" line. hc_m4 escaped only because its reader is read on a path where the module was imported after the stamp. Resolve the flag lazily on the FIRST read (which is the graphbank cache install, after the overrides are applied), then cache; the _QSA_SPARSE_DECODE module global stays (tests force it to a bool) and the native-gated default in the server auto-arm is unchanged. The env is frozen once serving starts, so a lazy first read is still a single cached bool on the hot path. Regression tests, the shape that would have caught this: - the reader picks up a stamp applied AFTER import (an import-frozen reader fails it), - the fixed-M4 auto-arm block stamps the lane when the native extension is built, or prints the declined-to-stock verdict and leaves it unstamped when it is not.
Second arming failure (battery, 2026-09-07): served as `mtplx serve` launches it, hc_m4 was OFF for the same reason the QSA decode lane was in commit 3 -- its reader froze the environment at import (default off) while the fixed-M4 auto-arm stamps the lane keys into the environment AFTER runtime_options is imported. The earlier claim that hc_m4 read on a post-stamp path did not hold for the served path. Resolve every remainder-lane flag at USE (the install / route path, which runs after the overrides are applied), never at import: - runtime_options: qwen4_hc_m4_enabled, qsa_sparse_decode_tile and qsa_sparse_decode_splits (qsa_sparse_decode_enabled was fixed in commit 3); each keeps its module global as a test override (None = read env). - models/qwen4_exp: _prefill_mask_fuse_enabled drops @lru_cache (its body already reads os.environ), so a stamp landing after import is seen. - qwen4_prefill_chunk.resolve_query_tile_rows already read at use. The native-gated default and the MTPLX_FABLE_* aliases are unchanged. Upstream's own MTPLX_QWEN4_OPDIET / MTPLX_QWEN4_VERIFY_GLUE readers are left as-is (not part of this remainder set). Test reproducing the served order (tests/test_qwen4_remainder_arming.py): import mtplx.runtime + mtplx.server.openai FIRST, assert all four readers off, run _server_runtime_env_overrides for the fixed-M4 pack, apply it to os.environ, then assert all four arm -- the decode lane armed when the native extension is built, else an explicit declined-to-stock verdict, never silent absence. The hc_m4 read-once test is rewritten to assert read-at-use, and the mask-fuse test drops its now-defunct cache_clear() calls.
8c53d81 to
72f58d4
Compare
… 391 remainder port Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391 remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K sparse-GQA decode extension. The two lanes here add to the same fixed-M4 lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa), all cleanly additive). PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never merged, but the maintainer independently re-landed the Flash-Next stack under the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane (mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel (mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is re-siting the arming and native loading off the absent full_stack_env onto upstream's own machinery. Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out): - ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram rows and the auxiliary embedding plane is produced with mx.async_eval outside the compiled verifier. The stock owner-side row cache is preserved; declines to stock with a printed reason when the ple_cpu_rows extension is not built. - qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the pool kernel metadata once per indexer and shares one inv_freq object. Rebase changes vs the closed-PR commit: - mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_* names kept as aliases (primary wins when both set). - mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary. - mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET). - mtplx/runtime.py: the two installs run after the fixed-M4 verify install, logging instead of the removed _print_install_receipt. - mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so normalize_runtime_env_overrides accepts the server-stamped values. - mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py). - scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized release wheel carries a signed ple_cpu_rows Mach-O. - scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows. 772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default off) touches the same _SidecarGather rows but at the disk-layout layer; it is orthogonal to this runtime-scheduling lane and does not subsume it. CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files plus the wheel-bundler test.
Arming audit (battery, 2026-09-07): `mtplx serve` imports generation / runtime / model modules before parse_args stamps the auto-arm env, so any flag whose reader resolves at module import freezes its default before the stamp lands and the auto-arm's setdefault is a silent no-op as launched. A column-0 scan of every module holding a stamped key's reader found exactly four such readers among the ~31 auto-armed keys; every other stamped key reads the environment at use or is consumed from config.json at model load. All four are decode-verify lanes, so the release control (71.17 tok/s at 16K) ran without them -- a plausible slice of the 71->81 gap. Resolve all four at use (read the environment each call; the module global stays a test/force override; the env is frozen once serving starts, so two traces of one graph still read the same value): - MTPLX_QWEN4_DRAFT_K20_PRESCATTER: qwen4_draft_k20_prescatter._ENABLED, and generation.py's cached _QWEN4_DRAFT_K20_PRESCATTER (removed; the one draft consult site calls the reader). - MTPLX_QWEN4_BLOCK_VERIFY: qwen4_block_verify._ENABLED, and generation.py's cached _QWEN4_BLOCK_VERIFY (removed; the accept-loop consult calls the reader). - MTPLX_QWEN4_OPDIET (+ _ITEMS): runtime_options. - MTPLX_QWEN4_VERIFY_GLUE (+ _ITEMS): runtime_options (reset hook kept, now forcing the globals). No default value changed; keys that already read at use are untouched. Upstream's STRICT_CLAIMS and BATCH_PAGED_OFFSETS are also import-frozen but are not auto-armed (operator sets them pre-launch), so they are left as-is. Tests: tests/test_qwen4_remainder_arming.py extended to assert all four arm in the served order (import first, stamp, read) and that the fixed-M4 auto-arm stamps OPDIET / BLOCK_VERIFY / VERIFY_GLUE and their readers then arm. The block-verify and draft-k20 source-inspection tests and the opdiet read-once test are rewritten to assert read-at-use.
The arming audit's lesson: gate on the install verdict, not the env. Three
decode-verify lanes had no per-window observable in
/health qwen4_install_reports -- draft_k20_prescatter, block_verify, opdiet --
so a served window could not confirm they engaged. Add read-only reports (no
behaviour change, no defaults touched):
- draft_k20_prescatter: {armed (read at use), engaged (first-use latch set when
claim_draft_route installs the route), receipt (the last install receipt)}.
- block_verify: {armed, engaged (latched when a block verifier is built for the
accept loop)} plus a one-shot "[mtplx] MTPLX_QWEN4_BLOCK_VERIFY armed:" log.
- opdiet: {armed, items (configured selection), applied (first-use latch of the
items that actually ran at a gated site)}.
Each appears only when ARMED (read at use, gate-able without a request), so an
unarmed lane stays absent (== off) like the other lanes; the engaged/applied
latch rides inside the armed report.
CPU test: tests/test_qwen4_remainder_arming.py asserts the three reports are
absent when off / =0 and present with armed True under a served-order stamp.
… 391 remainder port Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391 remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K sparse-GQA decode extension. The two lanes here add to the same fixed-M4 lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa), all cleanly additive). PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never merged, but the maintainer independently re-landed the Flash-Next stack under the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane (mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel (mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is re-siting the arming and native loading off the absent full_stack_env onto upstream's own machinery. Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out): - ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram rows and the auxiliary embedding plane is produced with mx.async_eval outside the compiled verifier. The stock owner-side row cache is preserved; declines to stock with a printed reason when the ple_cpu_rows extension is not built. - qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the pool kernel metadata once per indexer and shares one inv_freq object. Rebase changes vs the closed-PR commit: - mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_* names kept as aliases (primary wins when both set). - mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary. - mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET). - mtplx/runtime.py: the two installs run after the fixed-M4 verify install, logging instead of the removed _print_install_receipt. - mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so normalize_runtime_env_overrides accepts the server-stamped values. - mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py). - scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized release wheel carries a signed ple_cpu_rows Mach-O. - scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows. 772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default off) touches the same _SidecarGather rows but at the disk-layout layer; it is orthogonal to this runtime-scheduling lane and does not subsume it. CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files plus the wheel-bundler test.
72f58d4 to
8f0256d
Compare
… 391 remainder port Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391 remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K sparse-GQA decode extension. The two lanes here add to the same fixed-M4 lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa), all cleanly additive). PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never merged, but the maintainer independently re-landed the Flash-Next stack under the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane (mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel (mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is re-siting the arming and native loading off the absent full_stack_env onto upstream's own machinery. Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out): - ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram rows and the auxiliary embedding plane is produced with mx.async_eval outside the compiled verifier. The stock owner-side row cache is preserved; declines to stock with a printed reason when the ple_cpu_rows extension is not built. - qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the pool kernel metadata once per indexer and shares one inv_freq object. Rebase changes vs the closed-PR commit: - mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_* names kept as aliases (primary wins when both set). - mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary. - mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET). - mtplx/runtime.py: the two installs run after the fixed-M4 verify install, logging instead of the removed _print_install_receipt. - mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so normalize_runtime_env_overrides accepts the server-stamped values. - mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py). - scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized release wheel carries a signed ple_cpu_rows Mach-O. - scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows. 772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default off) touches the same _SidecarGather rows but at the disk-layout layer; it is orthogonal to this runtime-scheduling lane and does not subsume it. CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files plus the wheel-bundler test.
8f0256d to
cb32386
Compare
… 391 remainder port Rebase of PR youssofal#475 (cached async PLE + pooled-key rowsel) onto the PR 391 remainder port head (perf/qwen38-391-remainder-main = upstream main 2.11.2 + the four remainder lanes: HC_M4, prefill mask fuse, QSA query tile, and the QSA split-K sparse-GQA decode extension. The two lanes here add to the same fixed-M4 lane_defaults / _QWEN4_PORT_KEYS block, and their native loader (ple_cpu_rows) and wheel-signer entry union with the QSA sparse-decode extension's (mtplx_native_qsa), all cleanly additive). PR 391 is closed and its Fable delivery stack (mtplx/full_stack_env.py, the turbo-full-stack profile, mtplx/native's QSA sparse-GQA loader) was never merged, but the maintainer independently re-landed the Flash-Next stack under the MTPLX_QWEN4_*/MTPLX_QSA_* namespace, so both lanes' base contracts are present upstream and both apply: the fixed-M4 compiled-verify auxiliary plane (mtplx/qwen4_fixed_verify.py) and the QSA indexer pooled-key kernel (mtplx/kernels/qsa_indexer_prepare._pool_keys_kernel). The work here is re-siting the arming and native loading off the absent full_stack_env onto upstream's own machinery. Lanes (armed by default for a served fixed-M4 Flash-Next pack, per-lane opt-out): - ple_cached_aux: a native CPU-stream provider stages the fixed 64 M4 n-gram rows and the auxiliary embedding plane is produced with mx.async_eval outside the compiled verifier. The stock owner-side row cache is preserved; declines to stock with a printed reason when the ple_cpu_rows extension is not built. - qsa_pooled_rowsel: the twelve QSA indexers' pooled-key preparation binds the pool kernel metadata once per indexer and shares one inv_freq object. Rebase changes vs the closed-PR commit: - mtplx/qwen4_aux_lanes.py rewritten off full_stack_env: primary keys are MTPLX_QWEN4_PLE_CACHED_AUX / MTPLX_QSA_POOLED_ROWSEL, the PR 391 MTPLX_FABLE_* names kept as aliases (primary wins when both set). - mtplx/server/openai.py: the two keys join the fixed-M4 lane_defaults and _QWEN4_PORT_KEYS (so the existing pop-loop kill-switch honours KEY=0), with an alias pre-step mirroring an operator's MTPLX_FABLE_* export onto the primary. - mtplx/qsa_pooled_rowsel.py: op-diet contract re-pointed from the absent fable_opdiet_enabled to upstream's qwen4_opdiet_enabled (MTPLX_QWEN4_OPDIET). - mtplx/runtime.py: the two installs run after the fixed-M4 verify install, logging instead of the removed _print_install_receipt. - mtplx/profiles.py: the two keys added to MODEL_RUNTIME_ENV_OVERRIDE_KEYS so normalize_runtime_env_overrides accepts the server-stamped values. - mtplx/native/__init__.py: a minimal PLE-only loader (load_ple_cpu_rows_extension / ple_cpu_rows_unavailable_reason); PR 391's qsa_sparse_gqa loader is not reproduced (upstream loads native QSA via kernels/qsa_prefill_direct.py). - scripts/bundle_native_runtime_wheel.py: accepts and Developer-ID signs the new mtplx_native_ple_cpu_rows extension alongside mtplx_qsa_kernels, so a notarized release wheel carries a signed ple_cpu_rows Mach-O. - scripts/fable/setup_over100_venv.sh: builds only ple_cpu_rows. 772f5be's opt-in interleaved n-gram row cache (MTPLX_NGRAM_ROW_FILE, default off) touches the same _SidecarGather rows but at the disk-layout layer; it is orthogonal to this runtime-scheduling lane and does not subsume it. CPU tests (venv mlx 0.32.2, no GPU): 84 passed across the six lane test files plus the wheel-bundler test.
cb32386 to
27d5ff6
Compare
|
David, thank you for this. I read the full diff, ran your test files (562 passed, 23 skipped on CPU), built both native extensions in a scratch directory, and checked the C++ n-gram row hash against the numpy one on 3,690 cases (0 mismatches). Here is where I landed. Things I want to take:
Things that need fixing before they can merge:
Two smaller notes: the HC-M4 kernel reads bf16 weights, not an 8-bit copy as the description says, and the On the speed claim: the +15% at 16K is the whole stack including the frozen-key re-engagement, and your own knock-out table attributes roughly 6 to 8 points to the six lanes themselves. I have already landed a general fix for the frozen keys on our side (one refresh after the profile environment is applied), so that part of the gain will ship regardless. For the four rounding-class lanes (HC read, mask fuse, sparse decode, and the query tile when it is active) I want the HumanEval/MBPP cells in the table filled in before they go on by default, and I will run an A/B on this machine with each lane switched individually. If you can rebase with the four fixes above and split the docs artifacts into their own commit, I can take the exact lanes right away and evaluate the rest with numbers. |
|
A short update on the exact lanes. The pooled row selection lane is ported on our side with your authorship, gated so that a daemon with op-diet off still boots on the stock path, with your two test files and our environment and health tests green. On this M5 Max with the Flash-Next Speed pack, an 8.8k-token prompt decoded greedily to 1,024 tokens is byte-identical with the lane on and off, with the same acceptance ladder and 71.8 against 72.5 tok/s, so at that length it changes nothing either way. It needs the 16K and 100K cells before it goes on by default. The cached PLE rows lane waits on the native pins from the review. |
Add the cascade context-ladder decode chart (cascade_decode_by_context.svg): exact pairing and cascade alpha 0.0/0.5/1.0/2.0 across 1K-128K, fastest of seeds with min-max variance bands; 16,384 merged from the arm-G sweep; 261,120 absent (every arm OOMs on the youssofal#475 base without youssofal#482). Update the charts manifest.
cascade_decode_by_context now carries eight arms: the exact pairing, the OPT rule at alpha 0.0/0.25/0.5/0.75/1.0/2.0, and the TokenV3 rule at alpha 0.95 across 1K-128K (fastest of three seeds, min-max bands). 261,120 is absent because every arm exceeds the memory knob on the youssofal#475 base without youssofal#482. Chart and manifest only; nothing under mtplx/.
Two charts comparing this pull request's operating points against youssofal#475 and youssofal#478, both re-derived from the receipt json rather than from any table: cascade_vs_475_478_16k.svg decode at 16,384 tokens for release 2.11.2 (exact), youssofal#475 (exact), youssofal#478 (typical 0.09), OPT alpha 0.25, and TokenV3 alpha 0.75 and 0.95, each bar annotated with its own HumanEval strict pass@1 cascade_vs_475_478_ladder.svg decode against context size, 1,024 to 131,072, for the same four arms that have a full ladder Bar height and line point are the fastest seed, every band is min-max, and acceptance mode is in each label: a tok/s figure is not readable without it. The alphas of the two rules are different quantities and the captions say so. No arm carries a 261,120 point, because on the youssofal#475 base without youssofal#482 every arm on this pack exceeds the memory knob.
Two charts comparing this pull request against youssofal#475 and youssofal#478 on BOTH model packs, so the pack and the optimization set are separated rather than confounded. Both are re-derived from the receipt json: bare_vs_475_478_16k.svg decode at 16,384 tokens, the Optimized-Speed pack (release exact, youssofal#475 exact, youssofal#478 typical 0.09) beside the Bare-Speed pack (the same three, then this pull request at the exact law, at typical 0.09, at cascade OPT alpha 0.5 and at cascade TokenV3 alpha 0.95) bare_vs_475_478_ladder.svg decode against context size, 1,024 to 261,120, for youssofal#475 and youssofal#478 on both packs and this pull request on Bare-Speed The two colour families are the two packs, not the two pull requests. Bar height and line point are the fastest seed, every band is min-max, and the acceptance mode is in each label: a tok/s figure is not readable without it. At 261,120 the Bare-Speed arms carry a point and the Optimized-Speed arms do not, because the latter exceed the memory knob on the youssofal#475 base without youssofal#482.
… as the reference David's rulings applied to the youssofal#485 chart set: RULING 1 ("youssofal#475 is NOT exact"): no chart, legend, reference line, caption or manifest labels youssofal#475 or release 2.11.2 as "exact". The acceptance-off state is now named "acceptance mode off" / "cascade off"; the youssofal#475 arm is "youssofal#475 (base)". * cascade_decode_vs_alpha: reference lines relabeled "cascade off, youssofal#475 base (82.80)" and (see below) "typical 0.2 (99.16)". * cascade_decode_by_context: arm-C-caspair legend "exact (no cascade)" -> "cascade off (youssofal#475 base)". * cascade_vs_475_478_16k / _ladder: "youssofal#475 (base)", "acceptance mode off" / "cascade off"; the exact-acceptance disclaimer now names the ordinary speculative-decoding acceptance law, not any arm. RULING 2 (typical 0.2, not 0.09, is the youssofal#478 reference): every youssofal#485-vs-youssofal#478 comparison now references youssofal#478 at typical threshold 0.2 (pooled 16,384 window, 99.16 tok/s fastest of n=9, HumanEval strict pass@1 0.9695), replacing typical 0.09. Applied to cascade_decode_vs_alpha (dotted reference line), the 16K bars and the context ladder. The §2.4 ABAB (typical 0.09 vs TokenV3 0.95) is a separate measurement and is unchanged. Charts re-rendered from receipts (fastest-of-seeds, min-max band); manifest bytes/sha256 refreshed. Docs only; no receipts touched. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…typical 0.2 David's rulings applied to the youssofal#488 chart set: RULING 1 ("youssofal#475 is NOT exact"): the acceptance-off state is named "acceptance mode off" (bars/ladders/legends), the youssofal#475 arm is "youssofal#475 (base)", and release 2.11.2 is never labeled "exact". * bare_16k_bars, bare_decode_16k_windows: arm A/B legends "..., exact" -> "..., acceptance mode off". * decode_tok_s, prefill_tok_s, ttft_s, peak_memory_gb, summary (context ladder): A/B legends "..., exact (FR-Spec dark)" -> "..., acceptance mode off (FR-Spec dark)". * bare_vs_475_478_16k / _ladder: release/youssofal#475 relabeled; the disclaimer now names the ordinary speculative-decoding acceptance law, not any arm. * pr391-charts/pr391-typical-acceptance-decode: "exact (lane off)" -> "typical off"; footnote "+/-1 task of exact" -> "of typical off". RULING 2 (typical 0.2 is the youssofal#478 reference): the Optimized-Speed youssofal#478 comparison arm in bare_vs_475_478_16k and _ladder now references typical 0.2 (16,384 fastest 99.16 tok/s, n=9), replacing typical 0.09. The Bare-Speed pack has no typical-0.2 measurement, so its youssofal#478 arm stays typical 0.09, explicitly labeled (data gap, not a substitution). Layout: the longer mode labels needed room. bare_vs_475_478_16k now scales its width with the bar count and drops the pack group-labels below their bracket lines clear of the caption; bare_vs_475_478_ladder widens the right gutter so the legend is not clipped; the summary grid falls back to a two-column legend when labels are long so nothing is cut off. Charts re-rendered from receipts (fastest-of-seeds, min-max band). Docs only. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…al off"
David's RULING 1 ("youssofal#475 is NOT exact"): the typical sweep's acceptance-off arm
must not be labeled "exact". It is the ordinary acceptance law with typical
acceptance turned off, so it is now named "typical off".
* acc_vs_length, acc_vs_speed, acc_vs_length_mbpp, acc_vs_speed_mbpp: series
label "exact (off)" -> "typical off"; the MBPP single-series title/caption
"exact arm" -> "typical-off arm".
* tok_s_vs_passk: the base point "exact (off)" -> "typical off".
* pr391-charts/pr391-typical-acceptance-decode: legend "exact (lane off)" ->
"typical off"; footnote "+/-1 task of exact (153/164)" -> "of typical off
(153/164)".
* manifest.json alt/caption "exact arm" -> "typical-off arm"; bytes refreshed.
RULING 2 does not change this PR's charts: the typical sweep shows every
threshold (off / 0.09 / 0.2 / 0.4) as its own series, not a youssofal#485-vs-youssofal#478
comparison, so there is no single youssofal#478 reference to switch here.
Charts re-rendered from the sweep receipts. Docs only.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…typical 0.2 David's rulings applied to the youssofal#488 chart set: RULING 1 ("youssofal#475 is NOT exact"): the acceptance-off state is named "acceptance mode off" (bars/ladders/legends), the youssofal#475 arm is "youssofal#475 (base)", and release 2.11.2 is never labeled "exact". * bare_16k_bars, bare_decode_16k_windows: arm A/B legends "..., exact" -> "..., acceptance mode off". * decode_tok_s, prefill_tok_s, ttft_s, peak_memory_gb, summary (context ladder): A/B legends "..., exact (FR-Spec dark)" -> "..., acceptance mode off (FR-Spec dark)". * bare_vs_475_478_16k / _ladder: release/youssofal#475 relabeled; the disclaimer now names the ordinary speculative-decoding acceptance law, not any arm. * pr391-charts/pr391-typical-acceptance-decode: "exact (lane off)" -> "typical off"; footnote "+/-1 task of exact" -> "of typical off". RULING 2 (typical 0.2 is the youssofal#478 reference): the Optimized-Speed youssofal#478 comparison arm in bare_vs_475_478_16k and _ladder now references typical 0.2 (16,384 fastest 99.16 tok/s, n=9), replacing typical 0.09. The Bare-Speed pack has no typical-0.2 measurement, so its youssofal#478 arm stays typical 0.09, explicitly labeled (data gap, not a substitution). Layout: the longer mode labels needed room. bare_vs_475_478_16k now scales its width with the bar count and drops the pack group-labels below their bracket lines clear of the caption; bare_vs_475_478_ladder widens the right gutter so the legend is not clipped; the summary grid falls back to a two-column legend when labels are long so nothing is cut off. Charts re-rendered from receipts (fastest-of-seeds, min-max band). Docs only. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… as the reference David's rulings applied to the youssofal#485 chart set: RULING 1 ("youssofal#475 is NOT exact"): no chart, legend, reference line, caption or manifest labels youssofal#475 or release 2.11.2 as "exact". The acceptance-off state is now named "acceptance mode off" / "cascade off"; the youssofal#475 arm is "youssofal#475 (base)". * cascade_decode_vs_alpha: reference lines relabeled "cascade off, youssofal#475 base (82.80)" and (see below) "typical 0.2 (99.16)". * cascade_decode_by_context: arm-C-caspair legend "exact (no cascade)" -> "cascade off (youssofal#475 base)". * cascade_vs_475_478_16k / _ladder: "youssofal#475 (base)", "acceptance mode off" / "cascade off"; the exact-acceptance disclaimer now names the ordinary speculative-decoding acceptance law, not any arm. RULING 2 (typical 0.2, not 0.09, is the youssofal#478 reference): every youssofal#485-vs-youssofal#478 comparison now references youssofal#478 at typical threshold 0.2 (pooled 16,384 window, 99.16 tok/s fastest of n=9, HumanEval strict pass@1 0.9695), replacing typical 0.09. Applied to cascade_decode_vs_alpha (dotted reference line), the 16K bars and the context ladder. The §2.4 ABAB (typical 0.09 vs TokenV3 0.95) is a separate measurement and is unchanged. Charts re-rendered from receipts (fastest-of-seeds, min-max band); manifest bytes/sha256 refreshed. Docs only; no receipts touched.
…al off"
David's RULING 1 ("youssofal#475 is NOT exact"): the typical sweep's acceptance-off arm
must not be labeled "exact". It is the ordinary acceptance law with typical
acceptance turned off, so it is now named "typical off".
* acc_vs_length, acc_vs_speed, acc_vs_length_mbpp, acc_vs_speed_mbpp: series
label "exact (off)" -> "typical off"; the MBPP single-series title/caption
"exact arm" -> "typical-off arm".
* tok_s_vs_passk: the base point "exact (off)" -> "typical off".
* pr391-charts/pr391-typical-acceptance-decode: legend "exact (lane off)" ->
"typical off"; footnote "+/-1 task of exact (153/164)" -> "of typical off
(153/164)".
* manifest.json alt/caption "exact arm" -> "typical-off arm"; bytes refreshed.
RULING 2 does not change this PR's charts: the typical sweep shows every
threshold (off / 0.09 / 0.2 / 0.4) as its own series, not a youssofal#485-vs-youssofal#478
comparison, so there is no single youssofal#478 reference to switch here.
Charts re-rendered from the sweep receipts. Docs only.
…typical 0.2 David's rulings applied to the youssofal#488 chart set: RULING 1 ("youssofal#475 is NOT exact"): the acceptance-off state is named "acceptance mode off" (bars/ladders/legends), the youssofal#475 arm is "youssofal#475 (base)", and release 2.11.2 is never labeled "exact". * bare_16k_bars, bare_decode_16k_windows: arm A/B legends "..., exact" -> "..., acceptance mode off". * decode_tok_s, prefill_tok_s, ttft_s, peak_memory_gb, summary (context ladder): A/B legends "..., exact (FR-Spec dark)" -> "..., acceptance mode off (FR-Spec dark)". * bare_vs_475_478_16k / _ladder: release/youssofal#475 relabeled; the disclaimer now names the ordinary speculative-decoding acceptance law, not any arm. * pr391-charts/pr391-typical-acceptance-decode: "exact (lane off)" -> "typical off"; footnote "+/-1 task of exact" -> "of typical off". RULING 2 (typical 0.2 is the youssofal#478 reference): the Optimized-Speed youssofal#478 comparison arm in bare_vs_475_478_16k and _ladder now references typical 0.2 (16,384 fastest 99.16 tok/s, n=9), replacing typical 0.09. The Bare-Speed pack has no typical-0.2 measurement, so its youssofal#478 arm stays typical 0.09, explicitly labeled (data gap, not a substitution). Layout: the longer mode labels needed room. bare_vs_475_478_16k now scales its width with the bar count and drops the pack group-labels below their bracket lines clear of the caption; bare_vs_475_478_ladder widens the right gutter so the legend is not clipped; the summary grid falls back to a two-column legend when labels are long so nothing is cut off. Charts re-rendered from receipts (fastest-of-seeds, min-max band). Docs only.
…n from the README The README carried the same vocabulary defects the pull request body did, and was fixed the same way, against pr-bodies/GLOSSARY.md. "Exact" describes the acceptance LAW, never an arm. This pack's arms A and B are release 2.11.2 and youssofal#475, which are rounding-class builds that happen to run exact acceptance, so calling them "(A, exact)" reads as a claim about their kernels. Every arm label now states which lossy mode is off: the five four-arm tables, the all-kernels table, the HumanEval rows, and the two prose sites that said "youssofal#475 exact code" and "16K exact". Section 6 claimed the four-arm sweep was still being measured while Sections 2.1 to 2.5 published it in full, and listed the queue order and the GPU-lock protocol to get there. It now names only the three real gaps. The Section 2 banner made the same stale claim and is gone, the pack footnote no longer names an internal role, and the method row states what the GPU lock guarantees rather than how it is taken. Checked with pr-bodies/check_glossary.py, which now gates files by path.
…senses of "exact" Against pr-bodies/GLOSSARY.md, the same pass the pull request bodies had. The arm table gave an acceptance mode for D and E only, so A and C read as though they had none. Both run with typical off, and their rows now say so. "Exact" appears once here, in the optimization-class sense: an exact optimization is byte-for-byte the stock path, against a rounding-class one that differs only by floating-point rounding. Nothing in the file bound that, so it was indistinguishable from the exact acceptance LAW that youssofal#478 and youssofal#485 use the same word for. The paragraph now binds it and states that it never labels an arm, since arms A and C are rounding-class builds overall. The prose mixed short context labels with exact counts; prose now gives the counts and the short form is left to table row labels.
…verflow The long-context verdict table blamed a QSA-indexer prefill transient for the 261,120-token OOM on all four arms. The prefill completes on every arm; the overflow is the first decode step's speculative-verify KV-cache write, where TensorOffsetKVCache.update_and_fetch used the functional mx.slice_update and reallocated the full per-layer KV buffers (about 6.4 GB in one command buffer). The probes W1 and W2 targeted prefill and so could not have fit. This matches the youssofal#475 body's Section 1.5 and the youssofal#482 fix arm, which fits 261,120 on all three cold seeds at 100.82 GB.
cascade_humaneval_by_alpha.svg: strict and completed-task pass@1 as grouped bars for cascade off (youssofal#475 base), the OPT rule at alpha 0.0/0.25/0.5/0.75 and the TokenV3 rule at 0.75/0.95, with the cascade-off strict level as a reference line, so the quality cost of the OPT rule is visible next to TokenV3. Data: evalsweep478/armG_summary.json and the youssofal#478 sweep's typical-off cell; nothing re-measured.
Four hot-path gates are read once at import so decode never touches os.environ and two traces of one compiled graph cannot disagree: MTPLX_QWEN4_OPDIET, MTPLX_QWEN4_VERIFY_GLUE, MTPLX_QWEN4_DRAFT_K20_PRESCATTER and MTPLX_QWEN4_BLOCK_VERIFY (generation.py keeps copies of the last two). `mtplx serve` imports those modules before ServerState stamps the model family's runtime env, and the Flash-Next lane defaults arm all four, so every served daemon since 2.11.1 reported the keys as configured on /health while running with all of them off: the PR #391 ports measured on 2026-09-03 never reached a user, and the K20 prescatter and op diet receipts were harness-only. davidtai's PR #475 found the same four frozen readers. apply_profile_env now re-reads the gates in their owning modules right after it writes os.environ (runtime_options.refresh_env_flags, which also re-copies generation's constants) and returns the live values as a receipt. Every call site of apply_profile_env runs before a model load, so nothing has been compiled when the values change and the traces-agree invariant holds. A test mapping passed as `environ` leaves module state alone. Block verification is re-engaged only now that its law is exact (previous commit).
|
2.11.3 ships the fix for the frozen setting readers this PR found, with credit in the release notes. The pooled row-selection kernel measured a tie on the 9k-token prompt (79.4 to 79.6 tok/s with it, 79.6 to 79.7 without), so it is not in this release and the PR stays open. Release: https://github.com/youssofal/MTPLX/releases/tag/v2.11.3 |
Six serving optimizations for Qwen3.8 Flash-Next: the PR 391 remainder and two exact decode optimizations
This pull request adds six serving optimizations to the Qwen3.8 Flash-Next path. The base is upstream MTPLX 2.11.2. Four are the PR 391 optimizations that 2.11.2 did not re-land: a hyper-connection read kernel, a prefill causal-mask fuse, a QSA prefill query tile, and a QSA split-K sparse decode kernel. Two are exact decode optimizations: cached async PLE rows and pooled QSA row selection. The four remainder optimizations are rounding-class, so the output differs from the stock path only by floating-point rounding. The two aux optimizations are exact, so their output is byte-for-byte the base. The base commits also re-engage four upstream verify optimizations that 2.11.2 shipped but left silently off when served (Section 3).
The comparison has two arms: release 2.11.2 alone (A) and release plus this pull request's six optimizations (C, this pull request). The headline is C against A.
Headline figures at 16,384 tokens of context, 3 seeds, cold prefill:
xhigh, one seed): HumanEval strict pass@1 0.9695 on release against 0.9634 with the six optimizations, completed-task 1.0000 against 0.9814; MBPP at that sampler ran for one arm only (Section 4).Terms used in this document:
1. Benchmark results
Two arms ran in one benchmark run on one machine. Each arm received the same request body per cell. Section 2 gives the settings.
1.1 Summary
Summary grid: decode, prefill, time-to-first-token and peak memory by context size, one line per arm, each point the fastest of its seeds. Bars show the slowest-to-fastest spread.
1.2 Decode tok/s
The 261,120-token cell is in Section 1.5; both arms exceed the memory knob there.
1.3 Prefill tok/s
The two exact decode optimizations do not change prefill; the candidate column is arm C prefill.
1.4 Time-to-first-token and peak memory
Time-to-first-token by context size.
Peak memory by context size. The 100 GiB cap holds on every arm that reaches the cell.
The per-context tables for time-to-first-token and peak memory are in the perf report under
docs/perf/.1.5 261,120-token context
Both release 2.11.2 and this pull request exceed the 100 GiB memory knob at 261,120 tokens on this machine under the served launch. Metal returns
insufficient_memoryat the same point on both, at a sampled peak of 100.82 GB on release and 100.79 GB on this pull request.The prefill is not what overflows. Every arm's receipt shows the full prompt read and tokens already emitted before the failure:
new_prefill_tokens261120, time to first token 243.5 s on release and 241.9 s on this pull request, thencompletion_tokens12 on release and 3 on this pull request, withfinish_reasonerrorand[METAL] Command buffer execution failed: Insufficient Memory. The overflow is in the first decode step.That decode step is the fixed four-row speculative verify, and the allocation that overflows is its KV-cache write. The verify holds its KV in
TensorOffsetKVCache, whoseupdate_and_fetchwrites the new rows with the functionalmx.slice_update. That op returns a new array, so it reallocates the whole[1, 2, capacity, 256]bf16 buffer for keys and again for values, about 0.27 GB each at the 262K capacity, in every one of the 12 full-attention layers: about 6.4 GB inside one verify command buffer, on a steady state already near 100.8 GB against a 107.374 GB limit. The stockKVCacheused by plain autoregressive decode writes in place and allocates nothing, which is why the same prompt fits with native MTP off. A direct probe of that cache at this geometry measures 6.42 GB resident and, once the write is in place, a 0.00 GB update transient.A second, smaller term sits in the same step: with a query length of 3 to 6 rows and 12 query heads per key/value head, query length times 12 is past the 32-row limit of MLX's fused vector attention kernel, so MLX declines the fused call and the unfused route materializes a
[24, S, 261120]score tensor per layer. Removing only that term does not make the cell fit. A head-chunked build engaged and still failed at the same 100.82 GB peak. The prefill mask-fuse optimization declines in exactly this band, query length 3 to 8, so it does not help either; the wide 256-row prefill chunks still fuse. The analysis trail is inover100-reports/oom255k/report.md.The earlier #391 tree only looked like it fit. That fit was one cold seed; its other two seeds also ran out of memory at the same cell.
This pull request does not change 261,120-token behaviour; the fix is #482, which writes the verify KV in place and keeps the verify attention on the fused kernel. Two windows confirm the cause and the fix:
KVCache)stopWith the fix the cell decodes at 62.84 tok/s fastest (56.82-62.84), prefill 1082.9 tok/s, and sits about 6.5 GB under the knob, while 16,384-token output stays byte-identical to release.
1.6 Interleaved 16,384-token protocol
The arms ran interleaved at the 16,384 / 1,024 cell in the order A C D E, three windows per arm (D and E are the #478 typical-acceptance arms). The paired delta compares each arm-C window against its neighbouring arm-A windows, which cancels the slow warm-up drift across the run.
Release 71.97 tok/s and candidate 82.84 tok/s are each the fastest window. The fastest-window gap is +15.1%. The paired per-seed 95% interval on the mean delta is [+1.76, +8.89] tok/s. Drift bracket (slowest-to-fastest spread between the release windows) +0.26 tok/s.
1.7 Supplementary measurements
These two sets are supplementary, not headline. Both use the fastest-of-seeds rule.
Greedy code-eval on the release arm, kept for reference only. It is not the quality gate: greedy is temperature 0, which is not the setting anything here was benchmarked at.
Three per-optimization knock-outs at 16,384 tokens, on this pull request's tree, each one window of three seeds. The plan was ten leave-one-out windows; the run was cancelled after three, which had already finished and passed their gates. The other seven were not measured, and each of these three is a single window, so its per-seed range is wide: read the change as indicative, not a paired result. This pull request's fastest ABAB decode is 82.84 tok/s.
The remainder-only arm: release plus the four ported optimizations and the four re-engaged keys, without the two exact decode optimizations. The battery calls it arm B. It was stopped after 32K, so it is not a headline. Decode is the fastest of its seeds with the slowest-to-fastest range.
2. Benchmark method
Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed, revision29ba90f82124961d0d902a9ea9bbb1034972af2fmtplx serveselects for this pack; a fresh SSD session cache directory per windowxhighxhigh)Per-optimization attribution was not measured on this base; the numbers of record are the whole-PR deltas from the ladder (Section 1) and the 16K interleave (Section 1.6). A partial remainder-only arm (release plus the four remainder optimizations, at 1K, 8K and 32K) is kept as supplementary data in the docs appendix under
docs/perf/; it is not one of the two arms in these tables.3. What each optimization does
MTPLX_QWEN4_HC_M4MTPLX_QSA_SPARSE_DECODEMTPLX_QWEN4_PLE_CACHED_AUXMTPLX_QSA_POOLED_ROWSELMTPLX_QWEN4_PREFILL_MASK_FUSEMTPLX_QSA_PREFILL_QUERY_TILEDecode
Hyper-connection read at M4 (
MTPLX_QWEN4_HC_M4)mtplx/kernels/qwen4_m4_hyper_read.py,mtplx/models/qwen4_exp.py.MTPLX_QWEN4_HC_M4=0turns it off.QSA split-K sparse decode (
MTPLX_QSA_SPARSE_DECODE)mtplx_native_qsa; if the extension is not built the optimization declines to stock, prints the reason, and still serves.mtplx/kernels/qsa_sparse_decode.py,native_extensions/qsa_sparse_gqa/(mtplx_native_qsa).MTPLX_QSA_SPARSE_DECODE=0turns it off.Cached async PLE rows (
MTPLX_QWEN4_PLE_CACHED_AUX)mx.async_evalso it overlaps the compiled replay instead of sitting on the critical path. It needs the native extensionmtplx_native_ple_cpu_rows; if the extension is not built the optimization declines to stock and still serves, still exact.mtplx/ple_cached_aux.py,mtplx/ple_cached_row_handoff.py,native_extensions/ple_cpu_rows/(mtplx_native_ple_cpu_rows).MTPLX_QWEN4_PLE_CACHED_AUX=0turns it off.Pooled QSA row selection (
MTPLX_QSA_POOLED_ROWSEL)qsa_indexer_pool_keys_metalkernel metadata once per indexer at construction and shares oneinv_freqobject across all twelve indexers, so the repeated per-step setup leaves the decode step. Stock MLX, no native extension. The install validates the 48-layer QSA layout, the per-indexer geometry, the RMS-norm epsilon, the RoPE scale and the sharedinv_freqidentity, and a contract failure fails the model load.mtplx/qsa_pooled_rowsel.py.MTPLX_QSA_POOLED_ROWSEL=0turns it off.Prefill
Prefill causal-mask fuse (
MTPLX_QWEN4_PREFILL_MASK_FUSE)mtplx/models/qwen4_exp.py.MTPLX_QWEN4_PREFILL_MASK_FUSE=0turns it off.QSA prefill query tile (
MTPLX_QSA_PREFILL_QUERY_TILE)mtplx/qwen4_prefill_chunk.py,mtplx/models/qwen4_exp.py.MTPLX_QSA_PREFILL_QUERY_TILE=0turns it off.Every old
MTPLX_FABLE_*name still works as an alias when the new key is unset.Four upstream optimizations that were silently off
These are not new optimizations. They are upstream 2.11.2's own re-landed verify optimizations, now actually engaged; an audit of the ~31 auto-armed keys found four whose readers were dead as served.
MTPLX_QWEN4_DRAFT_K20_PRESCATTERqwen4_draft_k20_prescatter.py:200(re-cachedgeneration.py:323)MTPLX_QWEN4_BLOCK_VERIFYqwen4_block_verify.py:132(re-cachedgeneration.py:337)MTPLX_QWEN4_OPDIETruntime_options.py:61MTPLX_QWEN4_VERIFY_GLUEruntime_options.py:159Cause:
mtplx serveimports the generation and runtime modules before it stamps the auto-armed keys, so these four readers froze their default (off) at import and the auto-arm was a silent no-op; the arm A (release) numbers above were measured with the four off, which is the release as a user runs it. Fix: this pull request's base commits change the four readers to resolve the environment at use, change no defaults, and addtests/test_qwen4_remainder_arming.py, which replays the served import order. Arm C carries the fix, so it includes the four re-engaged optimizations.What is not ported
The fifth PR 391 optimization, verify graph-build overlap (
MTPLX_QWEN4_GRAPH_BUILD_OVERLAP), is not in this pull request. Upstream 2.11.2 re-landed the fixed-M4 verify as a single compiled graph. The optimization rode a prefix and suffix split of that graph, and that split does not exist upstream. Porting it is a redesign of the compiled verify, and its whole claim is a submission-timing reorder that cannot be proven on the CPU path.Feasibility, from the port report: on the PR 391 campaign the optimization targeted 1.934 ms per cycle of host-late GPU idle at 16,384 tokens; at the default one-layer split the binding term is about 0.53 ms per cycle, about +1.4 tok/s (about 1.4%); at a three-to-four-layer split it peaks near 1.7 to 1.8 ms per cycle, about +3.6 to +4.7 tok/s (about 4.5%). The optimization is exact by construction, so it is a speed-only optimization, not a quality risk. Recommendation: defer it to a dedicated split-compilation task with a bit-exact A/B gate. The one-page feasibility is in
over100-reports/remainder-port-report.md.4. Quality
The four remainder optimizations are rounding-class, so this pull request ships on the quality gate, not on bit-identity. The two aux optimizations are exact and do not affect it. The gate is the repo's own code-eval, HumanEval and MBPP, served, paired release against this pull request. Every cell runs at the benchmark's own sampler: temperature 1, top-p 0.95, top-k 20, reasoning effort
xhigh, one seed (20260829). That is the setting the speed numbers were measured at, so the quality cells answer the question the speed cells raise.Each cell reports three numbers per arm: strict pass@1 counts a task cut off at the output cap as a failure; completed-task pass@1 excludes the cut-off tasks; and the truncation rate carries the mean and max completion tokens. With reasoning
xhighat temperature 1 the release cut off about 10% of HumanEval tasks at an 8,192-token cap while still thinking, so the strict number measured verbosity, not correctness; the cap is now 32,768 tokens. Truncation is reported in its own column and never re-run, so a non-zero rate is data rather than a failed cell; the strict and completed-task columns bracket what it costs.Provenance of the candidate rows: this pull request's own tree was not re-run for quality. The candidate cells are the typical-acceptance sweep's exact cell, served from
264e0835with the typical rule switched off. That tree is this pull request's code plus the typical rule, so with the rule off it executes this pull request's optimizations and nothing else. Release MBPP at this sampler was never run, and the programme was cut after the exact arm, so that cell reads "not run"; the greedy release MBPP figure is in the supplementary row of Section 1.7.Scoring: prompt + solution + tests, solution = last fenced block defining the entry point (evalplus sanitize). Re-scored offline from saved completions after the live gate was found to drop the prompt's helper definitions (HumanEval/38, /50).
Each optimization carries a
=0opt-out, the kill switch if an optimization moves quality.5. Changes that did not work
These candidates were measured and rejected on the earlier #391 stack. Their code is not in this pull request. The numbers are from that stack, at the 17,408-token shape.
The full rejected inventory with per-row numbers is in the perf report under
docs/perf/.6. How to run and how to disable
Build the venv and the native extensions with the setup script, then serve the pack. The server arms all six optimizations by default:
scripts/fable/setup_over100_venv.sh mtplx serve \ --model ~/.mtplx/models/Youssofal--Qwen3.8-Flash-Next-MTPLX-Optimized-Speed \ --model-id mtplx-flash-next-optimized-speedRead
GET /healthand confirm the install verdicts name the optimizations. For the QSA sparse decode optimization, confirm the[mtplx] qsa_sparse_decode armed:line; a serve that printeddeclined to stockran the stock path.Turn one optimization off with its key set to
0:The two native optimizations need their extensions built in the serve environment:
mtplx_native_qsafor QSA sparse decode andple_cpu_rowsfor cached PLE. The setup script builds both. Without an extension its optimization declines to the stock path and still serves.7. File map and provenance
mtplx/models/qwen4_exp.py,mtplx/runtime_options.py,mtplx/qwen4_prefill_chunk.py,mtplx/runtime.py,mtplx/ple_cached_aux.py,mtplx/ple_cached_row_handoff.py,mtplx/qsa_pooled_rowsel.py,mtplx/qwen4_aux_lanes.pymtplx/kernels/qwen4_m4_hyper_read.py,mtplx/kernels/qsa_sparse_decode.py,mtplx/native/__init__.pynative_extensions/qsa_sparse_gqa/(mtplx_native_qsa),native_extensions/ple_cpu_rows/(mtplx_native_ple_cpu_rows)mtplx/server/openai.py,mtplx/profiles.pyscripts/bundle_native_runtime_wheel.pytests/test_qwen4_hc_m4.py,tests/test_qwen4_prefill_mask_fuse.py,tests/test_qsa_sparse_decode.py,tests/test_qsa_sparse_decode_wiring.py,tests/test_qsa_sparse_gqa_native.py,tests/test_pr391_ple_cached_aux_cpu.py,tests/test_pr391_ple_cached_row_handoff_cpu.py,tests/test_pr391_fixed_m4_pool_install_cpu.py,tests/test_qwen4_aux_lanes.py,tests/test_bundle_native_runtime_wheel.pydocs/perf/21be78b3(MTPLX v2.11.2).perf/qwen38-aux-lanes-main@72f58d41= upstream main, then four remainder commits (hyper-connection read, prefill mask fuse, QSA prefill query tile, QSA split-K sparse decode), then one aux commit (cached async PLE rows, pooled QSA row selection).a5e38bb7.History. On the PR 391 campaign at 16,384 tokens the release default reached 71.17 tok/s, the full PR 391 stack reached 80.92 tok/s, and upstream 2.10.2 reached 57.65 tok/s on the same instrument. The nine keys 2.11.2 re-landed recovered about 60% of the full-stack gain; these six optimizations carry the rest and the two exact decode optimizations. On the earlier #391 stack the two aux optimizations together gave +1.59% decode over the same build with both off, byte-identical, in a same-build pair at 16,384 tokens. Those numbers are superseded by the two-arm battery above, which is measured on 2.11.2.