Repository navigation
Conversation
…tZuijdwijk b10641 controller Per-seq acceptance EMA with censoring-aware probe sizes each draft from measured acceptance (n_min..n_max) instead of a fixed n_max. Wired into dflash and mtp draft loops. Greedy-lossless: spec-adaptive vs no-spec bit-identical at temp 0. Measured: +23% counting, +7% JSON, par on decaying content (ledger 2026-08-26/27). Co-authored-by: LaurentZuijdwijk (controller design, fork b10641)
|
First end-to-end agent-build runs on this branch — receipts from a 4-cell N-test (Qwen3.8-27B UD-Q4_K_XL + DFlash2 Q4_K_M,
Happy to run any cell you'd want to see before merge. |
|
Thanks for the port and the careful validation. I'm not going to merge this here, for two reasons. First, development is moving to the halo-box community fork (https://github.com/halo-box/strix-llama.cpp), which has already merged this controller. Second, when we tested it there we found a problem with several requests in flight ( |
PR: spec: adaptive draft sizing (--spec-draft-adaptive) - port of LaurentZuijdwijk's controller
Controller design and constants are LaurentZuijdwijk/llama.cpp b10641 - full credit for the idea. The censoring-aware EMA (full accepts probe additively, partial accepts average the exact stopping point), the constants (alpha 0.25 / probe 1.0 / init 2.0), the n_min semantics, and the content-class framing are all from his fork. This PR is the port onto strix-halo-vulkan (rebased on current tip df1671a, one commit, ~75 lines in
common/only, no backend changes) plus fresh validation on the rebased base.Mechanism
Per-seq EMA of accepted-tokens-per-draft, censoring-aware:
Each draft is sized
clamp(round(EMA), n_min, n_max)instead of always n_max. Defaults off; two new flags (--spec-draft-adaptive,--spec-draft-n-min), wired into the dflash and mtp draft loops.Validation on the REBASED base (all fresh, this week)
PPL (spec-independent): wikitext-2, 580x512, Qwen3.8-27B Q5_K_XL-v2:
7.0801 +/- 0.046267.0838 +/- 0.04628Greedy determinism (honest revision of the earlier "bit-identical" claim):
get_prev_tokensport landed in between) - reported as superseded rather than silently dropped.Throughput (content-class ratchet, from issue #11, same-binary A/B, every cell measured):
Not a universal win - it drafts what the acceptance histogram supports. Default recommendation stays fixed n4 for prose-heavy work; adaptive for agent/emission-heavy sessions.
In-harness (issue #11): three full game-build runs (low/med/high effort) with adaptive on, try-1 success each under a runtime smoke gate, decode aggregate 18-26 t/s.
Relation to adaptive-mtp-port
Complementary, not competing: that branch is Stew's independent controller (hysteresis climb-counter with a rise-cost table, MTP-scoped, with tests); this one is the EMA design, draft-type-agnostic (dflash path included - the Qwen agent stack here runs dflash). Both independently land on the same content-class behavior; happy to run a controller-vs-controller A/B on the same binary if useful, or to rebase the dflash wiring onto that branch instead of carrying two controllers.
Honest limits