…a, every number measured (#233)
The loop ROCm Hyperloom (AMD-AGI) runs on Instinct GPUs, applied to how the engine
already works: the harness measures, a setting is kept with its measurement, and
the next model starts from what is known.
- config/recipes.json + app/recipes.{h,cpp}: tuned llama-server settings per
model and route, each with why / measured / source. `1bit serve` applies the
ones that match (device, architecture, MoE, Hadamard stamp, drafter), never
over a flag it already set, and prints each one; --recipes FILE replaces the
built-in set, --no-recipes turns them off, a malformed file stops serve with
the reason. The two tuned settings that were code move in:
rotated-moe-ub1024 (#205, #229) and dflash-p-min-0 (#194). What follows
from the files themselves (Hadamard env, DFlash draft length) stays in code.
- tools/bench.py: A/B-measures `1bit serve` configurations, interleaved rounds,
the first config as the baseline in the same run, an optional box flock,
best/median, JSON record. Qwen3.8-27B-H32, 2 rounds: base 494 tok/s prompt,
13.1 / 13.1 / 13.5 decode; --dflash 467 (-5.6%), 40.7 / 26.4 / 14.7.
- tools/kernelforge/gdn-prefill: the delta-net prefill kernel as a KernelForge
task for gfx1151 (standalone kernel, float64 oracle: 139 dB, driver with
KernelForge's contract, program.md on what differs from Instinct, run.sh).
- tests: recipes_route (10 checks) and bench_selftest (4) run without a GPU;
fake_backend.py reports llama-server timings.
- docs/recipes.md, docs/bench.md; serve.md, README, site nav point to them.
Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This PR gets the one-server prompt speed past 451 t/s while the DFlash2 drafter is attached.
Measured with Qwen3.8-27B-H32 through
1bit serve. The prompt is 1,838 tokens (best / median of 5). Decode is code / prose / short, 256 tokens greedy, best of 3. The box was quiet and I held the shared box lock.--dflash, before (#194)--dflash, this PRChanges
serve --dflashreads the drafter'sdflash.block_sizeand drafts block − 1 tokens. If the file has no block size, it falls back to 16. The draft length also sets the rollback depth, so asking for 16 against a block of 8 wrote twice as many snapshots as needed.-ub 1024from serve: Hadamard-rotated files run 1024-token micro-batches #205 applies to MoE files only. A rotated file counts as MoE when its<arch>.expert_countis above 0. Dense files keep 512: the 27B serves 491-495 t/s at 512 and 471-473 at 1024.tests/hadamard_route.shadds dense vs MoE micro-batch cases and a block-size case; 13/13 pass.docs/lean.md,docs/serve.md, the README status and the site description are updated with these numbers.Checks
test-backend-ops -o GATED_DELTA_NET: 54/54 on gfx1151.🤖 Generated with Claude Code