…a, every number measured
The loop ROCm Hyperloom (AMD-AGI) runs on Instinct GPUs, applied to how the engine
already works: the harness measures, a setting is kept with its measurement, and
the next model starts from what is known.
- config/recipes.json + app/recipes.{h,cpp}: tuned llama-server settings per
model and route, each with why / measured / source. `1bit serve` applies the
ones that match (device, architecture, MoE, Hadamard stamp, drafter), never
over a flag it already set, and prints each one; --recipes FILE replaces the
built-in set, --no-recipes turns them off, a malformed file stops serve with
the reason. The two tuned settings that were code move in:
rotated-moe-ub1024 (#205, #229) and dflash-p-min-0 (#194). What follows
from the files themselves (Hadamard env, DFlash draft length) stays in code.
- tools/bench.py: A/B-measures `1bit serve` configurations, interleaved rounds,
the first config as the baseline in the same run, an optional box flock,
best/median, JSON record. Qwen3.8-27B-H32, 2 rounds: base 494 tok/s prompt,
13.1 / 13.1 / 13.5 decode; --dflash 467 (-5.6%), 40.7 / 26.4 / 14.7.
- tools/kernelforge/gdn-prefill: the delta-net prefill kernel as a KernelForge
task for gfx1151 (standalone kernel, float64 oracle: 139 dB, driver with
KernelForge's contract, program.md on what differs from Instinct, run.sh).
- tests: recipes_route (10 checks) and bench_selftest (4) run without a GPU;
fake_backend.py reports llama-server timings.
- docs/recipes.md, docs/bench.md; serve.md, README, site nav point to them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This PR brings over what fits from ROCm Hyperloom, AMD's agent harness for Instinct GPUs: the harness runs the benchmarks, each kept setting carries its measurement, and each run starts from a fresh baseline.
Recipes: tuned settings as data
config/recipes.jsonholds tuned llama-server settings per model and route. Each entry says what it matches, what it adds, why, and the measurement behind it. A recipe without a measurement does not load.1bit servecompiles the file in and applies every recipe that matches the launch: device, architecture, MoE or not, Hadamard stamp, drafter. It never adds a flag the command line already has, and it prints what it added:--recipes FILEreplaces the built-in set and--no-recipesturns recipes off. A malformed file stops serve and names the problem, for exampleunknown match key "archtecture".rotated-moe-ub1024(serve: Hadamard-rotated files run 1024-token micro-batches #205, One server: the prompt keeps 95% of its speed with the DFlash2 drafter (445 -> 477 t/s) #229) anddflash-p-min-0(One server: W4A4 prompts and DFlash2 decode on the lean ROCm route #194). Settings that follow from the files themselves stay in code: the Hadamard environment and the DFlash draft length.tools/bench.py: the harness runs the measurements
It A/B-measures
1bit serveconfigurations. The runs are interleaved round by round, and the first config is the baseline, measured in the same run. It can hold a box lock, reports best and median, and writes a JSON record. For example, 2 rounds on strixhalo, Qwen3.8-27B-H32:tools/kernelforge/gdn-prefill: a KernelForge task for gfx1151
This packages the delta-net prefill kernel for KernelForge. It is our kernel, and about 9% of the 27B's prompt time. The task has:
program.mdon what differs from Instinct (wave32, WMMA not MFMA, no MFMA counters);run.shto launch a campaign.It isn't run yet. KernelForge drives its agent through a logged-in
claudeCLI, billed to that account, and strixhalo has none.Tests
recipes_route(10 checks) andbench_selftest(4) run without a GPU.hadamard_routestill passes 13/13 with the settings now coming from recipes.🤖 Generated with Claude Code