Skip to content

Recipes, tools/bench.py and a KernelForge task: tuned settings as data, every number measured - #233

Merged
bong-water-water-bong merged 1 commit into
mainfrom
recipes
Sep 29, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
recipes

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

This PR brings over what fits from ROCm Hyperloom, AMD's agent harness for Instinct GPUs: the harness runs the benchmarks, each kept setting carries its measurement, and each run starts from a fresh baseline.

Recipes: tuned settings as data

  • What a recipe is. config/recipes.json holds tuned llama-server settings per model and route. Each entry says what it matches, what it adds, why, and the measurement behind it. A recipe without a measurement does not load.
  • How serve applies them. 1bit serve compiles the file in and applies every recipe that matches the launch: device, architecture, MoE or not, Hadamard stamp, drafter. It never adds a flag the command line already has, and it prints what it added:
    1bit serve: recipe rotated-moe-ub1024: -ub 1024
    
  • Overrides. --recipes FILE replaces the built-in set and --no-recipes turns recipes off. A malformed file stops serve and names the problem, for example unknown match key "archtecture".
  • What moved into recipes. Two tuned settings that were in code: rotated-moe-ub1024 (serve: Hadamard-rotated files run 1024-token micro-batches #205, One server: the prompt keeps 95% of its speed with the DFlash2 drafter (445 -> 477 t/s) #229) and dflash-p-min-0 (One server: W4A4 prompts and DFlash2 decode on the lean ROCm route #194). Settings that follow from the files themselves stay in code: the Hadamard environment and the DFlash draft length.

tools/bench.py: the harness runs the measurements

It A/B-measures 1bit serve configurations. The runs are interleaved round by round, and the first config is the baseline, measured in the same run. It can hold a box lock, reports best and median, and writes a JSON record. For example, 2 rounds on strixhalo, Qwen3.8-27B-H32:

config prompt decode code decode prose decode short
base 501.6 / 494.3 13.2 / 13.1 13.2 / 13.1 13.6 / 13.5
dflash 470.3 / 466.8 (-5.6%) 40.9 / 40.7 (+210.6%) 26.7 / 26.4 (+100.9%) 15.0 / 14.7 (+8.4%)

tools/kernelforge/gdn-prefill: a KernelForge task for gfx1151

This packages the delta-net prefill kernel for KernelForge. It is our kernel, and about 9% of the 27B's prompt time. The task has:

  • a standalone copy of the kernel;
  • a float64 oracle, which the current kernel matches at 139 dB SNR;
  • a driver that follows KernelForge's contract;
  • a program.md on what differs from Instinct (wave32, WMMA not MFMA, no MFMA counters);
  • run.sh to launch a campaign.

It isn't run yet. KernelForge drives its agent through a logged-in claude CLI, billed to that account, and strixhalo has none.

Tests

  • recipes_route (10 checks) and bench_selftest (4) run without a GPU.
  • The existing hadamard_route still passes 13/13 with the settings now coming from recipes.
  • The whole no-GPU ctest set passes 12/12.

🤖 Generated with Claude Code

…a, every number measured

The loop ROCm Hyperloom (AMD-AGI) runs on Instinct GPUs, applied to how the engine
already works: the harness measures, a setting is kept with its measurement, and
the next model starts from what is known.

- config/recipes.json + app/recipes.{h,cpp}: tuned llama-server settings per
  model and route, each with why / measured / source. `1bit serve` applies the
  ones that match (device, architecture, MoE, Hadamard stamp, drafter), never
  over a flag it already set, and prints each one; --recipes FILE replaces the
  built-in set, --no-recipes turns them off, a malformed file stops serve with
  the reason. The two tuned settings that were code move in:
  rotated-moe-ub1024 (#205, #229) and dflash-p-min-0 (#194). What follows
  from the files themselves (Hadamard env, DFlash draft length) stays in code.
- tools/bench.py: A/B-measures `1bit serve` configurations, interleaved rounds,
  the first config as the baseline in the same run, an optional box flock,
  best/median, JSON record. Qwen3.8-27B-H32, 2 rounds: base 494 tok/s prompt,
  13.1 / 13.1 / 13.5 decode; --dflash 467 (-5.6%), 40.7 / 26.4 / 14.7.
- tools/kernelforge/gdn-prefill: the delta-net prefill kernel as a KernelForge
  task for gfx1151 (standalone kernel, float64 oracle: 139 dB, driver with
  KernelForge's contract, program.md on what differs from Instinct, run.sh).
- tests: recipes_route (10 checks) and bench_selftest (4) run without a GPU;
  fake_backend.py reports llama-server timings.
- docs/recipes.md, docs/bench.md; serve.md, README, site nav point to them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 29, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 5f02718

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

205 - Partially compliant

Compliant requirements:

  • W4A4 prefill +6-10% for free on llama-bench
  • Through 1bit serve --long-model, the 16K first token comes at 12.96 s instead of 13.89 s
  • tests/hadamard_route.sh checks the flag
  • flash attention off drops pp16384 to 337, so llama-server's default -fa auto (on) must stay

Non-compliant requirements:

  • None

Requires further human verification:

  • None

229 - Partially compliant

Compliant requirements:

  • This PR gets the one-server prompt speed past 451 t/s while the DFlash2 drafter is attached
  • Measured with Qwen3.8-27B-H32 through 1bit serve
  • Prompt is 1,838 tokens (best / median of 5)
  • Decode is code / prose / short, 256 tokens greedy, best of 3
  • The box was quiet and I held the shared box lock
  • Prompt (t/s) No drafter: 497 / 495, --dflash, before (One server: W4A4 prompts and DFlash2 decode on the lean ROCm route #194): 445 / 439, --dflash, this PR: 477 / 472
  • Decode (tok/s) No drafter: 13.2 / 13.2 / 13.6, --dflash, before (One server: W4A4 prompts and DFlash2 decode on the lean ROCm route #194): 40.9 / 26.8 / 13.7, --dflash, this PR: 41.1 / 26.7 / 15.0
  • ROCmFPX pin moves to b82c8b8 (ROCmFPX#4)
  • Two fixes: chunked delta-net prefill kernel now writes the state snapshots that speculative rollback keeps
  • Layer inputs are now written at their micro-batch's offset
  • Draft length follows the drafter
  • -ub 1024 from serve: Hadamard-rotated files run 1024-token micro-batches #205 applies to MoE files only
  • Tests: tests/hadamard_route.sh adds dense vs MoE micro-batch cases and a block-size case; 13/13 pass
  • Docs: docs/lean.md, docs/serve.md, the README status and the site description are updated with these numbers

Non-compliant requirements:

  • None

Requires further human verification:

  • None

194 - Partially compliant

Compliant requirements:

  • With this PR, one 1bit serve runs Qwen3.8-27B with Hadamard W4A4 prompt processing and DFlash2 decode
  • 1bit serve -m Qwen3.8-27B-Q4_0-H32.gguf --dflash Qwen3.8-27B-DFlash2-q8_0.gguf
  • Pin bump: third_party/llama.cpp-rocmfpx moves d572668 → 8abd563 (ROCmFPX#2)
  • That brings upstream's DFlash2 drafters (spec : add DFlash2 support (local convolution + candidate selector) (#27342) ggml-org/llama.cpp#27816), ported to the ROCm tree
  • A HIP top-k for large rows (TOP_K 445/445)
  • Bounded recurrent-state rollback for draft-dflash
  • app/serve.cpp: --dflash now passes --spec-draft-p-min 0 unless --mtp-p-min is given
  • The ROCm tree defaults to 0.75, which cuts DFlash2 blocks from 6.7 to 5.4 tokens per step
  • Upstream's default is already 0, so the Vulkan route doesn't change
  • tests/hadamard_route.sh: new checks that --dflash on a rotated file stays on ROCm0 with the drafter, n-max 16, p-min 0 and the W4A4 environment (10/10 pass)
  • Docs: docs/lean.md gets a one-server section with measurements; docs/serve.md gets a pointer to it
  • Measured: Strix Halo, through 1bit serve, 2026-09-28, quiet box
  • Prompt is the 1,838-token prompt (best / median of 5)
  • Decode is code / prose / short, 256 tokens greedy, best of 3
  • Prompt t/s No drafter: 502 / 500, --dflash: 445 / 439
  • Acceptance: 6.71 tokens per step on code and 4.11 on prose, matching upstream
  • Greedy output vs no drafter: identical on code and short prompts. On prose, one near-tie word flips at character 470
  • Prompt cost of the drafter: about 0.5 s on this prompt

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 4 🔵🔵🔵🔵⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Recipe Parsing Error Handling

The recipe parsing logic in Recipes::parse throws a std::runtime_error with a descriptive message when a recipe lacks a "measured" field, but it does not handle the case where the "measured" field is present but empty. This could lead to a recipe being silently ignored if the measurement is an empty string, which might be a valid state in some edge cases.

if (rec.measured.empty()) throw std::runtime_error("recipe " + rec.id + " has no \"measured\"");
Recipe Application Order

The order in which recipes are applied in launch_for function is not explicitly defined. If multiple recipes match, their application order could affect the final command line arguments and environment variables, especially if there are conflicts between recipes. This could lead to inconsistent behavior if the order of application is not deterministic.

for (const auto& line : recipes.apply(facts, argv, env))
Timing Measurement Consistency

The bench.py tool uses fixed timing values for prompt and predicted per second in the fake backend. While this is acceptable for testing, it could mask inconsistencies in actual timing measurements if the real backend does not behave exactly as expected. This could lead to misleading performance comparisons.

body = {"messages": [{"role": "user", "content": content}], "temperature": 0, "max_tokens": max_tokens,
        "cache_prompt": False, "chat_template_kwargs": {"enable_thinking": False}}
req = urllib.request.Request(f"http://127.0.0.1:{port}/v1/chat/completions", json.dumps(body).encode(),

@bong-water-water-bong
bong-water-water-bong merged commit 46fdaca into main Sep 29, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the recipes branch September 29, 2026 18:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant