Skip to content

One server: W4A4 prompts and DFlash2 decode on the lean ROCm route - #194

Merged
bong-water-water-bong merged 1 commit into
mainfrom
one-server
Sep 28, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
one-server

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

With this PR, one 1bit serve runs Qwen3.8-27B with Hadamard W4A4 prompt processing and DFlash2 decode:

1bit serve -m Qwen3.8-27B-Q4_0-H32.gguf --dflash Qwen3.8-27B-DFlash2-q8_0.gguf

Changes

  • Pin bump: third_party/llama.cpp-rocmfpx moves d572668 → 8abd563 (ROCmFPX#2). That brings:
  • app/serve.cpp: --dflash now passes --spec-draft-p-min 0 unless --mtp-p-min is given. The ROCm tree defaults to 0.75, which cuts DFlash2 blocks from 6.7 to 5.4 tokens per step. Upstream's default is already 0, so the Vulkan route doesn't change.
  • tests/hadamard_route.sh: new checks that --dflash on a rotated file stays on ROCm0 with the drafter, n-max 16, p-min 0 and the W4A4 environment (10/10 pass).
  • Docs: docs/lean.md gets a one-server section with measurements; docs/serve.md gets a pointer to it.

Measured

Strix Halo, through 1bit serve, 2026-09-28, quiet box. Prompt is the 1,838-token prompt (best / median of 5). Decode is code / prose / short, 256 tokens greedy, best of 3.

Prompt t/s Decode tok/s
No drafter 502 / 500 13.0 / 13.0 / 13.4
--dflash 445 / 439 40.9 / 26.8 / 13.7
  • Acceptance: 6.71 tokens per step on code and 4.11 on prose, matching upstream.
  • Greedy output vs no drafter: identical on code and short prompts. On prose, one near-tie word flips at character 470.
  • Prompt cost of the drafter: about 0.5 s on this prompt. That is the drafter's encoder pass (0.13 s) plus larger context checkpoints. It is the next thing to work on.

🤖 Generated with Claude Code

Bumps third_party/llama.cpp-rocmfpx to 8abd563 (ROCmFPX#2): upstream's DFlash2
drafters ported onto the ROCm tree, a HIP top-k for the drafter's
vocabulary-wide candidate pick, and bounded recurrent-state rollback for
draft-dflash.

`1bit serve -m Qwen3.8-27B-Q4_0-H32.gguf --dflash <drafter>` now runs the
Hadamard W4A4 prompt path and DFlash2 decode in one llama-server. `--dflash`
passes --spec-draft-p-min 0 unless --mtp-p-min is given (the ROCm tree's
default of 0.75 cut blocks from 6.7 to 5.4 tokens a step).
tests/hadamard_route.sh checks the argv and environment of that route.

Strix Halo, through `1bit serve`, 1,838-token prompt / 256-token greedy
decode (code / prose / short): no drafter 502 t/s, 13.0 / 13.0 / 13.4 tok/s;
--dflash 445 t/s, 40.9 / 26.8 / 13.7 tok/s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 28, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 33227c8

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

2 - Partially compliant

Compliant requirements:

  • Tokenizer implementation in C++ without Python or ICU dependencies
  • Reads vocabulary from GGUF metadata
  • Exact matching against HF tokenizers 0.22.2 for pre-token pieces, ids, and decoded bytes
  • Implementation of byte-level BPE from GGUF
  • NFC normalization and Qwen2 split regex
  • GPT-2 byte-level mapping and BPE merges
  • Decode functionality
  • UTF-8 codec, character classes and NFC from Python unicodedata 16.0.0
  • Documentation in docs/tokenizer.md explaining design, verification, and known limits

Non-compliant requirements:

  • Verification with 481 committed cases and 20,081 random cases with 0 mismatches (this is a test requirement, not a code implementation requirement)

Requires further human verification:

  • None
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Default p-min for DFlash

The change introduces a default --spec-draft-p-min 0 for DFlash models when no explicit --mtp-p-min is provided. This is important for ensuring consistent behavior across different backends, particularly ROCm where the default is 0.75, which would otherwise reduce block size from 6.7 to 5.4 tokens per step. The change ensures that DFlash2 blocks are kept whole unless explicitly overridden.

// a DFlash block is kept whole unless --mtp-p-min says otherwise: upstream's default p-min
// is 0, the ROCm tree's is 0.75, which cuts DFlash2 blocks from 6.7 to 5.4 tokens a step
else if (!o.dflash.empty()) argv.insert(argv.end(), {"--spec-draft-p-min", "0"});
New test case for DFlash

A new test case has been added to verify that when using --dflash on a rotated file, the system correctly routes to ROCm0 with the DFlash drafter, full blocks (n-max 16), and p-min 0. This ensures that the intended behavior for the DFlash functionality is properly tested.

run "$scratch/h32.gguf" "$scratch/df.json" --dflash "$scratch/draft.gguf"
after() { field "$scratch/df.json" "r[\"argv\"][r[\"argv\"].index(\"$1\")+1]"; }
check "--dflash on a rotated file stays on ROCm0" '[ "$(after --device)" = ROCm0 ]'
check "  with the DFlash drafter" '[ "$(after --spec-type)" = draft-dflash ] && [ "$(after -md)" = "$scratch/draft.gguf" ]'
check "  full blocks: n-max 16, p-min 0" '[ "$(after --spec-draft-n-max)" = 16 ] && [ "$(after --spec-draft-p-min)" = 0 ]'
check "  and the Hadamard W4A4 environment" '[ "$(field "$scratch/df.json" "r[\"env\"].get(\"GGML_W4A4_TENSORS\")")" = all ]'

@bong-water-water-bong
bong-water-water-bong merged commit 3a9bc89 into main Sep 28, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the one-server branch September 28, 2026 14:28
bong-water-water-bong added a commit that referenced this pull request Sep 28, 2026
… (#195)

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 29, 2026
…a, every number measured (#233)

The loop ROCm Hyperloom (AMD-AGI) runs on Instinct GPUs, applied to how the engine
already works: the harness measures, a setting is kept with its measurement, and
the next model starts from what is known.

- config/recipes.json + app/recipes.{h,cpp}: tuned llama-server settings per
  model and route, each with why / measured / source. `1bit serve` applies the
  ones that match (device, architecture, MoE, Hadamard stamp, drafter), never
  over a flag it already set, and prints each one; --recipes FILE replaces the
  built-in set, --no-recipes turns them off, a malformed file stops serve with
  the reason. The two tuned settings that were code move in:
  rotated-moe-ub1024 (#205, #229) and dflash-p-min-0 (#194). What follows
  from the files themselves (Hadamard env, DFlash draft length) stays in code.
- tools/bench.py: A/B-measures `1bit serve` configurations, interleaved rounds,
  the first config as the baseline in the same run, an optional box flock,
  best/median, JSON record. Qwen3.8-27B-H32, 2 rounds: base 494 tok/s prompt,
  13.1 / 13.1 / 13.5 decode; --dflash 467 (-5.6%), 40.7 / 26.4 / 14.7.
- tools/kernelforge/gdn-prefill: the delta-net prefill kernel as a KernelForge
  task for gfx1151 (standalone kernel, float64 oracle: 139 dB, driver with
  KernelForge's contract, program.md on what differs from Instinct, run.sh).
- tests: recipes_route (10 checks) and bench_selftest (4) run without a GPU;
  fake_backend.py reports llama-server timings.
- docs/recipes.md, docs/bench.md; serve.md, README, site nav point to them.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant