Skip to content

One server: the prompt keeps 95% of its speed with the DFlash2 drafter (445 -> 477 t/s) - #229

Merged
bong-water-water-bong merged 1 commit into
mainfrom
dflash-block
Sep 29, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
dflash-block

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

This PR gets the one-server prompt speed past 451 t/s while the DFlash2 drafter is attached.

Measured with Qwen3.8-27B-H32 through 1bit serve. The prompt is 1,838 tokens (best / median of 5). Decode is code / prose / short, 256 tokens greedy, best of 3. The box was quiet and I held the shared box lock.

Prompt (t/s) Decode (tok/s)
No drafter 497 / 495 13.2 / 13.2 / 13.6
--dflash, before (#194) 445 / 439 40.9 / 26.8 / 13.7
--dflash, this PR 477 / 472 41.1 / 26.7 / 15.0

Changes

  • ROCmFPX pin moves to b82c8b8 (ROCmFPX#4). Two fixes:
    • The chunked delta-net prefill kernel now writes the state snapshots that speculative rollback keeps. Before, any drafter pushed the prompt onto the slower per-token kernel: 2,966 ms vs 2,588 ms for 1,322 tokens. With the fix it's 2,762 ms.
    • Layer inputs are now written at their micro-batch's offset. Before, prompts longer than one micro-batch gave the DFlash (and EAGLE3) encoder stale hidden states.
  • Draft length follows the drafter. serve --dflash reads the drafter's dflash.block_size and drafts block − 1 tokens. If the file has no block size, it falls back to 16. The draft length also sets the rollback depth, so asking for 16 against a block of 8 wrote twice as many snapshots as needed.
  • -ub 1024 from serve: Hadamard-rotated files run 1024-token micro-batches #205 applies to MoE files only. A rotated file counts as MoE when its <arch>.expert_count is above 0. Dense files keep 512: the 27B serves 491-495 t/s at 512 and 471-473 at 1024.
  • Tests: tests/hadamard_route.sh adds dense vs MoE micro-batch cases and a block-size case; 13/13 pass.
  • Docs: docs/lean.md, docs/serve.md, the README status and the site description are updated with these numbers.

Checks

  • Acceptance: 6.54 tokens per step on code, 4.23 on prose (upstream 6.71 / 4.25).
  • Greedy output vs no drafter: identical on code and short prompts. On prose, one near-tie phrase differs.
  • test-backend-ops -o GATED_DELTA_NET: 54/54 on gfx1151.

🤖 Generated with Claude Code

Qwen3.8-27B-H32 through `1bit serve --dflash`, 1,838-token prompt: 445 ->
477 tok/s best (439 -> 472 median; 497 without a drafter), decode unchanged
(41.1 / 26.7 / 15.0 tok/s code / prose / short).

- ROCmFPX pin -> b82c8b8 (ROCmFPX#4): the chunked delta-net prefill kernel
  writes the rollback snapshots speculative decoding keeps, so a drafter no
  longer drops prompts onto the per-token kernel; layer inputs land at their
  micro-batch's offset, so the DFlash encoder reads the right hidden states on
  prompts longer than one micro-batch.
- serve --dflash drafts the drafter's dflash.block_size - 1 tokens (16 when the
  file does not say) instead of 16. That is also the rollback depth: 16
  against DFlash2's block of 8 wrote twice the snapshots.
- -ub 1024 (#205) only for rotated MoE files: the dense 27B serves 491-495
  tok/s with 512 and 471-473 with 1024.
- tests/hadamard_route.sh: dense/MoE micro-batch and block-size cases (13/13).
- docs/lean.md, docs/serve.md, README status, site description.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 29, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit b082638

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

194 - Partially compliant

Compliant requirements:

  • Pin bump to ROCmFPX#2 with DFlash2 drafters and HIP top-k
  • app/serve.cpp correctly sets --spec-draft-p-min 0 for --dflash unless --mtp-p-min is given
  • Tests added in tests/hadamard_route.sh verify the behavior
  • Documentation updated in docs/lean.md and docs/serve.md with measurements

Non-compliant requirements:

  • None

Requires further human verification:

  • None

4 - Partially compliant

Compliant requirements:

  • 1bit-server implements Lemonade's backend protocol
  • Launch flags are supported
  • Health endpoint implemented
  • Endpoints /v1/models, /v1/chat/completions, /v1/completions implemented
  • Telemetry with OpenAI usage and llama-server timings
  • Chat with Jinja template and reasoning content routing
  • Stop strings and end-of-turn detection
  • Prompt cache with shared prefix reuse
  • Pinned dependencies with sha256
  • Documentation in docs/server.md

Non-compliant requirements:

  • None

Requires further human verification:

  • None

205 - Partially compliant

Compliant requirements:

  • W4A4 prefill +6-10% for free with 1bit serve --long-model
  • 1024-token micro-batches for MoE files implemented in app/serve.cpp
  • Tests added in tests/hadamard_route.sh to verify the flag

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 3 🔵🔵🔵⚪⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Draft length calculation

The draft length calculation in app/serve.cpp now reads the drafter's dflash.block_size and drafts block - 1 tokens. This is a good change, but it's important to ensure that the fallback to 16 tokens when block_size is not present is handled correctly and consistently across all code paths. The logic should be robust to avoid any potential issues with draft length being set incorrectly.

else if (!o.dflash.empty()) {
    const long long block = gguf_int(o.dflash, gguf_architecture(o.dflash) + ".block_size");
    argv.insert(argv.end(), {"--spec-draft-n-max", block > 1 ? std::to_string(block - 1) : "16"});
}
Micro-batch size for MoE files

The code now sets -ub 1024 for MoE files, which is intended to improve prompt processing speed. However, it's crucial to verify that this change doesn't negatively impact performance on dense models or introduce any regressions. The change should be validated with performance benchmarks on both MoE and dense models.

if (gguf_int(o.model, gguf_architecture(o.model) + ".expert_count") > 0)
    argv.insert(argv.end(), {"-ub", "1024"});

@bong-water-water-bong
bong-water-water-bong merged commit 9f6a0de into main Sep 29, 2026
11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the dflash-block branch September 29, 2026 14:59
bong-water-water-bong added a commit that referenced this pull request Sep 29, 2026
…a, every number measured (#233)

The loop ROCm Hyperloom (AMD-AGI) runs on Instinct GPUs, applied to how the engine
already works: the harness measures, a setting is kept with its measurement, and
the next model starts from what is known.

- config/recipes.json + app/recipes.{h,cpp}: tuned llama-server settings per
  model and route, each with why / measured / source. `1bit serve` applies the
  ones that match (device, architecture, MoE, Hadamard stamp, drafter), never
  over a flag it already set, and prints each one; --recipes FILE replaces the
  built-in set, --no-recipes turns them off, a malformed file stops serve with
  the reason. The two tuned settings that were code move in:
  rotated-moe-ub1024 (#205, #229) and dflash-p-min-0 (#194). What follows
  from the files themselves (Hadamard env, DFlash draft length) stays in code.
- tools/bench.py: A/B-measures `1bit serve` configurations, interleaved rounds,
  the first config as the baseline in the same run, an optional box flock,
  best/median, JSON record. Qwen3.8-27B-H32, 2 rounds: base 494 tok/s prompt,
  13.1 / 13.1 / 13.5 decode; --dflash 467 (-5.6%), 40.7 / 26.4 / 14.7.
- tools/kernelforge/gdn-prefill: the delta-net prefill kernel as a KernelForge
  task for gfx1151 (standalone kernel, float64 oracle: 139 dB, driver with
  KernelForge's contract, program.md on what differs from Instinct, run.sh).
- tests: recipes_route (10 checks) and bench_selftest (4) run without a GPU;
  fake_backend.py reports llama-server timings.
- docs/recipes.md, docs/bench.md; serve.md, README, site nav point to them.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant