Skip to content

spec : add DFlash2 support (local convolution + candidate selector) (#27342) - #27816

Merged
ngxson merged 2 commits into
masterfrom
xsn/dflash2
Aug 27, 2026
Merged

ngxson merged 2 commits into
masterfrom
xsn/dflash2

Conversation

@ngxson

@ngxson ngxson commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

I cannot push to #27342 because the PR was created from an org, applying @ORippler 's request to revert top-k.cu here

for full changes, please refer to #27342

SubSir and others added 2 commits August 27, 2026 19:05
…27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------

Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>
@github-actions github-actions Bot added model Model specific testing Everything test related conversion labels Aug 27, 2026
@ngxson
ngxson marked this pull request as ready for review August 27, 2026 17:16
@ngxson
ngxson requested review from a team, CISC and ggerganov as code owners August 27, 2026 17:16
@ngxson
ngxson merged commit b10f9ca into master Aug 27, 2026
25 of 28 checks passed
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
…gml-org#27342) (ggml-org#27816)

* spec : add DFlash2 support (local convolution + candidate selector) (ggml-org#27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------

Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

* revert top-k.cu changes

---------

Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
wanghqc pushed a commit to qualcomm/llama.cpp that referenced this pull request Sep 7, 2026
…gml-org#27342) (ggml-org#27816)

* spec : add DFlash2 support (local convolution + candidate selector) (ggml-org#27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------

Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

* revert top-k.cu changes

---------

Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
(cherry picked from commit b10f9ca)
@anfedoro

anfedoro commented Sep 8, 2026

Copy link
Copy Markdown

senseless until split-mode = tensor is stable.

zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…gml-org#27342) (ggml-org#27816)

* spec : add DFlash2 support (local convolution + candidate selector) (ggml-org#27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------

Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

* revert top-k.cu changes

---------

Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…gml-org#27342) (ggml-org#27816)

* spec : add DFlash2 support (local convolution + candidate selector) (ggml-org#27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------

Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

* revert top-k.cu changes

---------

Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
…gml-org#27342) (ggml-org#27816)

* spec : add DFlash2 support (local convolution + candidate selector) (ggml-org#27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------

Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

* revert top-k.cu changes

---------

Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
bri-prism pushed a commit to PrismML-Eng/llama.cpp that referenced this pull request Sep 25, 2026
…gml-org#27342) (ggml-org#27816) (#261)

* spec : add DFlash2 support (local convolution + candidate selector) (ggml-org#27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------




* revert top-k.cu changes

---------



(cherry picked from commit b10f9ca)

Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>
Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
delneg added a commit to delneg/qwen-dflash that referenced this pull request Sep 27, 2026
DFlash 2 is merged upstream (ggml-org/llama.cpp#27816), so CI now builds
llama-server from ggml-org/llama.cpp at b11213 instead of the z-lab fork.
That picks up the fused DFlash encoder/KV injection (#27310), which removes
the prompt-processing slowdown the fork had with speculation on (M1 Max,
4.4k-token prompt: 109 tok/s with DFlash vs 88 without; the fork was ~2x
slower with DFlash).

start.sh now runs a single slot (-np 1) with a RAM prompt cache, so an
agent's long prompt is reused between turns instead of re-read, and sizes
context to the machine's RAM. QUANT=Q8_0 selects the 8-bit target and draft.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
bong-water-water-bong added a commit to 1bit-MONSTER/ROCmFPX that referenced this pull request Sep 28, 2026
… for large rows (#2)

Ports upstream DFlash2 support (b10f9ca, "spec : add DFlash2 support (local
convolution + candidate selector)") onto this fork's DFlash code:

- dflash.cpp: DFlash2 hparams (block/conv/selector keys), per-layer local
  convolutions around attention and FFN, the target's logit scale/softcap on
  the draft logits, and the selector lattice (top-k candidates + pairwise
  transition scores per block position) written to the pre-norm output
  (this fork's name for upstream's h_nextn)
- speculative.cpp: the greedy lattice walk (p_min on the softmax of the
  winning score, n_min); block tokens request no logits for DFlash2
- llama-context: DFlash2 graph node budget
- ggml-cuda: upstream's HIP radix top-k, so TOP_K over a 248k vocab stays on
  the GPU instead of falling back to the CPU (test-backend-ops TOP_K 445/445
  on gfx1151, rows up to 524288)
- common: keep bounded recurrent-state rollback (n_rs_seq) for draft-dflash,
  as upstream does; without it every partial accept restores a checkpoint and
  replays the accepted tokens

M-RoPE handling from the upstream commit is not needed here: the Qwen3.8
drafter has one 64-wide section, identical to NEOX with its positions.

Qwen3.8-27B (Hadamard-rotated Q4_0 + W4A4 prompt path) + Qwen3.8-27B-DFlash2
q8_0 on gfx1151, -ctxcp 0, --spec-draft-p-min 0, 256 tokens greedy: mean
acceptance length 6.71 code / 4.11 prose, matching upstream (6.71 / 4.25).
Greedy output identical to no-draft on code and short prompts; prose flips
one near-tie word at char 470 (batched verification rounding).

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion model Model specific testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants