Skip to content

DeepseekV4: Add fused hyper-connection ops - #25585

Merged
am17an merged 6 commits into
ggml-org:masterfrom
am17an:hc_ops
Jul 16, 2026
Merged

am17an merged 6 commits into
ggml-org:masterfrom
am17an:hc_ops

Conversation

@am17an

@am17an am17an commented Jul 12, 2026

Copy link
Copy Markdown
Contributor

Overview

Add the sinkhorn ops to the DeepseekV4 graph. Graph nodes go from 29k to 8k with this change. Big increases in the TG + PP

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, CUDA kernels and backend tests are auto-generated, I have reviewed and tested them

@am17an
am17an requested review from a team, CISC, JohannesGaessler and ggerganov as code owners July 12, 2026 11:54
@github-actions github-actions Bot added model Model specific testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Jul 12, 2026
Comment thread ggml/src/ggml-cuda/dsv4-hc.cu Outdated
Comment thread ggml/src/ggml-cuda/dsv4-hc.cu Outdated
Comment thread ggml/src/ggml-cuda/dsv4-hc.cu Outdated
@jadenmach2

Copy link
Copy Markdown
Contributor

@am17an looks like there's an overlap with #25421 . I added the Sinkhorn step as a standalone SINKHORN_NORM op, you've fused the whole HC block, which is the better call.

I came at this from the AMD side. Profiling deepseek-v4-flash on an MI250X, the decomposed Sinkhorn was ~30% of gpu kernel time across ~250k tiny 4×4 dispatches, fusing it gave me +60.9% decode / +7.4% prefill on 4 GCDs.

Since we're solving the same thing, want to join forces? I'm happy to close mine and help on this one, especially validating/tuning it on MI250X if that's not already done/ any other additions you would like to delegate ?

@am17an

am17an commented Jul 13, 2026

Copy link
Copy Markdown
Contributor Author

@jadenmach2 sure, thanks! I've only tested these on nvidia gpus

@ggerganov ggerganov self-assigned this Jul 16, 2026
Comment thread src/models/deepseek4.cpp Outdated
Comment thread src/models/deepseek4.cpp Outdated
Comment thread tests/test-backend-ops.cpp Outdated
Comment thread ggml/include/ggml.h Outdated
@am17an
am17an requested a review from fairydreaming July 16, 2026 15:53
@am17an
am17an merged commit 0dc74e3 into ggml-org:master Jul 16, 2026
26 of 31 checks passed
@ggerganov

Copy link
Copy Markdown
Member

@fairydreaming

Copy link
Copy Markdown
Contributor

@am17an The new ops are producing nans sometimes:

https://github.com/ggml-org/llama.cpp/actions/runs/29572046650/job/87857884185?pr=25816#step:3:1965

This is just my wild guess, but isn't that because of uninitialized sentinel tensors? test_dsv4_hc implements its own initialize_tensors() where it skips initialization if tensor_range() returns false - and it does it based on the tensor name.

@am17an

am17an commented Jul 17, 2026

Copy link
Copy Markdown
Contributor Author

Yeah I think that's the most likely cause. AFK in case @fairydreaming you want to pick this up, otherwise I will fix in a few hours

@fairydreaming

Copy link
Copy Markdown
Contributor

@ggerganov @am17an This should fix it: #25822

gianni-cor pushed a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Jul 25, 2026
* dsv4 hc-ops

* add missing files;

* add cparams

* update rpc version

* address review comments

* address review comments
kmbandy added a commit to kmbandy/llama.cpp that referenced this pull request Jul 31, 2026
Merge upstream ggml-org/llama.cpp master (876a432) into fork master.
Fork point was 2026-07-06 (20a04b2): 330 commits behind, 965 ahead.
37 files conflicted; all resolved by hand (never --theirs), plus fixes for
cleanly-merged files that referenced APIs changed elsewhere.

Brings in DSpark speculative decoding (8407527) and the DeepSeek V4 work
that landed after our fork point, including the fused hyper-connection ops
(0dc74e3, ggml-org#25585), DS4 seq_rm fix, graph-split reduction, MTP tensor
loading, dflash K/V rotation and sidecar auto-download.

Upstream removals that forced fork-side ports:
  * -sm row / CUDA split buffers removed (74976e1). Dropped our copy of the
    split-buffer implementation (no fork code in it) and rewrote
    ggml_cuda_mul_mat onto upstream's early-return dispatch, re-inserting the
    ML8_FP8 route, the F8_E4M3 guard and the TQ4_1S/TQ3_1S kernels.
    ggml_cuda_Memcpy2DPeerAsync is kept -- it carries our no-P2P host staging.
  * mmq.cuh rewritten upstream (<type,J,fallback>, y_scale, stream-k helper).
    Ported the MAD-88 routed-expert pointer hook to the new kernel signature,
    args struct and both launch sites.
  * WMMA flash-attention kernel deleted upstream; our RDNA4 path now leads.
  * use_mmap/use_mlock/use_direct_io collapsed into llama_load_mode. Weight
    paging now strips only the mmap bit instead of forcing no-mmap.
  * llama_context auto-FA/GDN resolution folded into resolve_fused_ops(); the
    WP attention-island guard is re-injected there, scoped to flash-attn.
  * server: draft/MTP context now built by common_speculative_init_from_params,
    so --spec-draft-n-ctx and the MTP tier-disable move onto params_dft; slot
    memory ops go through slot.mem; migrated two subprocess sites to
    common_subproc; folded our byte-backpressure into upstream's server_pipe
    (upstream's drops the oldest item, which would corrupt a proxied body).

Fork-visible behaviour changes:
  * GGML_TYPE_Q2_0 is 56 here, not upstream's 42 -- 42 is our TURBO3_0 and type
    ids are on-disk. Our turbo/ml8 GGUFs stay readable; an upstream Q2_0 GGUF
    needs reconversion. gguf-py kept in sync.
  * DS4 per-layer output renamed l_out -> l_last upstream; added l_last to the
    WP FFN-island pin or it silently stops firing on DeepSeek V4.
  * Dropped our -ffast-math on HIP: upstream sets -funsafe-math-optimizations
    instead because -ffast-math implies -ffinite-math-only, which breaks ggml's
    INFINITY masking. Do not reinstate.
  * --flash-attn off with a turbo/quantized cache is now an error rather than a
    silent override, matching upstream.
  * DFlash conversion delegates to the target model's vocab class instead of
    hardcoding the deepseek-v3 pre-tokenizer.

ABI: llama_model_params gained load_mtp. Every binary that consumes it must be
rebuilt on BOTH machines (llama, llama-server, llama-wp-expert-worker,
test-wp-expert-worker) -- a partial target list is what crash-looped the fleet
on 2026-07-30.

Verified: no conflict markers; fork-marker counts vs backup flat or up
(kv-tier, wp_, weight_pager, mt_pagedattn, mt::, MAD-, routed_expert, ml8,
spec-draft-n-ctx); every hip_xdev function survives; all fork flags still
register in arg.cpp; CPU build (llama, llama-server, llama-cli) clean.
Backup: backup/pre-upstream-sync-2026-07-31 (e6b6856).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018kbRS3KuJquSpjpZXdtSNp
NighmareGit pushed a commit to NighmareGit/atomic-llama-cpp-turboquant that referenced this pull request Aug 8, 2026
Merges 251 upstream commits on top of the fork's 392. Base was 22b208b
(2026-07-15).

What this brings in for DeepSeek V4:

- CUDA kernels for the hyper-connection ops and the lightning indexer
  (dsv4-hc.cu, lightning-indexer.cu, upstream ggml-org#25585 and ggml-org#25545). These
  landed upstream after our base, so the graph no longer needs a CPU
  fallback for those ops.
- MTP and DSpark support (ggml-org#25784), the wo_a reshape fix on load, and the
  same-K/V-cache-type enforcement (ggml-org#25871).
- Exclusion of the i32 ffn_gate_tid2eid routing table from quantization,
  which the fork did not carry.

Conflict resolution kept both architectures everywhere the two sides
touched the same code:

- llama-kv-cache: kept the fork's default-off attention-rotation policy
  and its env overrides, took upstream's GLM_DSA addition to the DSA
  indexer arch list.
- llama-context: moved the TurboQuant flash-attention auto-enable above
  upstream's generic quantized-V check, which would otherwise reject
  turbo cache types under -fa off, and dropped the fork's older V-cache
  check in favour of upstream's.
- mmq.cuh: kept the fork's int64 offsets in all three of upstream's new
  NVFP4 branches.
- fattn.cu: dropped the WMMA block, since upstream removed that kernel
  and its helpers entirely; kept the RDNA4 turbo path.
- ggml-cuda.cu: kept the host-staged cross-device copy and routed its
  peer copy through upstream's new virtual-to-physical device mapping.
- chat.cpp: rebuilt on upstream's file with the fork's Inkling and
  Laguna parsers and the leading-whitespace tolerance reapplied;
  thinking_end_tag became thinking_end_tags upstream.
- laguna.cpp/laguna.py and mtmd-image.cpp: took upstream, which already
  carries the fork's own upstreamed review fixes plus later refinements.
- Removed the inherited upstream workflows again, per 0c9a069.

GGML_OP_COUNT is 103: upstream's 101 plus the fork's TURBO_WHT and
FLASH_ATTN_EXT_BANDED.

Also drops a duplicate LLM_ARCH_LAGUNA case in test-llama-archs that the
merge would otherwise have left in moe_mandatory.
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
* dsv4 hc-ops

* add missing files;

* add cparams

* update rpc version

* address review comments

* address review comments
TrevorS added a commit to TrevorS/llama.cpp that referenced this pull request Sep 4, 2026
Upstream's meta backend aborts on any op its split-state switch does not
name, and ggml-org#25585 added only the hyper-connection ops upstream itself took:

  ggml-backend-meta.cpp:1064: ggml op not implemented: DSV4_HC_FUSED

That kills test-llama-archs on the Meta device for deepseek4, and any path
that estimates memory through that backend. The eight ops left are the fused
HC pair, the fused DSV4 hyper-connection, the three lightning-indexer ops,
the QAT set_rows and the split-attention LSE merge. All of them mix or gather
across a whole row, so they take the same conservative arm as their siblings:
a split source resolves to UNKNOWN rather than a guessed axis.

deepseek4 on the Meta device goes from abort to OK (1.11e-07).

Note for later: GGML_OP_COL2IM_1D is missing from the same switch upstream,
which is upstream's own gap and is left alone here.
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* dsv4 hc-ops

* add missing files;

* add cparams

* update rpc version

* address review comments

* address review comments
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* dsv4 hc-ops

* add missing files;

* add cparams

* update rpc version

* address review comments

* address review comments
TrevorS added a commit to TrevorS/llama.cpp that referenced this pull request Sep 16, 2026
Upstream's meta backend aborts on any op its split-state switch does not
name, and ggml-org#25585 added only the hyper-connection ops upstream itself took:

  ggml-backend-meta.cpp:1064: ggml op not implemented: DSV4_HC_FUSED

That kills test-llama-archs on the Meta device for deepseek4, and any path
that estimates memory through that backend. The eight ops left are the fused
HC pair, the fused DSV4 hyper-connection, the three lightning-indexer ops,
the QAT set_rows and the split-attention LSE merge. All of them mix or gather
across a whole row, so they take the same conservative arm as their siblings:
a split source resolves to UNKNOWN rather than a guessed axis.

deepseek4 on the Meta device goes from abort to OK (1.11e-07).

Note for later: GGML_OP_COL2IM_1D is missing from the same switch upstream,
which is upstream's own gap and is left alone here.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* dsv4 hc-ops

* add missing files;

* add cparams

* update rpc version

* address review comments

* address review comments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants