DeepseekV4: Add fused hyper-connection ops - #25585
Conversation
|
@am17an looks like there's an overlap with #25421 . I added the Sinkhorn step as a standalone SINKHORN_NORM op, you've fused the whole HC block, which is the better call. I came at this from the AMD side. Profiling deepseek-v4-flash on an MI250X, the decomposed Sinkhorn was ~30% of gpu kernel time across ~250k tiny 4×4 dispatches, fusing it gave me +60.9% decode / +7.4% prefill on 4 GCDs. Since we're solving the same thing, want to join forces? I'm happy to close mine and help on this one, especially validating/tuning it on MI250X if that's not already done/ any other additions you would like to delegate ? |
|
@jadenmach2 sure, thanks! I've only tested these on nvidia gpus |
|
@am17an The new ops are producing nans sometimes: |
This is just my wild guess, but isn't that because of uninitialized sentinel tensors? |
|
Yeah I think that's the most likely cause. AFK in case @fairydreaming you want to pick this up, otherwise I will fix in a few hours |
|
@ggerganov @am17an This should fix it: #25822 |
* dsv4 hc-ops * add missing files; * add cparams * update rpc version * address review comments * address review comments
Merge upstream ggml-org/llama.cpp master (876a432) into fork master. Fork point was 2026-07-06 (20a04b2): 330 commits behind, 965 ahead. 37 files conflicted; all resolved by hand (never --theirs), plus fixes for cleanly-merged files that referenced APIs changed elsewhere. Brings in DSpark speculative decoding (8407527) and the DeepSeek V4 work that landed after our fork point, including the fused hyper-connection ops (0dc74e3, ggml-org#25585), DS4 seq_rm fix, graph-split reduction, MTP tensor loading, dflash K/V rotation and sidecar auto-download. Upstream removals that forced fork-side ports: * -sm row / CUDA split buffers removed (74976e1). Dropped our copy of the split-buffer implementation (no fork code in it) and rewrote ggml_cuda_mul_mat onto upstream's early-return dispatch, re-inserting the ML8_FP8 route, the F8_E4M3 guard and the TQ4_1S/TQ3_1S kernels. ggml_cuda_Memcpy2DPeerAsync is kept -- it carries our no-P2P host staging. * mmq.cuh rewritten upstream (<type,J,fallback>, y_scale, stream-k helper). Ported the MAD-88 routed-expert pointer hook to the new kernel signature, args struct and both launch sites. * WMMA flash-attention kernel deleted upstream; our RDNA4 path now leads. * use_mmap/use_mlock/use_direct_io collapsed into llama_load_mode. Weight paging now strips only the mmap bit instead of forcing no-mmap. * llama_context auto-FA/GDN resolution folded into resolve_fused_ops(); the WP attention-island guard is re-injected there, scoped to flash-attn. * server: draft/MTP context now built by common_speculative_init_from_params, so --spec-draft-n-ctx and the MTP tier-disable move onto params_dft; slot memory ops go through slot.mem; migrated two subprocess sites to common_subproc; folded our byte-backpressure into upstream's server_pipe (upstream's drops the oldest item, which would corrupt a proxied body). Fork-visible behaviour changes: * GGML_TYPE_Q2_0 is 56 here, not upstream's 42 -- 42 is our TURBO3_0 and type ids are on-disk. Our turbo/ml8 GGUFs stay readable; an upstream Q2_0 GGUF needs reconversion. gguf-py kept in sync. * DS4 per-layer output renamed l_out -> l_last upstream; added l_last to the WP FFN-island pin or it silently stops firing on DeepSeek V4. * Dropped our -ffast-math on HIP: upstream sets -funsafe-math-optimizations instead because -ffast-math implies -ffinite-math-only, which breaks ggml's INFINITY masking. Do not reinstate. * --flash-attn off with a turbo/quantized cache is now an error rather than a silent override, matching upstream. * DFlash conversion delegates to the target model's vocab class instead of hardcoding the deepseek-v3 pre-tokenizer. ABI: llama_model_params gained load_mtp. Every binary that consumes it must be rebuilt on BOTH machines (llama, llama-server, llama-wp-expert-worker, test-wp-expert-worker) -- a partial target list is what crash-looped the fleet on 2026-07-30. Verified: no conflict markers; fork-marker counts vs backup flat or up (kv-tier, wp_, weight_pager, mt_pagedattn, mt::, MAD-, routed_expert, ml8, spec-draft-n-ctx); every hip_xdev function survives; all fork flags still register in arg.cpp; CPU build (llama, llama-server, llama-cli) clean. Backup: backup/pre-upstream-sync-2026-07-31 (e6b6856). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018kbRS3KuJquSpjpZXdtSNp
Merges 251 upstream commits on top of the fork's 392. Base was 22b208b (2026-07-15). What this brings in for DeepSeek V4: - CUDA kernels for the hyper-connection ops and the lightning indexer (dsv4-hc.cu, lightning-indexer.cu, upstream ggml-org#25585 and ggml-org#25545). These landed upstream after our base, so the graph no longer needs a CPU fallback for those ops. - MTP and DSpark support (ggml-org#25784), the wo_a reshape fix on load, and the same-K/V-cache-type enforcement (ggml-org#25871). - Exclusion of the i32 ffn_gate_tid2eid routing table from quantization, which the fork did not carry. Conflict resolution kept both architectures everywhere the two sides touched the same code: - llama-kv-cache: kept the fork's default-off attention-rotation policy and its env overrides, took upstream's GLM_DSA addition to the DSA indexer arch list. - llama-context: moved the TurboQuant flash-attention auto-enable above upstream's generic quantized-V check, which would otherwise reject turbo cache types under -fa off, and dropped the fork's older V-cache check in favour of upstream's. - mmq.cuh: kept the fork's int64 offsets in all three of upstream's new NVFP4 branches. - fattn.cu: dropped the WMMA block, since upstream removed that kernel and its helpers entirely; kept the RDNA4 turbo path. - ggml-cuda.cu: kept the host-staged cross-device copy and routed its peer copy through upstream's new virtual-to-physical device mapping. - chat.cpp: rebuilt on upstream's file with the fork's Inkling and Laguna parsers and the leading-whitespace tolerance reapplied; thinking_end_tag became thinking_end_tags upstream. - laguna.cpp/laguna.py and mtmd-image.cpp: took upstream, which already carries the fork's own upstreamed review fixes plus later refinements. - Removed the inherited upstream workflows again, per 0c9a069. GGML_OP_COUNT is 103: upstream's 101 plus the fork's TURBO_WHT and FLASH_ATTN_EXT_BANDED. Also drops a duplicate LLM_ARCH_LAGUNA case in test-llama-archs that the merge would otherwise have left in moe_mandatory.
* dsv4 hc-ops * add missing files; * add cparams * update rpc version * address review comments * address review comments
Upstream's meta backend aborts on any op its split-state switch does not name, and ggml-org#25585 added only the hyper-connection ops upstream itself took: ggml-backend-meta.cpp:1064: ggml op not implemented: DSV4_HC_FUSED That kills test-llama-archs on the Meta device for deepseek4, and any path that estimates memory through that backend. The eight ops left are the fused HC pair, the fused DSV4 hyper-connection, the three lightning-indexer ops, the QAT set_rows and the split-attention LSE merge. All of them mix or gather across a whole row, so they take the same conservative arm as their siblings: a split source resolves to UNKNOWN rather than a guessed axis. deepseek4 on the Meta device goes from abort to OK (1.11e-07). Note for later: GGML_OP_COL2IM_1D is missing from the same switch upstream, which is upstream's own gap and is left alone here.
* dsv4 hc-ops * add missing files; * add cparams * update rpc version * address review comments * address review comments
* dsv4 hc-ops * add missing files; * add cparams * update rpc version * address review comments * address review comments
Upstream's meta backend aborts on any op its split-state switch does not name, and ggml-org#25585 added only the hyper-connection ops upstream itself took: ggml-backend-meta.cpp:1064: ggml op not implemented: DSV4_HC_FUSED That kills test-llama-archs on the Meta device for deepseek4, and any path that estimates memory through that backend. The eight ops left are the fused HC pair, the fused DSV4 hyper-connection, the three lightning-indexer ops, the QAT set_rows and the split-attention LSE merge. All of them mix or gather across a whole row, so they take the same conservative arm as their siblings: a split source resolves to UNKNOWN rather than a guessed axis. deepseek4 on the Meta device goes from abort to OK (1.11e-07). Note for later: GGML_OP_COL2IM_1D is missing from the same switch upstream, which is upstream's own gap and is left alone here.
* dsv4 hc-ops * add missing files; * add cparams * update rpc version * address review comments * address review comments
Overview
Add the sinkhorn ops to the DeepseekV4 graph. Graph nodes go from 29k to 8k with this change. Big increases in the TG + PP
Additional information
Requirements