Repository navigation
llama : fix tensor split for fused qkv with uneven K/V head sizes - #29294
Conversation
|
The MiMo-V2.6-Flash-RL on 2x DGX Sparks proceeds further with this change but still crashes with another error: |
24a8465 to
a048481
Compare
|
The tensor split over RPC seems to work now. However the MTP conversion is still not correct: I run with: llama-server -hf ggml-org/MiMo-V2.6-Flash-RL-GGUF:MXFP4 \
--spec-default --spec-type draft-mtp \
-ub 2048 -c 262144 -lv 4 \
--top-p 0.95 --top-k 0 --temp 1.0 \
--rpc 192.168.100.1:50052,192.168.100.2:50052 -dev RPC0,RPC1 \
--host 0.0.0.0 --port 8013 -lm dio --alias agent-dsv4 -sm tensor 2>&1 | tee ~/llama.logFor some reason the converted MTP sidecar has block indices starting from 48, instead of 0: Used models: https://huggingface.co/ggml-org/MiMo-V2.6-Flash-RL-GGUF |
It looks like
This is normal, the sidecar's blocks should complement main model blocks. |
|
Yeah, just needed 6e0962b |
ggerganov
left a comment
There was a problem hiding this comment.
I confirm it is working. Thanks
…ml-org#29294) * llama : fix tensor split for fused qkv with uneven K/V head sizes Assisted-by: Qwen3.8-27B * fix v granularity * convert: fix mtp conversion * convert: add support for mtp flags * fix loader
Merge upstream commits: - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364) - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294) - metal: split fa kernels into per-dtype libraries (ggml-org#29329) - metal: FWHT kernels for block widths above 512 (ggml-org#29095) - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393) - common: extract shared unicode path/string helpers (ggml-org#29415) - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432) - rpc: include nb in the get_alloc_size cache key (ggml-org#29283) - [SYCL] support sparse FA (ggml-org#28796) - musa: fix PH1 operator failures and build issues (ggml-org#29193) - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231) - opencl: add q5_k bin kernel (ggml-org#29401) - hexagon: add q5_k quant type support (ggml-org#29123) - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404) - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403) - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409) - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422) - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417) Assisted-by: Pi
…ml-org#29294) * llama : fix tensor split for fused qkv with uneven K/V head sizes Assisted-by: Qwen3.8-27B * fix v granularity * convert: fix mtp conversion * convert: add support for mtp flags * fix loader
…ml-org#29294) * llama : fix tensor split for fused qkv with uneven K/V head sizes Assisted-by: Qwen3.8-27B * fix v granularity * convert: fix mtp conversion * convert: add support for mtp flags * fix loader
…ml-org#29294) * llama : fix tensor split for fused qkv with uneven K/V head sizes Assisted-by: Qwen3.8-27B * fix v granularity * convert: fix mtp conversion * convert: add support for mtp flags * fix loader (cherry picked from commit f805c57)
Overview
Additional information
Requirements