Skip to content

vulkan: optimize IQ4_XS matmul kernels - #28440

Merged
0cc4m merged 1 commit into
ggml-org:0cc4m/vulkan-coopmat-int8from
pwilkin:vulkan-opt-iq4xs-coopmat
Sep 16, 2026
Merged

0cc4m merged 1 commit into
ggml-org:0cc4m/vulkan-coopmat-int8from
pwilkin:vulkan-opt-iq4xs-coopmat

Conversation

@pwilkin

@pwilkin pwilkin commented Sep 5, 2026

Copy link
Copy Markdown
Member

Overview

The coopmat part of the Vulkan IQ4_XS changes from #28415

Additional information

Stack on top of #27952

Requirements

@pwilkin
pwilkin requested review from a team and ggerganov as code owners September 5, 2026 14:04
@github-actions github-actions Bot added testing Everything test related Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning labels Sep 5, 2026
@0cc4m
0cc4m force-pushed the 0cc4m/vulkan-coopmat-int8 branch from 224807f to 8c7611a Compare September 10, 2026 06:35
Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that
expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8
and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling.

Assisted-by: OpenAI Codex

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@pwilkin
pwilkin force-pushed the vulkan-opt-iq4xs-coopmat branch from 5468601 to 6b056d3 Compare September 15, 2026 16:17
@0cc4m
0cc4m merged commit 9ff173d into ggml-org:0cc4m/vulkan-coopmat-int8 Sep 16, 2026
24 of 27 checks passed
0cc4m pushed a commit that referenced this pull request Sep 17, 2026
)

Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that
expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8
and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling.

Assisted-by: OpenAI Codex

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
0cc4m pushed a commit that referenced this pull request Sep 18, 2026
)

Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that
expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8
and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling.

Assisted-by: OpenAI Codex

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
0cc4m pushed a commit that referenced this pull request Sep 24, 2026
)

Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that
expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8
and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling.

Assisted-by: OpenAI Codex

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
0cc4m added a commit that referenced this pull request Sep 24, 2026
…27952)

* vulkan: add int8 coopmat quantized matmul shader

* apply scales inline

* use scalar sums

* probe and directly access coopmat values instead of going through shmem

* add q8_0 support

* add BK_STEP to shader, default to 2

* use larger workgroups

* double buffering

* preload scales

* coopmat load first, then wmma

* use float for scales

* add faster RDNA int->float conversion

* workgroup scheduling for cache proximity

* clean up

* use wave32

* restructure for vgpr use

* skip computation for inactive tiles

* only force subgroup size 32 on AMD RDNA

* use BK_STEP 4

* fix compilation

* move quant-specific prefetch function out of main file

* add q4_1, q5_0, q5_1 support

* restructure mmq cm1 functions

* enable mul_mat_id support

* fix segfault

* fix mul_mat_id bug

* support iq4_nl and mxfp4

* remove elem row/col fast path, invalid for RDNA4

* use shmem arrays for LUTs

* use 4-byte loads where possible

* add q3_k, q4_k, q5_k, q6_k and nvfp4 support

* fix l warptile

* improve performance

* improve performance

* improvements

* dedup b scales

* merge shmem arrays

* undo uint8_t, gate to RDNA3/4

* add RDNA4 architecture, use for hardcoded coopmat elem thread access, set BK_STEP back to 4

* improve offset application

* clean up

* fix iq4_nl and nvfp4 performance

* rdna4 tuning

* use BK_STEP 2 on MUL_MAT_ID

* adapt to upstream changes

* fix shmem support function, clean up comments

* fix warptile logic

Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>

* vulkan: add IQ4_XS support to the coopmat1 integer matmul shader (#28440)

Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that
expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8
and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling.

Assisted-by: OpenAI Codex

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* avoid compiling f16 acc shader variants

---------

Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning testing Everything test related Vulkan Issues specific to the Vulkan backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants