Skip to content

vulkan: Use native e2m1 and e4m3 conversions for mxfp4/nvfp4 - #25338

Merged
jeffbolznv merged 1 commit into
ggml-org:masterfrom
jeffbolznv:fp4_fp8
Jul 13, 2026
Merged

jeffbolznv merged 1 commit into
ggml-org:masterfrom
jeffbolznv:fp4_fp8

Conversation

@jeffbolznv

Copy link
Copy Markdown
Contributor

Overview

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type promotions, and also uses the float8 extension to do ue4m3 promotions for nvfp4. It's reasonable to assume that an implementation that supports fp4 will also support fp8, so we don't need to handle all possible combinations of support.

For NVIDIA, this extension is currently only supported in our Vulkan developer driver (https://developer.nvidia.com/vulkan-driver). It also requires some not-yet-merged glslang support (KhronosGroup/glslang#4325).

Perf results:

before:

Z:\github\jeffbolznv\llama.cpp\build\bin\RelWithDebInfo>llama-bench.exe -fa 1 -n 0 -p 512 -r 10 --prio 1 -m c:\models\gpt-oss-20b-mxfp4.gguf -m C:\models\gemma-4-26B-A4B-it-NVFP4.gguf -m C:\models\gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2v
| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | Vulkan     |  -1 |   1 |           pp512 |    11630.52 ± 162.11 |
| gemma4 26B.A4B NVFP4           |  16.45 GiB |    25.23 B | Vulkan     |  -1 |   1 |           pp512 |     8202.02 ± 544.78 |
| gemma4 31B NVFP4               |  17.97 GiB |    30.70 B | Vulkan     |  -1 |   1 |           pp512 |      3373.52 ± 11.27 |

Z:\github\jeffbolznv\llama.cpp\build\bin\RelWithDebInfo>llama-bench.exe -fa 1 -n 128 -p 0 -r 10 --prio 1 -m c:\models\gpt-oss-20b-mxfp4.gguf -m C:\models\gemma-4-26B-A4B-it-NVFP4.gguf -m C:\models\gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2v
| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | Vulkan     |  -1 |   1 |           tg128 |       321.90 ± 28.33 |
| gemma4 26B.A4B NVFP4           |  16.45 GiB |    25.23 B | Vulkan     |  -1 |   1 |           tg128 |        135.62 ± 2.01 |
| gemma4 31B NVFP4               |  17.97 GiB |    30.70 B | Vulkan     |  -1 |   1 |           tg128 |         44.76 ± 0.16 |

after:

Z:\github\jeffbolznv\llama.cpp\build\bin\RelWithDebInfo>llama-bench.exe -fa 1 -n 0 -p 512 -r 10 --prio 1 -m c:\models\gpt-oss-20b-mxfp4.gguf -m C:\models\gemma-4-26B-A4B-it-NVFP4.gguf -m C:\models\gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 1 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2v
| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | Vulkan     |  -1 |   1 |           pp512 |     12067.91 ± 96.60 |
| gemma4 26B.A4B NVFP4           |  16.45 GiB |    25.23 B | Vulkan     |  -1 |   1 |           pp512 |     8565.77 ± 329.21 |
| gemma4 31B NVFP4               |  17.97 GiB |    30.70 B | Vulkan     |  -1 |   1 |           pp512 |      3490.61 ± 18.59 |

Z:\github\jeffbolznv\llama.cpp\build\bin\RelWithDebInfo>llama-bench.exe -fa 1 -n 128 -p 0 -r 10 --prio 1 -m c:\models\gpt-oss-20b-mxfp4.gguf -m C:\models\gemma-4-26B-A4B-it-NVFP4.gguf -m C:\models\gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 1 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2v
| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | Vulkan     |  -1 |   1 |           tg128 |       340.07 ± 27.19 |
| gemma4 26B.A4B NVFP4           |  16.45 GiB |    25.23 B | Vulkan     |  -1 |   1 |           tg128 |        137.94 ± 7.67 |
| gemma4 31B NVFP4               |  17.97 GiB |    30.70 B | Vulkan     |  -1 |   1 |           tg128 |         51.74 ± 0.77 |

coopmat1 before:

Z:\github\jeffbolznv\llama.cpp\build\bin\RelWithDebInfo>llama-bench.exe -fa 1 -n 0 -p 512 -r 10 --prio 1 -m c:\models\gpt-oss-20b-mxfp4.gguf -m C:\models\gemma-4-26B-A4B-it-NVFP4.gguf -m C:\models\gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | Vulkan     |  -1 |   1 |           pp512 |     7148.97 ± 153.38 |
| gemma4 26B.A4B NVFP4           |  16.45 GiB |    25.23 B | Vulkan     |  -1 |   1 |           pp512 |      6332.03 ± 81.12 |
| gemma4 31B NVFP4               |  17.97 GiB |    30.70 B | Vulkan     |  -1 |   1 |           pp512 |       2100.50 ± 5.29 |

coopmat1 after:

Z:\github\jeffbolznv\llama.cpp\build\bin\RelWithDebInfo>llama-bench.exe -fa 1 -n 0 -p 512 -r 10 --prio 1 -m c:\models\gpt-oss-20b-mxfp4.gguf -m C:\models\gemma-4-26B-A4B-it-NVFP4.gguf -m C:\models\gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 5090 (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 1 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
| model                          |       size |     params | backend    | ngl |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE          |  11.27 GiB |    20.91 B | Vulkan     |  -1 |   1 |           pp512 |     7600.39 ± 142.37 |
| gemma4 26B.A4B NVFP4           |  16.45 GiB |    25.23 B | Vulkan     |  -1 |   1 |           pp512 |      6630.41 ± 57.77 |
| gemma4 31B NVFP4               |  17.97 GiB |    30.70 B | Vulkan     |  -1 |   1 |           pp512 |       2411.14 ± 7.43 |

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, used codex to do much of the implementation. I reviewed/fixed things.

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type
promotions, and also uses the float8 extension to do ue4m3 promotions for
nvfp4. It's reasonable to assume that an implementation that supports fp4 will
also support fp8, so we don't need to handle all possible combinations of
support.
@jeffbolznv
jeffbolznv requested a review from a team as a code owner July 6, 2026 02:53
@github-actions github-actions Bot added Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning labels Jul 6, 2026
@0cc4m

0cc4m commented Jul 13, 2026

Copy link
Copy Markdown
Contributor

Haven't tested the native path, waiting for dependencies to catch up. @ggml-org/maintainers Another approval needed.

@jeffbolznv
jeffbolznv merged commit e920c52 into ggml-org:master Jul 13, 2026
39 of 40 checks passed
RehanQasim-dev pushed a commit to aifoundry-org/llama.cpp that referenced this pull request Jul 23, 2026
…g#25338)

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type
promotions, and also uses the float8 extension to do ue4m3 promotions for
nvfp4. It's reasonable to assume that an implementation that supports fp4 will
also support fp8, so we don't need to handle all possible combinations of
support.
RehanQasim-dev pushed a commit to aifoundry-org/llama.cpp that referenced this pull request Jul 23, 2026
…g#25338)

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type
promotions, and also uses the float8 extension to do ue4m3 promotions for
nvfp4. It's reasonable to assume that an implementation that supports fp4 will
also support fp8, so we don't need to handle all possible combinations of
support.
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
…g#25338)

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type
promotions, and also uses the float8 extension to do ue4m3 promotions for
nvfp4. It's reasonable to assume that an implementation that supports fp4 will
also support fp8, so we don't need to handle all possible combinations of
support.
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…g#25338)

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type
promotions, and also uses the float8 extension to do ue4m3 promotions for
nvfp4. It's reasonable to assume that an implementation that supports fp4 will
also support fp8, so we don't need to handle all possible combinations of
support.
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…g#25338)

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type
promotions, and also uses the float8 extension to do ue4m3 promotions for
nvfp4. It's reasonable to assume that an implementation that supports fp4 will
also support fp8, so we don't need to handle all possible combinations of
support.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…g#25338)

This uses the new VK_EXT_shader_ocp_microscaling_types extension to do fp4 type
promotions, and also uses the float8 extension to do ue4m3 promotions for
nvfp4. It's reasonable to assume that an implementation that supports fp4 will
also support fp8, so we don't need to handle all possible combinations of
support.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants