Skip to content

cuda: mmvf v16 bf16 unpack builds on pre-Ampere / CUDA 12.0 - #1

Merged
neurall merged 1 commit into
neurall:releasefrom
homeofe:fix-mmvf-bf16-pre-ampere
Oct 2, 2026
Merged

neurall merged 1 commit into
neurall:releasefrom
homeofe:fix-mmvf-bf16-pre-ampere

Conversation

@homeofe

@homeofe homeofe commented Oct 2, 2026

Copy link
Copy Markdown

Problem

The release branch does not build for pre-Ampere GPUs with CUDA 12.0. mmvf_v16_unpack calls __bfloat1622float2 directly, which is not available there:

ggml/src/ggml-cuda/mmvf.cu(435): error: identifier "__bfloat1622float2" is undefined

Change

One line: use ggml_cuda_cast<float2>(h[k]) from convert.cuh. It already wraps __bfloat1622float2 behind __CUDA_ARCH__ >= 800 (and has the HIP / MUSA paths), and falls back to two scalar __bfloat162float conversions otherwise. Ampere and newer compile to the same intrinsic as before.

Tests

  • Before: cmake -B build -G Ninja -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=75 -DGGML_NATIVE=ON -DLLAMA_CURL=OFF -DCMAKE_BUILD_TYPE=Release && cmake --build build --target llama-server llama-bench llama-cli fails with the error above.
  • After: the same build completes (282/282).
  • Runtime: MiMo-V2.6-Flash IQ2_M with the expert cache on this machine is being measured now; I will add the numbers here.

Test machine

RTX 2080 Ti 11 GB (sm_75, PCIe 3.0 x16), Threadripper 3960X, 128 GB DDR4-3200 quad channel, Ubuntu 24.04, CUDA 12.0.140, gcc 13.3, driver 580.178.04.

__bfloat1622float2 is not available for sm_75 with CUDA 12.0. Use ggml_cuda_cast<float2>,
which already guards it (__CUDA_ARCH__ >= 800) and falls back to two scalar conversions.
@homeofe

homeofe commented Oct 2, 2026

Copy link
Copy Markdown
Author

Runtime check done: with this fix the branch builds and runs MiMo-V2.6-Flash IQ2_M on the RTX 2080 Ti (sm_75, CUDA 12.0). Two more findings from that run, both posted with numbers in ggml-org#27861:

@homeofe

homeofe commented Oct 2, 2026

Copy link
Copy Markdown
Author

More data from the same CPU/RAM, now on an RX 9070 XT 16 GB (HIP, gfx1201, Windows 11), MiMo-V2.6-Flash IQ2_M. With ~15% of the model in VRAM your fork's auto placement (21 slots/layer, 8.1 GiB, hit ~37%) matches stock decode (9.1 vs 9.1 tok/s) and reads prompts 1.75x faster (241 vs 138 tok/s at 28K tokens), also ahead of the Vulkan release (196). On the 2080 Ti (~10% in VRAM) it was 0.43x at decode. Details in ggml-org#27861.

Building the fork for Windows HIP needed nothing fork-specific beyond this PR's fix. The one obstacle was HIP SDK 7.2's clang with MSVC >= 14.40 ('isgreater' cannot overload, ggml-org#22570), which is a toolchain issue and not about the fork.

@neurall neurall left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks right: ggml_cuda_cast wraps the same intrinsic on sm_80+ and has the pre-Ampere fallback; builds on sm_75 / CUDA 12.0 per the author's 282/282. Thanks.

@neurall
neurall merged commit 43cf876 into neurall:release Oct 2, 2026
neurall pushed a commit that referenced this pull request Oct 4, 2026
__bfloat1622float2 is not available for sm_75 with CUDA 12.0. Use ggml_cuda_cast<float2>,
which already guards it (__CUDA_ARCH__ >= 800) and falls back to two scalar conversions.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants