Repository navigation
cuda: mmvf v16 bf16 unpack builds on pre-Ampere / CUDA 12.0 - #1
Conversation
__bfloat1622float2 is not available for sm_75 with CUDA 12.0. Use ggml_cuda_cast<float2>, which already guards it (__CUDA_ARCH__ >= 800) and falls back to two scalar conversions.
|
Runtime check done: with this fix the branch builds and runs MiMo-V2.6-Flash IQ2_M on the RTX 2080 Ti (sm_75, CUDA 12.0). Two more findings from that run, both posted with numbers in ggml-org#27861:
|
|
More data from the same CPU/RAM, now on an RX 9070 XT 16 GB (HIP, gfx1201, Windows 11), MiMo-V2.6-Flash IQ2_M. With ~15% of the model in VRAM your fork's auto placement (21 slots/layer, 8.1 GiB, hit ~37%) matches stock decode (9.1 vs 9.1 tok/s) and reads prompts 1.75x faster (241 vs 138 tok/s at 28K tokens), also ahead of the Vulkan release (196). On the 2080 Ti (~10% in VRAM) it was 0.43x at decode. Details in ggml-org#27861. Building the fork for Windows HIP needed nothing fork-specific beyond this PR's fix. The one obstacle was HIP SDK 7.2's clang with MSVC >= 14.40 ( |
neurall
left a comment
There was a problem hiding this comment.
Looks right: ggml_cuda_cast wraps the same intrinsic on sm_80+ and has the pre-Ampere fallback; builds on sm_75 / CUDA 12.0 per the author's 282/282. Thanks.
__bfloat1622float2 is not available for sm_75 with CUDA 12.0. Use ggml_cuda_cast<float2>, which already guards it (__CUDA_ARCH__ >= 800) and falls back to two scalar conversions.
Problem
The
releasebranch does not build for pre-Ampere GPUs with CUDA 12.0.mmvf_v16_unpackcalls__bfloat1622float2directly, which is not available there:Change
One line: use
ggml_cuda_cast<float2>(h[k])fromconvert.cuh. It already wraps__bfloat1622float2behind__CUDA_ARCH__ >= 800(and has the HIP / MUSA paths), and falls back to two scalar__bfloat162floatconversions otherwise. Ampere and newer compile to the same intrinsic as before.Tests
cmake -B build -G Ninja -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=75 -DGGML_NATIVE=ON -DLLAMA_CURL=OFF -DCMAKE_BUILD_TYPE=Release && cmake --build build --target llama-server llama-bench llama-clifails with the error above.Test machine
RTX 2080 Ti 11 GB (sm_75, PCIe 3.0 x16), Threadripper 3960X, 128 GB DDR4-3200 quad channel, Ubuntu 24.04, CUDA 12.0.140, gcc 13.3, driver 580.178.04.