Repository navigation
Conversation
|
Shouldn't the new code be in some |
|
You are certainly right because I do not encounter the condition CUDA_ARCH >= 800 in this control flow in my case but this must possibly cause a function cuda error already present, I will review this. |
|
I think the current code is likely to result in lots of compile failures with cuda compute cap >= 8.0. |
|
I hope that the latest additions will allow to work on |
33dcb9a to
d5f31ea
Compare
|
With this, we can notice for a similar token/s in f16 or bf16 (fot short sentence), the results bf16 are identical to an original model bf16 despite fallbacks, numerical fidelity preserved, model behavior unchanged. Tests:
Issue:
|
78b6108 to
d5f31ea
Compare
de5de34 to
633c9c6
Compare
|
I force-pushed a cleaned up version of the branch and would appreciate another review pass. The history is now split into focused commits for capability detection, legacy BF16/FP8 support, cuDNN fallbacks, MoE runtime changes, and regression coverage. I also removed the local The goal of this PR is to make Candle run more reliably on older NVIDIA GPUs by adding legacy BF16/FP8 fallback paths, improving capability detection, preserving numerical behavior where possible, and falling back from failing cuDNN convolution launches. This is mainly aimed at pre-Ampere cards. Example invocation on older GPU using ALLOW_LEGACY all or bf16,fp8: |
|
If you're going all the way with this @guoqingbao now has an emulated FP4 impl as well which works great on SM70 anyway :-) |
|
Trying to use Candle on a GTX 1070, and I was struggling so much (First time rust user, first time Candle user). All that was needed was in my Cargo.toml It would be much nicer if this PR would finally get merged :/ |
|
I understand your frustration. There are a lot of changes in this branch that we definitely want merged into main, but in that lies the issue - there are a lot of changes Reviewing and verifying this is not easy. "It works on my machine" is unfortunately not good enough, because it may break what previously worked on another machine. There are type shims both at cuda compile time and detected at runtime (do we need both?), new code paths to launch kernels, I see a moe kernel from vllm. There's even some new cuda types to aid in vectorized loading - not a bad idea to have but seems a but random to include a performance optimization in this PR? A hundred moving parts, a hundred ways to break. But! I think having this branch as the hub where backwards cuda compatibility is developed - and then extract only the required code into isolated PRs is a good way to get the improvements merged. |
|
@guoqingbao got fp4 working too, and WELL on sm70 at least |
035223f to
dd24dca
Compare
07980b7 to
418f0bb
Compare
|
I cleaned up the legacy CUDA work again. This PR remains intentionally BF16-only, enabled via the cuda-legacy-bf16 feature. The broader backwards-CUDA compatibility work now lives separately in haricot/candle:cuda_legacy, validated with sm_61 CUDA 12.9.x and cuDNN 9.10.2. and combines BF16, FP8, MXFP4, cuDNN fallback and pre-Volta MoE SIMT . I am not proposing cuda_legacy as one large PR. It is an integration/validation branch. I will wait for this BF16 PR to be merged before submitting the changes (FP8, MXFP4, cuDNN fallback, and pre-Volta MoE SIMT), because otherwise it won't compile on my card without the legacy BF16 atomic fallback mechanism. |
|
@sempervictus |
tested and works with:
related EricLBuehler#57