Force NVFP4 W4A8 path for NVFP4_W4A16 layers on Blackwell, where NVFP4 normally uses the native W4A4 path. - #24364
Conversation
| mul_mat_q_case<GGML_TYPE_NVFP4, true>(ctx, args, stream); | ||
| break; | ||
| } | ||
| #endif // GGML_CUDA_HAS_BLACKWELL_TARGET |
There was a problem hiding this comment.
How do we opt out from this? A drop from 5486.02 to 4492.20 is very severe.
If you want a higher precision, there's a variety of Q4 quants that are just as small and even more precise (see #23572 for detailed comparisons).
There was a problem hiding this comment.
There is no need to opt out from this, if you want to run W4A4 take a checkpoint with W4A4.
The intention of this PR is if a checkpoint has W4A16 layers it should have activations in higher precision.
This PR doesn't cause regression on pure W4A4_NVFP4 checkpoints.
You can confirm there is no regression by testing llama-bench with this PR on below checkpoints:
- Pure W4A4_NVFP4: https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-NVFP4
- W4A16_NVFP4: https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4
There was a problem hiding this comment.
There may be a case that there aren't other checkpoints or somebody likes one's calibration over another. I think many prefer increased speed and that is why they pick NVFP4. Usually Q4_K~ will always be better precision than NVFP4 if that is what they are going for and then it may end up faster than skipping the native FP4. We're blending checkpoints with Q_K quants and NVFP4 combined that can compensate for ppl loss. I think some selection control by the user would be a good balance to let them decide.
There was a problem hiding this comment.
ModelOpt, Redhat came up with W4A16_NVFP4 after through investigations on performance and accuracy. Checkout this weight-only-quantization-schemes.
We should honor the intention behind creating a recipe/checkpoint which was specifically designed for W4A16_NVFP4.
If someone intends to use W4A4 on Blackwell they should use a W4A4 checkpoint.
Few PR on vLLM for reference, that were added by ModelOpt to support W4A16_NVFP4 :
There was a problem hiding this comment.
I think some selection control by the user would be a good balance to let them decide.
I mean we can in principle add such a knob in the CUDA backend to give the ability to override model builder intents encoded in the GGUF, but the default should be what is encoded in the GGUF
There was a problem hiding this comment.
Added the knob for user controlled W4A4_NVFP4 fast path on Blackwell, even if the checkpoint specifies W4A16_NVFP4 layers
am17an
left a comment
There was a problem hiding this comment.
I didn't look into the PR in detail, but does using 8-bit activation disable the fp4 tensor core?
Yes. Basically you can think of this as a step towards the support of weight-only-quantization-schemes in llama.cpp. |
b72a8c9 to
18f1df3
Compare
ORippler
left a comment
There was a problem hiding this comment.
@ggerganov @CISC thoughts on this approach and the required granularity for W4A4 vs. W4A16?
| if (CMAKE_CUDA_ARCHITECTURES MATCHES "(^|;)12[0-9]a(-real|-virtual)?($|;)") | ||
| add_compile_definitions(GGML_CUDA_HAS_BLACKWELL_TARGET) | ||
| endif() |
There was a problem hiding this comment.
Just FYI this compile definition will be visible for all archs in CMAKE_CUDA_ARCHITECTURES (per arch specialization requires constructing nvcc commands by hand)
| if self._is_nvfp4: | ||
| for tensor_name, entry in quant_layers.items(): | ||
| if not isinstance(entry, dict) or entry.get("quant_algo") != "W4A16_NVFP4": | ||
| continue | ||
| if "lm_head" in tensor_name or "output" in tensor_name: | ||
| self._nvfp4_w4a16_output = True | ||
| continue | ||
| bid_m = re.search(r'\.layers\.(\d+)\.', tensor_name) | ||
| if bid_m: | ||
| self._nvfp4_w4a16_blocks.add(int(bid_m.group(1))) |
There was a problem hiding this comment.
https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4/blob/main/config.json#L455-L470
huggingface/modelopt seems to offer storing this information on a per-weight level, as opposed to per-transformer-block we parse here. Later on, we set the marker also on a per-weight/op level. I feel this is inconsistent and we should adapt the conversion script to support the finer level of granularity also
ggerganov
left a comment
There was a problem hiding this comment.
If I read correctly, this introduces a path to do w4a8 where we normally do w4a4. And the reason is that w4a8 has higher accuracy?
Is there a case where we would want to use w4a4? If no, we can probably simplify a lot of the logic.
Yet, if there is specific tensor hardware for that, it's probably useful.
| // NVFP4 W4A16: per-layer + LM head flag, true where NVFP4 weights skip activation quantization. | ||
| std::array<bool, LLAMA_MAX_LAYERS> nvfp4_w4a16_layer_arr = {}; | ||
| bool nvfp4_w4a16_output = false; | ||
|
|
There was a problem hiding this comment.
This does not feel like it belongs to hparams. It's too low-level, backend-specific information.
There was a problem hiding this comment.
Should we make this as a per-tensor attribute in llama-model at load time? This will also help acknowledge @ORippler's comment as well to have more flexibility for per-tensor information.
https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4/blob/main/config.json#L455-L470
huggingface/modelopt seems to offer storing this information on a per-weight level, as opposed to per-transformer-block we parse here. Later on, we set the marker also on a per-weight/op level. I feel this is inconsistent and we should adapt the conversion script to support the finer level of granularity also
There was a problem hiding this comment.
Yes, we should formulate this information to be per tensor. Also, it should abstract away the NVFP4 and make it more generic. For example, "allow 4-bit activations (bool)" seems generic enough.
We can represent this information with a map: tensor name -> bool. The values are false by default. The map is optional. To represent the map in GGUF, you'll need 2 arrays - the first one with the tensor names and the second one with the same size and the bool values.
There was a problem hiding this comment.
"allow 4-bit activations (bool)"
I intuitively think about this PR/feature as "enabling Weight-only-Quantization schemes", so naming should signal this intent if you agree -> "weight-only quantization (bool)" / "allow activation quantization (bool)".
W4A16 is effectively weight-only, as LLMs are trained in BF16 precision typically
There was a problem hiding this comment.
In some backends, we already quantize the activations by default (to 8-bits). So I'm not sure if it is not going to be a bit misleading to call it "weight-only quantization (bool)" / "allow activation quantization (bool)".
There was a problem hiding this comment.
Updated the change so allow-activation-quant is stored per-tensor. Since it is currently NVFP4-specific, activation quantization is enabled by default for all layers. For layers that use W4A16_NVFP4 as the quantization method, the activation-quant array stores the tensor name paired with false. Please review.
Yes on both. The motivation is that doing W4A16 (or weight-only-quantization in general) may allow one to quantize more layers than when one quantizes both activations and weights. This reduces memory-footprint of the model, and furthermore increases decode throughput - decode is effectively weight-streaming in local inference with small BS setting) |
|
Correct me if I am wrong, but I think in the CUDA backend we already do w4a8 for existing types like Q4 - we do it by default. If that's the case, adding support for w4a8+nvfp4 would make sense only if it is better quality-wise than the existing w4a8+q4 methods. |
I think that's going to be difficult to prove generally, though I can add the single model we have validated so far:
|
|
For my understanding, the nvidia/Qwen3.6-35B-A3B-NVFP4 readme says that the model was quantized with the Model Optimizer. Does the quantization process involve some form of quantization aware training? My understanding is that if a model is trained natively in NVFP4 format, the best thing to do is keep the format intact during inference. I.e. it's not beneficial to use Q4 quantizations in such case on hardware that supports NVFP4. But I would assume that if the model was trained in NVFP4 in the first place, then the recommended way to run would be If the model was not trained in NVFP4, and instead was quantized to NVFP4, then we have to justify and demonstrate in which cases it is worth it compared to the Q4 formats. One simple justification is that It might be worth taking a look at the KLD between these quantizations and a reference BF16 model. For example, let's take a look at KLD and speed for: |
No, training was involved, the checkpoint is yielded by PTQ only. QAT/QAD would be disclosed in the model card like this https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4#training-methodology (this is a W4A4 checkpoint).
As outlined above, the proposed benefit of weight-only quantization is to allow the quantization of weights more aggressive at iso-quality. We will try to collect & present some numbers. In general, the idea would be to go for NVFP4 with W4A16, the W4A8 path is just a transient step. |
18f1df3 to
2c7052a
Compare
For my understanding, is the expectation that W4A16 would be faster compared to W4A8 if implemented efficiently? |
| #### GGML_CUDA_FORCE_W4A4 | ||
|
|
||
| NVFP4 models that carry W4A16 layers request higher-precision activations, so on Blackwell those layers run through the W4A8 path instead of the native W4A4 path. Set `GGML_CUDA_DISABLE_FORCE_W4A8=1` to ignore that request and keep the native W4A4 path for faster prompt processing at the cost of accuracy. | ||
| NVFP4 models that carry W4A16 layers request higher-precision activations (W4A8), so on Blackwell those layers run through the W4A8 path instead of the native W4A4 path. Set `GGML_CUDA_FORCE_W4A4=1` to override that request and keep the native W4A4 path for faster prompt processing at the cost of accuracy. |
There was a problem hiding this comment.
At some point we can remove this compile-time option and toggle this functionality at runtime through a libllama argument. It will simply skip setting the quantization hints for the matrix multiplications (i.e. override the activation policy). The idea is to avoid "communicating" directly with backend.
| @@ -443,7 +443,7 @@ extern "C" { | |||
| enum ggml_op_hint { | |||
| GGML_HINT_NONE = 0, | |||
| GGML_HINT_SRC0_IS_HADAMARD = 1, | |||
| GGML_HINT_NO_QUANT_SRC1 = 2, // W4A16_NVFP4: keep activations higher precision. | |||
| GGML_HINT_NO_QUANT_SRC1 = 2, // keep src1 at higher precision (skip 4-bit activation quant). | |||
There was a problem hiding this comment.
If we agree on the "allow 4-bit activations", then we should consistently name where relevant. For example here: GGML_HINT_SRC1_ALLOW_4BIT.
| # Per-tensor activation precision policy (tensor name -> allow 4-bit activations). | ||
| ALLOW_4BIT_ACT_TENSOR = "general.allow_4bit_act.tensor" | ||
| ALLOW_4BIT_ACT_VALUE = "general.allow_4bit_act.value" |
There was a problem hiding this comment.
This can become more generic and future-proof to allow setting additional per-tensor options in the future. To do that, the array with the tensor names should be called something like "general.tensor_extra.name". And the "allow 4-bit activations" values should be stored in the bool array named "general.tensor_extra.allow_4bit_act".
| if (no_quant_src1_for_weight(act_policy, w)) { | ||
| ggml_mul_mat_set_hint(res, GGML_HINT_NO_QUANT_SRC1); | ||
| } | ||
|
|
There was a problem hiding this comment.
Should be:
| if (no_quant_src1_for_weight(act_policy, w)) { | |
| ggml_mul_mat_set_hint(res, GGML_HINT_NO_QUANT_SRC1); | |
| } | |
| if (llama_act_policy_allow_4bit(act_policy, w)) { | |
| ggml_mul_mat_set_hint(res, GGML_HINT_SRC1_ALLOW_4BIT); | |
| } | |
| @@ -18,6 +18,7 @@ struct ggml_tensor; | |||
|
|
|||
| struct llama_cparams; | |||
| struct llama_layer; | |||
| struct llama_weight_act_policy; | |||
There was a problem hiding this comment.
| struct llama_weight_act_policy; | |
| struct llama_act_policy; |
|
My opinion is that we should solve this differently at a ggml level. As of right now we have this for "precision": // precision
enum ggml_prec {
GGML_PREC_DEFAULT = 0, // stored as ggml_tensor.op_params, 0 by default
GGML_PREC_F32 = 10,
};I am interpreting "default" to mean that we don't really care about the precision and that the backends should just optimize for speed/memory use. So on Blackwell that would mean W4A4. But we can add something like To be clear: I think that the way we currently define "precision" in ggml is poorly defined and needs more maintainer attention. My opinion is that |
This makes sense to me. It would reduce the ambiguity created by having multiple precision-related signals and hints. |
|
I agree it would be nice to have more clear instructions to the backend about what precision to use, but who makes that choice and how? And I think we need to be more clear about weight/activation precision vs accumulator precision/range. |
0a1701e to
a470fa6
Compare
7d987b6 to
fabd334
Compare
|
As @ORippler is out this week and I have addressed his comments . Can we mark this PR merge-ready? |
|
@ynankani I can't push to your branch - could you apply this patch: diff --git a/src/llama-context.cpp b/src/llama-context.cpp
index e4c218d67b..795803b022 100644
--- a/src/llama-context.cpp
+++ b/src/llama-context.cpp
@@ -2559,7 +2559,7 @@ llm_graph_params llama_context::graph_params(
/*.loras =*/ loras.get(),
/*.mctx =*/ mctx,
/*.cross =*/ &cross,
- /*.act_policy =*/ &model.act_policy,
+ /*.prec_policy =*/ &model.prec_policy,
/*.samplers =*/ sampling.samplers,
/*.n_outputs =*/ n_outputs,
/*.cb =*/ graph_get_cb(),
diff --git a/src/llama-graph.cpp b/src/llama-graph.cpp
index 9b9766da55..0b3bab6123 100644
--- a/src/llama-graph.cpp
+++ b/src/llama-graph.cpp
@@ -1489,7 +1489,7 @@ llm_graph_context::llm_graph_context(const llm_graph_params & params) :
loras (params.loras),
mctx (params.mctx),
cross (params.cross),
- act_policy (params.act_policy),
+ prec_policy (params.prec_policy),
samplers (params.samplers),
cb_func (params.cb),
res (params.res),
@@ -1518,8 +1518,8 @@ ggml_tensor * llm_graph_context::build_lora_mm(
ggml_tensor * w_s) const {
ggml_tensor * res = ggml_mul_mat(ctx0, w, cur);
- if (act_policy) {
- act_policy->apply(res);
+ if (prec_policy) {
+ prec_policy->apply(res);
}
if (w_s) {
@@ -1554,8 +1554,8 @@ ggml_tensor * llm_graph_context::build_lora_mm_id(
ggml_tensor * w_s) const {
ggml_tensor * res = ggml_mul_mat_id(ctx0, w, cur, ids);
- if (act_policy) {
- act_policy->apply(res);
+ if (prec_policy) {
+ prec_policy->apply(res);
}
if (w_s) {
diff --git a/src/llama-graph.h b/src/llama-graph.h
index ef9b00a183..3daa425bc0 100644
--- a/src/llama-graph.h
+++ b/src/llama-graph.h
@@ -19,7 +19,7 @@ struct ggml_tensor;
struct llama_cparams;
struct llama_layer;
-struct llama_act_policy;
+struct llama_prec_policy;
struct llama_memory_context_i;
@@ -788,7 +788,7 @@ struct llm_graph_params {
const llama_memory_context_i * mctx;
const llama_cross * cross;
- const llama_act_policy * act_policy = nullptr;
+ const llama_prec_policy * prec_policy = nullptr;
std::map<llama_seq_id, llama_sampler *> samplers;
@@ -1030,7 +1030,7 @@ struct llm_graph_context {
const llama_memory_context_i * mctx;
const llama_cross * cross;
- const llama_act_policy * act_policy;
+ const llama_prec_policy * prec_policy;
std::map<llama_seq_id, llama_sampler *> samplers;
diff --git a/src/llama-model-saver.cpp b/src/llama-model-saver.cpp
index 5d2518f59e..6f58dd1509 100644
--- a/src/llama-model-saver.cpp
+++ b/src/llama-model-saver.cpp
@@ -194,12 +194,12 @@ void llama_model_saver::add_kv_from_model() {
// add_kv(LLM_KV_GENERAL_SAMPLING_MIROSTAT_ETA, ???);
add_kv(LLM_KV_GENERAL_NAME, model->name);
- if (!model->act_policy.prec_src1.empty()) {
+ if (!model->prec_policy.prec_src1.empty()) {
std::vector<std::string> tensor_names;
std::vector<int8_t> values;
- tensor_names.reserve(model->act_policy.prec_src1.size());
- values.reserve(model->act_policy.prec_src1.size());
- for (const auto & [w, prec] : model->act_policy.prec_src1) {
+ tensor_names.reserve(model->prec_policy.prec_src1.size());
+ values.reserve(model->prec_policy.prec_src1.size());
+ for (const auto & [w, prec] : model->prec_policy.prec_src1) {
tensor_names.push_back(ggml_get_name(w));
values.push_back(prec == GGML_PREC_Q8 ? 0 : 1);
}
diff --git a/src/llama-model.cpp b/src/llama-model.cpp
index 29c3bafcac..c7ce570178 100644
--- a/src/llama-model.cpp
+++ b/src/llama-model.cpp
@@ -1207,7 +1207,7 @@ struct llama_model::impl {
std::vector<float> tensor_split_owned;
};
-bool llama_act_policy::apply(ggml_tensor * res) const {
+bool llama_prec_policy::apply(ggml_tensor * res) const {
if (!res || !res->src[0]) {
return false;
}
@@ -1220,7 +1220,7 @@ bool llama_act_policy::apply(ggml_tensor * res) const {
return ggml_prec_set_src(res, it->second, 1);
}
-static void load_act_policy(llama_model_loader & ml, const llama_model & model, llama_act_policy & policy) {
+void llama_prec_policy::load(llama_model_loader & ml, const llama_model & model) {
std::vector<std::string> tensor_names;
if (!ml.get_arr(LLM_KV_GENERAL_TENSOR_EXTRA_NAME, tensor_names, false)) {
return;
@@ -1252,7 +1252,7 @@ static void load_act_policy(llama_model_loader & ml, const llama_model & model,
// resolve names to tensor pointers
for (const auto & [name, w] : model.tensors_by_name) {
if (want.count(name)) {
- policy.prec_src1.emplace(w, GGML_PREC_Q8);
+ prec_src1.emplace(w, GGML_PREC_Q8);
}
}
}
@@ -1782,7 +1782,7 @@ bool llama_model_base::load_tensors(llama_model_loader & ml) {
}
// per-tensor activation precision policy
- load_act_policy(ml, *this, act_policy);
+ prec_policy.load(ml, *this);
ml.init_mappings(true, use_mlock ? &pimpl->mlock_mmaps : nullptr);
pimpl->mappings.reserve(ml.mappings.size());
diff --git a/src/llama-model.h b/src/llama-model.h
index e9811224c8..a0f9f11423 100644
--- a/src/llama-model.h
+++ b/src/llama-model.h
@@ -17,7 +17,7 @@
struct llama_cparams;
struct llama_ubatch;
struct llama_model_loader;
-struct ggml_tensor;
+struct llama_model;
// available models
enum llm_type {
@@ -610,12 +610,17 @@ struct llama_meta_device_get_split_state_userdata {
struct ggml_backend_meta_split_state llama_meta_device_get_split_state(const struct ggml_tensor * tensor, void * userdata);
-// Per-tensor activation precision from GGUF, unlisted tensors default to native (4-bit) activations.
-struct llama_act_policy {
- // the key is the weight tensor `res->src[0]`
+struct llama_prec_policy {
+ // the key is the weight tensor `res->src[0]`, stores the recommended accumulation type of the op (unused for now)
+ // TODO: migrate ad-hoc ggml_prec_set_acc() calls to this container + update apply() to use it
+ std::unordered_map<const ggml_tensor *, ggml_prec> prec_acc;
+
+ // the key is the weight tensor `res->src[0]`, stores the recommended activation precision type
std::unordered_map<const ggml_tensor *, ggml_prec> prec_src1;
bool apply(ggml_tensor * res) const;
+
+ void load(llama_model_loader & ml, const llama_model & model);
};
struct llama_model {
@@ -628,7 +633,7 @@ struct llama_model {
llama_vocab vocab;
// per-tensor activation precision policy
- llama_act_policy act_policy;
+ llama_prec_policy prec_policy;
// for classifier models
std::vector<std::string> classifier_labels; |
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
fabd334 to
c5f5fb1
Compare
I tried this, but i feel the issue with this will be for the cache like kq and sparse attention score sites . The load time map won't know it and there's no pointer at load time for the cache, so it will be a miss for them. They need to be handled at the call site itself in my opinion. |
Yes, but not all precision policy has to be included in the model file. Some parts of the policy can come from other places. For example from user input or context parameters, etc. Even if we can't migrate all calls to a single place, it's OK. |
Merge upstream commits: - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364) - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294) - metal: split fa kernels into per-dtype libraries (ggml-org#29329) - metal: FWHT kernels for block widths above 512 (ggml-org#29095) - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393) - common: extract shared unicode path/string helpers (ggml-org#29415) - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432) - rpc: include nb in the get_alloc_size cache key (ggml-org#29283) - [SYCL] support sparse FA (ggml-org#28796) - musa: fix PH1 operator failures and build issues (ggml-org#29193) - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231) - opencl: add q5_k bin kernel (ggml-org#29401) - hexagon: add q5_k quant type support (ggml-org#29123) - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404) - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403) - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409) - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422) - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417) Assisted-by: Pi
Conflict: ggml/src/ggml-cuda/mmq.cu - upstream e9f824d (ggml-org#24364) adds a src1 precision policy (GGML_CUDA_MMQ_PREC, op_params slot 3) that selects the native W4A4 path for the FP4 types on Blackwell. Its helpers are taken as-is and prec_src1 is threaded through the fork's ggml_cuda_mul_mat_q_id pair path; on gfx1151 blackwell_mma_available is false, so the result is always GGML_PREC_Q8 and behaviour is unchanged. Upstream's own single-matmul ID body is dropped, as the fork replaces it with mul_mat_q_id.
) * Rebase and update based on ggml-org#26675 Signed-off-by: ynankani <ynankani@nvidia.com> * CI failure fix(launh_bounds overload on HIP) and cleanup Signed-off-by: ynankani <ynankani@nvidia.com> * Address review comments Signed-off-by: ynankani <ynankani@nvidia.com> * Use ggml tensor instead of name in act policy map Signed-off-by: ynankani <ynankani@nvidia.com> * Address review comments and cleanup Signed-off-by: ynankani <ynankani@nvidia.com> * Address review comments Signed-off-by: ynankani <ynankani@nvidia.com> * Rename changes Signed-off-by: ynankani <ynankani@nvidia.com> * Update ggml/src/ggml-cuda/mmq.cu Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * MXFP4 dispatch changes for higher src prec Signed-off-by: ynankani <ynankani@nvidia.com> * Refactor and address review comments Signed-off-by: ynankani <ynankani@nvidia.com> * Updates based on review comments Signed-off-by: ynankani <ynankani@nvidia.com> * Apply batched suggestions from code review Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Address review comments Signed-off-by: ynankani <ynankani@nvidia.com> * Apply patch from review Signed-off-by: ynankani <ynankani@nvidia.com> --------- Signed-off-by: ynankani <ynankani@nvidia.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
|
I experimented with changing a single group to W4A8 while leaving the rest at W4A4, to see the effects based on NVFP4 quant of Qwen3.8 27B:
(KLD against the bf16 base on 200 chunks of wikitext-2 at |
|
The same thing for
UD-Q6_K is only slightly larger (21.98 GB) and like 10 times less divergent, although W4A8 13% ahead by tg128 (74 vs 65 t/s) |
Overview
This PR adds support to force W4A8 path for W4A16_NVFP4 HF model layers on Blackwell, where NVFP4 normally uses the native W4A4 path.
This PR includes the below:
GGML_HINT_NO_QUANT_SRC1hint for NVFP4_W4A16 layersmul_mat_idas wellmul_matfor dense andmul_mat_idMoE modelsAdditional information
Observed a quality improvement for W4A8 compared to W4A4 and are able to meet stricter quality threshold of the test_mul_mat cases used by ggml.
Also it is mentioned that "W4A4 sometimes difficult to achieve for small LLMs" in this paper URL
Tested on : nvidia/Qwen3.6-35B-A3B-NVFP4
Force W4A8 for W4A16_NVFP4 layers:
Master Baseline:
Perplexity Improvement:
Requirements