Skip to content

Force NVFP4 W4A8 path for NVFP4_W4A16 layers on Blackwell, where NVFP4 normally uses the native W4A4 path. - #24364

Merged
ggerganov merged 14 commits into
ggml-org:masterfrom
ynankani:ynankani/Force_W4A16_NVFP4_to_W4A8
Sep 25, 2026
Merged

ggerganov merged 14 commits into
ggml-org:masterfrom
ynankani:ynankani/Force_W4A16_NVFP4_to_W4A8

Conversation

@ynankani

@ynankani ynankani commented Jun 9, 2026 •

Copy link
Copy Markdown
Contributor

Overview

This PR adds support to force W4A8 path for W4A16_NVFP4 HF model layers on Blackwell, where NVFP4 normally uses the native W4A4 path.
This PR includes the below:

  1. Adds a GGUF metadata for storing NVFP4_W4A16 layers, output weight.
  2. Loads the NVFP4_W4A16 layer metadata in model hparams .
  3. Use the metadata to set the new GGML_HINT_NO_QUANT_SRC1 hint for NVFP4_W4A16 layers
  4. On Blackwell GPU, dispatch the hinted layers through W4A8 path instead of native W4A4 path
  5. Allow the new hint for the MoE layer mul_mat_id as well
  6. Test case for mul_mat for dense and mul_mat_id MoE models

Additional information

Observed a quality improvement for W4A8 compared to W4A4 and are able to meet stricter quality threshold of the test_mul_mat cases used by ggml.
Also it is mentioned that "W4A4 sometimes difficult to achieve for small LLMs" in this paper URL
Tested on : nvidia/Qwen3.6-35B-A3B-NVFP4

Force W4A8 for W4A16_NVFP4 layers:

Model Size Params Backend ngl Test t/s
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 pp512 4492.20 +/- 25.66
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 pp1024 4689.26 +/- 47.87
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 pp2048 4822.40 +/- 75.42
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 tg128 148.25 +/- 1.42
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 tg256 164.13 +/- 6.55

Master Baseline:

Model Size Params Backend ngl Test t/s
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 pp512 5486.02 +/- 253.52
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 pp1024 5445.48 +/- 211.56
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 pp2048 5865.13 +/- 171.27
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 tg128 155.60 +/- 6.73
qwen35moe 35B.A3B NVFP4 21.09 GiB 35.51 B CUDA 99 tg256 165.52 +/- 6.55

Perplexity Improvement:

Config Command PPL
Force W4A8 for W4A16_NVFP4 layers llama-perplexity -c 2048 -b 2048 -ngl 99 6.0264 +/- 0.03777
Master baseline llama-perplexity -c 2048 -b 2048 -ngl 99 6.1234 +/- 0.03851

Requirements

@ynankani
ynankani requested review from a team, CISC and ggerganov as code owners June 9, 2026 14:32
@github-actions github-actions Bot added testing Everything test related Nvidia GPU Issues specific to Nvidia GPUs python python script changes ggml changes relating to the ggml tensor library for machine learning labels Jun 9, 2026
Comment thread ggml/src/ggml-cuda/mmq.cu Outdated
mul_mat_q_case<GGML_TYPE_NVFP4, true>(ctx, args, stream);
break;
}
#endif // GGML_CUDA_HAS_BLACKWELL_TARGET

@sanmai sanmai Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How do we opt out from this? A drop from 5486.02 to 4492.20 is very severe.

If you want a higher precision, there's a variety of Q4 quants that are just as small and even more precise (see #23572 for detailed comparisons).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no need to opt out from this, if you want to run W4A4 take a checkpoint with W4A4.
The intention of this PR is if a checkpoint has W4A16 layers it should have activations in higher precision.
This PR doesn't cause regression on pure W4A4_NVFP4 checkpoints.

You can confirm there is no regression by testing llama-bench with this PR on below checkpoints:

  1. Pure W4A4_NVFP4: https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-NVFP4
  2. W4A16_NVFP4: https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There may be a case that there aren't other checkpoints or somebody likes one's calibration over another. I think many prefer increased speed and that is why they pick NVFP4. Usually Q4_K~ will always be better precision than NVFP4 if that is what they are going for and then it may end up faster than skipping the native FP4. We're blending checkpoints with Q_K quants and NVFP4 combined that can compensate for ppl loss. I think some selection control by the user would be a good balance to let them decide.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ModelOpt, Redhat came up with W4A16_NVFP4 after through investigations on performance and accuracy. Checkout this weight-only-quantization-schemes.

We should honor the intention behind creating a recipe/checkpoint which was specifically designed for W4A16_NVFP4.
If someone intends to use W4A4 on Blackwell they should use a W4A4 checkpoint.

Few PR on vLLM for reference, that were added by ModelOpt to support W4A16_NVFP4 :

  1. Mixed precision W4A16_NVFP4 : PR
  2. W4A16 linear support : PR

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think some selection control by the user would be a good balance to let them decide.

I mean we can in principle add such a knob in the CUDA backend to give the ability to override model builder intents encoded in the GGUF, but the default should be what is encoded in the GGUF

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added the knob for user controlled W4A4_NVFP4 fast path on Blackwell, even if the checkpoint specifies W4A16_NVFP4 layers

@am17an am17an left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't look into the PR in detail, but does using 8-bit activation disable the fp4 tensor core?

@ORippler

Copy link
Copy Markdown
Collaborator

I didn't look into the PR in detail, but does using 8-bit activation disable the fp4 tensor core?

Yes. Basically you can think of this as a step towards the support of weight-only-quantization-schemes in llama.cpp.

@ynankani
ynankani force-pushed the ynankani/Force_W4A16_NVFP4_to_W4A8 branch from b72a8c9 to 18f1df3 Compare June 17, 2026 12:26
@github-actions github-actions Bot added documentation Improvements or additions to documentation CUDA Related to the CUDA backend labels Jun 17, 2026
@ggerganov ggerganov self-assigned this Jun 25, 2026
@ORippler ORippler self-assigned this Jul 8, 2026

@ORippler ORippler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@ggerganov @CISC thoughts on this approach and the required granularity for W4A4 vs. W4A16?

Comment thread docs/build.md Outdated
Comment thread ggml/src/ggml-cuda/CMakeLists.txt Outdated
Comment on lines +134 to +136
if (CMAKE_CUDA_ARCHITECTURES MATCHES "(^|;)12[0-9]a(-real|-virtual)?($|;)")
add_compile_definitions(GGML_CUDA_HAS_BLACKWELL_TARGET)
endif()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just FYI this compile definition will be visible for all archs in CMAKE_CUDA_ARCHITECTURES (per arch specialization requires constructing nvcc commands by hand)

Comment thread conversion/base.py Outdated
Comment on lines +830 to +839
if self._is_nvfp4:
for tensor_name, entry in quant_layers.items():
if not isinstance(entry, dict) or entry.get("quant_algo") != "W4A16_NVFP4":
continue
if "lm_head" in tensor_name or "output" in tensor_name:
self._nvfp4_w4a16_output = True
continue
bid_m = re.search(r'\.layers\.(\d+)\.', tensor_name)
if bid_m:
self._nvfp4_w4a16_blocks.add(int(bid_m.group(1)))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4/blob/main/config.json#L455-L470

huggingface/modelopt seems to offer storing this information on a per-weight level, as opposed to per-transformer-block we parse here. Later on, we set the marker also on a per-weight/op level. I feel this is inconsistent and we should adapt the conversion script to support the finer level of granularity also

@ggerganov ggerganov left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I read correctly, this introduces a path to do w4a8 where we normally do w4a4. And the reason is that w4a8 has higher accuracy?

Is there a case where we would want to use w4a4? If no, we can probably simplify a lot of the logic.

Yet, if there is specific tensor hardware for that, it's probably useful.

Comment thread src/llama-hparams.h Outdated
Comment on lines +164 to +167
// NVFP4 W4A16: per-layer + LM head flag, true where NVFP4 weights skip activation quantization.
std::array<bool, LLAMA_MAX_LAYERS> nvfp4_w4a16_layer_arr = {};
bool nvfp4_w4a16_output = false;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This does not feel like it belongs to hparams. It's too low-level, backend-specific information.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we make this as a per-tensor attribute in llama-model at load time? This will also help acknowledge @ORippler's comment as well to have more flexibility for per-tensor information.

https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4/blob/main/config.json#L455-L470
huggingface/modelopt seems to offer storing this information on a per-weight level, as opposed to per-transformer-block we parse here. Later on, we set the marker also on a per-weight/op level. I feel this is inconsistent and we should adapt the conversion script to support the finer level of granularity also

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, we should formulate this information to be per tensor. Also, it should abstract away the NVFP4 and make it more generic. For example, "allow 4-bit activations (bool)" seems generic enough.

We can represent this information with a map: tensor name -> bool. The values are false by default. The map is optional. To represent the map in GGUF, you'll need 2 arrays - the first one with the tensor names and the second one with the same size and the bool values.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"allow 4-bit activations (bool)"

I intuitively think about this PR/feature as "enabling Weight-only-Quantization schemes", so naming should signal this intent if you agree -> "weight-only quantization (bool)" / "allow activation quantization (bool)".

W4A16 is effectively weight-only, as LLMs are trained in BF16 precision typically

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In some backends, we already quantize the activations by default (to 8-bits). So I'm not sure if it is not going to be a bit misleading to call it "weight-only quantization (bool)" / "allow activation quantization (bool)".

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated the change so allow-activation-quant is stored per-tensor. Since it is currently NVFP4-specific, activation quantization is enabled by default for all layers. For layers that use W4A16_NVFP4 as the quantization method, the activation-quant array stores the tensor name paired with false. Please review.

@ORippler

ORippler commented Jul 8, 2026

Copy link
Copy Markdown
Collaborator

If I read correctly, this introduces a path to do w4a8 where we normally do w4a4. And the reason is that w4a8 has higher accuracy?

Yes on both. The motivation is that doing W4A16 (or weight-only-quantization in general) may allow one to quantize more layers than when one quantizes both activations and weights. This reduces memory-footprint of the model, and furthermore increases decode throughput - decode is effectively weight-streaming in local inference with small BS setting)

@ggerganov

Copy link
Copy Markdown
Member

Correct me if I am wrong, but I think in the CUDA backend we already do w4a8 for existing types like Q4 - we do it by default. If that's the case, adding support for w4a8+nvfp4 would make sense only if it is better quality-wise than the existing w4a8+q4 methods.

@ORippler

ORippler commented Jul 8, 2026

Copy link
Copy Markdown
Collaborator

If that's the case, adding support for w4a8+nvfp4 would make sense only if it is better quality-wise than the existing w4a8+q4 methods.

I think that's going to be difficult to prove generally, though I can add the single model we have validated so far:

model source size params effective BPW NVFP4 mode
qwen35moe 35B.A3B Q4_K - Medium unsloth/Qwen3.6-35B-A3B-GGUF 20.60 GiB 34.66 B 5.105 n/a
qwen35moe 35B.A3B NVFP4 nvidia/Qwen3.6-35B-A3B-NVFP4 19.51 GiB 34.66 B 4.835 W4A16

@ggerganov

Copy link
Copy Markdown
Member

For my understanding, the nvidia/Qwen3.6-35B-A3B-NVFP4 readme says that the model was quantized with the Model Optimizer. Does the quantization process involve some form of quantization aware training?

My understanding is that if a model is trained natively in NVFP4 format, the best thing to do is keep the format intact during inference. I.e. it's not beneficial to use Q4 quantizations in such case on hardware that supports NVFP4. But I would assume that if the model was trained in NVFP4 in the first place, then the recommended way to run would be w4a4 (otherwise it would mean that 4-bit tensor cores weren't utilized during the training?).

If the model was not trained in NVFP4, and instead was quantized to NVFP4, then we have to justify and demonstrate in which cases it is worth it compared to the Q4 formats. One simple justification is that w4a4 is going to be utilized for faster speed. But with w4a8 we are throwing this away, so it's no longer obvious what is the benefit of NVFP4+w4a8.

It might be worth taking a look at the KLD between these quantizations and a reference BF16 model. For example, let's take a look at KLD and speed for: Q4_K_S, Q4_K_M, NVFP4+wa4a, NVFP4+w4a8.

@ORippler

Copy link
Copy Markdown
Collaborator

For my understanding, the nvidia/Qwen3.6-35B-A3B-NVFP4 readme says that the model was quantized with the Model Optimizer. Does the quantization process involve some form of quantization aware training?

No, training was involved, the checkpoint is yielded by PTQ only. QAT/QAD would be disclosed in the model card like this https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4#training-methodology (this is a W4A4 checkpoint).

If the model was not trained in NVFP4, and instead was quantized to NVFP4, then we have to justify and demonstrate in which cases it is worth it compared to the Q4 formats [...]. it's no longer obvious what is the benefit of NVFP4+w4a8.

As outlined above, the proposed benefit of weight-only quantization is to allow the quantization of weights more aggressive at iso-quality. We will try to collect & present some numbers.

In general, the idea would be to go for NVFP4 with W4A16, the W4A8 path is just a transient step.

@ynankani
ynankani force-pushed the ynankani/Force_W4A16_NVFP4_to_W4A8 branch from 18f1df3 to 2c7052a Compare July 14, 2026 12:04
@ggerganov

Copy link
Copy Markdown
Member

In general, the idea would be to go for NVFP4 with W4A16, the W4A8 path is just a transient step.

For my understanding, is the expectation that W4A16 would be faster compared to W4A8 if implemented efficiently?

Comment thread docs/build.md Outdated
Comment on lines +278 to +280
#### GGML_CUDA_FORCE_W4A4

NVFP4 models that carry W4A16 layers request higher-precision activations, so on Blackwell those layers run through the W4A8 path instead of the native W4A4 path. Set `GGML_CUDA_DISABLE_FORCE_W4A8=1` to ignore that request and keep the native W4A4 path for faster prompt processing at the cost of accuracy.
NVFP4 models that carry W4A16 layers request higher-precision activations (W4A8), so on Blackwell those layers run through the W4A8 path instead of the native W4A4 path. Set `GGML_CUDA_FORCE_W4A4=1` to override that request and keep the native W4A4 path for faster prompt processing at the cost of accuracy.

@ggerganov ggerganov Jul 15, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At some point we can remove this compile-time option and toggle this functionality at runtime through a libllama argument. It will simply skip setting the quantization hints for the matrix multiplications (i.e. override the activation policy). The idea is to avoid "communicating" directly with backend.

Comment thread ggml/include/ggml.h Outdated
@@ -443,7 +443,7 @@ extern "C" {
enum ggml_op_hint {
GGML_HINT_NONE = 0,
GGML_HINT_SRC0_IS_HADAMARD = 1,
GGML_HINT_NO_QUANT_SRC1 = 2, // W4A16_NVFP4: keep activations higher precision.
GGML_HINT_NO_QUANT_SRC1 = 2, // keep src1 at higher precision (skip 4-bit activation quant).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we agree on the "allow 4-bit activations", then we should consistently name where relevant. For example here: GGML_HINT_SRC1_ALLOW_4BIT.

Comment thread gguf-py/gguf/constants.py Outdated
Comment on lines +28 to +30
# Per-tensor activation precision policy (tensor name -> allow 4-bit activations).
ALLOW_4BIT_ACT_TENSOR = "general.allow_4bit_act.tensor"
ALLOW_4BIT_ACT_VALUE = "general.allow_4bit_act.value"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This can become more generic and future-proof to allow setting additional per-tensor options in the future. To do that, the array with the tensor names should be called something like "general.tensor_extra.name". And the "allow 4-bit activations" values should be stored in the bool array named "general.tensor_extra.allow_4bit_act".

Comment thread src/llama-graph.cpp Outdated
Comment on lines +1395 to +1398
if (no_quant_src1_for_weight(act_policy, w)) {
ggml_mul_mat_set_hint(res, GGML_HINT_NO_QUANT_SRC1);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should be:

Suggested change
if (no_quant_src1_for_weight(act_policy, w)) {
ggml_mul_mat_set_hint(res, GGML_HINT_NO_QUANT_SRC1);
}
if (llama_act_policy_allow_4bit(act_policy, w)) {
ggml_mul_mat_set_hint(res, GGML_HINT_SRC1_ALLOW_4BIT);
}

Comment thread src/llama-graph.h Outdated
@@ -18,6 +18,7 @@ struct ggml_tensor;

struct llama_cparams;
struct llama_layer;
struct llama_weight_act_policy;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
struct llama_weight_act_policy;
struct llama_act_policy;

@JohannesGaessler

JohannesGaessler commented Jul 16, 2026 •

Copy link
Copy Markdown
Contributor

My opinion is that we should solve this differently at a ggml level. As of right now we have this for "precision":

    // precision
    enum ggml_prec {
        GGML_PREC_DEFAULT =  0, // stored as ggml_tensor.op_params, 0 by default
        GGML_PREC_F32     = 10,
    };

I am interpreting "default" to mean that we don't really care about the precision and that the backends should just optimize for speed/memory use. So on Blackwell that would mean W4A4. But we can add something like GGML_PREC_A8 or GGML_PREC_A16 to indicate that we want the backend to use at least that number of bits for the activations. I think this would be preferable over making the distinction an "op hint" because my interpretation of what that should be is a hint to the backend regarding what optimizations could be done without any negative side effects by exploiting that the data has some special structure. And it's also not clear to me what a backend is supposed to do for the combination of GGML_PREC_F32 and GGML_HINT_SRC1_ALLOW_4BIT.

To be clear: I think that the way we currently define "precision" in ggml is poorly defined and needs more maintainer attention. My opinion is that GGML_PREC_F32 should require strict FP32 arithmetic, not just the numerical range of FP32. For that we should add a new value.

@ynankani

Copy link
Copy Markdown
Contributor Author

My opinion is that we should solve this differently at a ggml level. As of right now we have this for "precision":

    // precision
    enum ggml_prec {
        GGML_PREC_DEFAULT =  0, // stored as ggml_tensor.op_params, 0 by default
        GGML_PREC_F32     = 10,
    };

I am interpreting "default" to mean that we don't really care about the precision and that the backends should just optimize for speed/memory use. So on Blackwell that would mean W4A4. But we can add something like GGML_PREC_A8 or GGML_PREC_A16 to indicate that we want the backend to use at least that number of bits for the activations. I think this would be preferable over making the distinction an "op hint" because my interpretation of what that should be is a hint to the backend regarding what optimizations could be done without any negative side effects by exploiting that the data has some special structure. And it's also not clear to me what a backend is supposed to do for the combination of GGML_PREC_F32 and GGML_HINT_SRC1_ALLOW_4BIT.

To be clear: I think that the way we currently define "precision" in ggml is poorly defined and needs more maintainer attention. My opinion is that GGML_PREC_F32 should require strict FP32 arithmetic, not just the numerical range of FP32. For that we should add a new value.

This makes sense to me. It would reduce the ambiguity created by having multiple precision-related signals and hints.

@jeffbolznv

Copy link
Copy Markdown
Contributor

I agree it would be nice to have more clear instructions to the backend about what precision to use, but who makes that choice and how? And I think we need to be more clear about weight/activation precision vs accumulator precision/range.

@ynankani
ynankani force-pushed the ynankani/Force_W4A16_NVFP4_to_W4A8 branch from 0a1701e to a470fa6 Compare July 16, 2026 15:36
@ynankani
ynankani force-pushed the ynankani/Force_W4A16_NVFP4_to_W4A8 branch from 7d987b6 to fabd334 Compare September 20, 2026 10:28
@ynankani

Copy link
Copy Markdown
Contributor Author

As @ORippler is out this week and I have addressed his comments . Can we mark this PR merge-ready?

@ynankani
ynankani requested a review from ORippler September 22, 2026 14:40
@ggerganov

Copy link
Copy Markdown
Member

@ynankani I can't push to your branch - could you apply this patch:

diff --git a/src/llama-context.cpp b/src/llama-context.cpp
index e4c218d67b..795803b022 100644
--- a/src/llama-context.cpp
+++ b/src/llama-context.cpp
@@ -2559,7 +2559,7 @@ llm_graph_params llama_context::graph_params(
         /*.loras       =*/ loras.get(),
         /*.mctx        =*/ mctx,
         /*.cross       =*/ &cross,
-        /*.act_policy  =*/ &model.act_policy,
+        /*.prec_policy =*/ &model.prec_policy,
         /*.samplers    =*/ sampling.samplers,
         /*.n_outputs   =*/ n_outputs,
         /*.cb          =*/ graph_get_cb(),
diff --git a/src/llama-graph.cpp b/src/llama-graph.cpp
index 9b9766da55..0b3bab6123 100644
--- a/src/llama-graph.cpp
+++ b/src/llama-graph.cpp
@@ -1489,7 +1489,7 @@ llm_graph_context::llm_graph_context(const llm_graph_params & params) :
     loras            (params.loras),
     mctx             (params.mctx),
     cross            (params.cross),
-    act_policy       (params.act_policy),
+    prec_policy      (params.prec_policy),
     samplers         (params.samplers),
     cb_func          (params.cb),
     res              (params.res),
@@ -1518,8 +1518,8 @@ ggml_tensor * llm_graph_context::build_lora_mm(
           ggml_tensor * w_s) const {
     ggml_tensor * res = ggml_mul_mat(ctx0, w, cur);

-    if (act_policy) {
-        act_policy->apply(res);
+    if (prec_policy) {
+        prec_policy->apply(res);
     }

     if (w_s) {
@@ -1554,8 +1554,8 @@ ggml_tensor * llm_graph_context::build_lora_mm_id(
           ggml_tensor * w_s) const {
     ggml_tensor * res = ggml_mul_mat_id(ctx0, w, cur, ids);

-    if (act_policy) {
-        act_policy->apply(res);
+    if (prec_policy) {
+        prec_policy->apply(res);
     }

     if (w_s) {
diff --git a/src/llama-graph.h b/src/llama-graph.h
index ef9b00a183..3daa425bc0 100644
--- a/src/llama-graph.h
+++ b/src/llama-graph.h
@@ -19,7 +19,7 @@ struct ggml_tensor;

 struct llama_cparams;
 struct llama_layer;
-struct llama_act_policy;
+struct llama_prec_policy;

 struct llama_memory_context_i;

@@ -788,7 +788,7 @@ struct llm_graph_params {
     const llama_memory_context_i * mctx;
     const llama_cross            * cross;

-    const llama_act_policy * act_policy = nullptr;
+    const llama_prec_policy * prec_policy = nullptr;

     std::map<llama_seq_id, llama_sampler *> samplers;

@@ -1030,7 +1030,7 @@ struct llm_graph_context {
     const llama_memory_context_i * mctx;
     const llama_cross            * cross;

-    const llama_act_policy * act_policy;
+    const llama_prec_policy * prec_policy;

     std::map<llama_seq_id, llama_sampler *> samplers;

diff --git a/src/llama-model-saver.cpp b/src/llama-model-saver.cpp
index 5d2518f59e..6f58dd1509 100644
--- a/src/llama-model-saver.cpp
+++ b/src/llama-model-saver.cpp
@@ -194,12 +194,12 @@ void llama_model_saver::add_kv_from_model() {
     // add_kv(LLM_KV_GENERAL_SAMPLING_MIROSTAT_ETA,     ???);
     add_kv(LLM_KV_GENERAL_NAME,                      model->name);

-    if (!model->act_policy.prec_src1.empty()) {
+    if (!model->prec_policy.prec_src1.empty()) {
         std::vector<std::string> tensor_names;
         std::vector<int8_t> values;
-        tensor_names.reserve(model->act_policy.prec_src1.size());
-        values.reserve(model->act_policy.prec_src1.size());
-        for (const auto & [w, prec] : model->act_policy.prec_src1) {
+        tensor_names.reserve(model->prec_policy.prec_src1.size());
+        values.reserve(model->prec_policy.prec_src1.size());
+        for (const auto & [w, prec] : model->prec_policy.prec_src1) {
             tensor_names.push_back(ggml_get_name(w));
             values.push_back(prec == GGML_PREC_Q8 ? 0 : 1);
         }
diff --git a/src/llama-model.cpp b/src/llama-model.cpp
index 29c3bafcac..c7ce570178 100644
--- a/src/llama-model.cpp
+++ b/src/llama-model.cpp
@@ -1207,7 +1207,7 @@ struct llama_model::impl {
     std::vector<float> tensor_split_owned;
 };

-bool llama_act_policy::apply(ggml_tensor * res) const {
+bool llama_prec_policy::apply(ggml_tensor * res) const {
     if (!res || !res->src[0]) {
         return false;
     }
@@ -1220,7 +1220,7 @@ bool llama_act_policy::apply(ggml_tensor * res) const {
     return ggml_prec_set_src(res, it->second, 1);
 }

-static void load_act_policy(llama_model_loader & ml, const llama_model & model, llama_act_policy & policy) {
+void llama_prec_policy::load(llama_model_loader & ml, const llama_model & model) {
     std::vector<std::string> tensor_names;
     if (!ml.get_arr(LLM_KV_GENERAL_TENSOR_EXTRA_NAME, tensor_names, false)) {
         return;
@@ -1252,7 +1252,7 @@ static void load_act_policy(llama_model_loader & ml, const llama_model & model,
     // resolve names to tensor pointers
     for (const auto & [name, w] : model.tensors_by_name) {
         if (want.count(name)) {
-            policy.prec_src1.emplace(w, GGML_PREC_Q8);
+            prec_src1.emplace(w, GGML_PREC_Q8);
         }
     }
 }
@@ -1782,7 +1782,7 @@ bool llama_model_base::load_tensors(llama_model_loader & ml) {
     }

     // per-tensor activation precision policy
-    load_act_policy(ml, *this, act_policy);
+    prec_policy.load(ml, *this);

     ml.init_mappings(true, use_mlock ? &pimpl->mlock_mmaps : nullptr);
     pimpl->mappings.reserve(ml.mappings.size());
diff --git a/src/llama-model.h b/src/llama-model.h
index e9811224c8..a0f9f11423 100644
--- a/src/llama-model.h
+++ b/src/llama-model.h
@@ -17,7 +17,7 @@
 struct llama_cparams;
 struct llama_ubatch;
 struct llama_model_loader;
-struct ggml_tensor;
+struct llama_model;

 // available models
 enum llm_type {
@@ -610,12 +610,17 @@ struct llama_meta_device_get_split_state_userdata {

 struct ggml_backend_meta_split_state llama_meta_device_get_split_state(const struct ggml_tensor * tensor, void * userdata);

-// Per-tensor activation precision from GGUF, unlisted tensors default to native (4-bit) activations.
-struct llama_act_policy {
-    // the key is the weight tensor `res->src[0]`
+struct llama_prec_policy {
+    // the key is the weight tensor `res->src[0]`, stores the recommended accumulation type of the op (unused for now)
+    // TODO: migrate ad-hoc ggml_prec_set_acc() calls to this container + update apply() to use it
+    std::unordered_map<const ggml_tensor *, ggml_prec> prec_acc;
+
+    // the key is the weight tensor `res->src[0]`, stores the recommended activation precision type
     std::unordered_map<const ggml_tensor *, ggml_prec> prec_src1;

     bool apply(ggml_tensor * res) const;
+
+    void load(llama_model_loader & ml, const llama_model & model);
 };

 struct llama_model {
@@ -628,7 +633,7 @@ struct llama_model {
     llama_vocab   vocab;

     // per-tensor activation precision policy
-    llama_act_policy act_policy;
+    llama_prec_policy prec_policy;

     // for classifier models
     std::vector<std::string> classifier_labels;

ynankani and others added 14 commits September 25, 2026 11:00
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: ynankani <ynankani@nvidia.com>
@ynankani
ynankani force-pushed the ynankani/Force_W4A16_NVFP4_to_W4A8 branch from fabd334 to c5f5fb1 Compare September 25, 2026 05:42
@ynankani

ynankani commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author
  • // the key is the weight tensor res->src[0], stores the recommended accumulation type of the op (unused for now)
  • // TODO: migrate ad-hoc ggml_prec_set_acc() calls to this container + update apply() to use it

I tried this, but i feel the issue with this will be for the cache like kq and sparse attention score sites . The load time map won't know it and there's no pointer at load time for the cache, so it will be a miss for them. They need to be handled at the call site itself in my opinion.

@ggerganov

Copy link
Copy Markdown
Member
  • // the key is the weight tensor res->src[0], stores the recommended accumulation type of the op (unused for now)
  • // TODO: migrate ad-hoc ggml_prec_set_acc() calls to this container + update apply() to use it

I tried this, but i feel the issue with this will be for the cache like kq and sparse attention score sites . The load time map won't know it and there's no pointer at load time for the cache, so it will be a miss for them. They need to be handled at the call site itself in opinion.

Yes, but not all precision policy has to be included in the model file. Some parts of the policy can come from other places. For example from user input or context parameters, etc. Even if we can't migrate all calls to a single place, it's OK.

@ggerganov
ggerganov merged commit e9f824d into ggml-org:master Sep 25, 2026
21 of 22 checks passed
feal87 added a commit to feal87/myllama.cpp that referenced this pull request Sep 25, 2026
Merge upstream commits:
  - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364)
  - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294)
  - metal: split fa kernels into per-dtype libraries (ggml-org#29329)
  - metal: FWHT kernels for block widths above 512 (ggml-org#29095)
  - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393)
  - common: extract shared unicode path/string helpers (ggml-org#29415)
  - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432)
  - rpc: include nb in the get_alloc_size cache key (ggml-org#29283)
  - [SYCL] support sparse FA (ggml-org#28796)
  - musa: fix PH1 operator failures and build issues (ggml-org#29193)
  - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231)
  - opencl: add q5_k bin kernel (ggml-org#29401)
  - hexagon: add q5_k quant type support (ggml-org#29123)
  - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404)
  - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403)
  - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409)
  - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422)
  - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417)

Assisted-by: Pi
Patt92 pushed a commit to Patt92/llama.cpp that referenced this pull request Sep 25, 2026
Conflict: ggml/src/ggml-cuda/mmq.cu - upstream e9f824d (ggml-org#24364) adds a src1
precision policy (GGML_CUDA_MMQ_PREC, op_params slot 3) that selects the native
W4A4 path for the FP4 types on Blackwell. Its helpers are taken as-is and
prec_src1 is threaded through the fork's ggml_cuda_mul_mat_q_id pair path; on
gfx1151 blackwell_mma_available is false, so the result is always GGML_PREC_Q8
and behaviour is unchanged. Upstream's own single-matmul ID body is dropped, as
the fork replaces it with mul_mat_q_id.
sky-mighty pushed a commit to sky-mighty/llama.cpp that referenced this pull request Sep 26, 2026
)

* Rebase and update based on ggml-org#26675

Signed-off-by: ynankani <ynankani@nvidia.com>

* CI failure fix(launh_bounds overload on HIP) and cleanup

Signed-off-by: ynankani <ynankani@nvidia.com>

* Address review comments

Signed-off-by: ynankani <ynankani@nvidia.com>

* Use ggml tensor instead of name in act policy map

Signed-off-by: ynankani <ynankani@nvidia.com>

* Address review comments and cleanup

Signed-off-by: ynankani <ynankani@nvidia.com>

* Address review comments

Signed-off-by: ynankani <ynankani@nvidia.com>

* Rename changes

Signed-off-by: ynankani <ynankani@nvidia.com>

* Update ggml/src/ggml-cuda/mmq.cu

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* MXFP4 dispatch changes for higher src prec

Signed-off-by: ynankani <ynankani@nvidia.com>

* Refactor and address review comments

Signed-off-by: ynankani <ynankani@nvidia.com>

* Updates based on review comments

Signed-off-by: ynankani <ynankani@nvidia.com>

* Apply batched suggestions from code review

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* Address review comments

Signed-off-by: ynankani <ynankani@nvidia.com>

* Apply patch from review

Signed-off-by: ynankani <ynankani@nvidia.com>

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
@sanmai

sanmai commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

I experimented with changing a single group to W4A8 while leaving the rest at W4A4, to see the effects based on NVFP4 quant of Qwen3.8 27B:

group %% in W4A8 KLD mean same top p pp4096
all W4A4 0% 0.064207 88.788% 6084
all W4A8 100% 0.040625 90.984% 3263
Q4_0 with imatrix - 0.029215 92.671% 4030
Q4_0 without imatrix - 0.036372 91.669% 4030
ffn_up 23.4% 0.058936 89.353% 5068
ffn_gate 23.4% 0.059741 89.125% 5082
ffn_down 23.4% 0.060563 89.171% 5005
attn_qkv 10.3% 0.063120 89.027% 5555
ssm_out 6.2% 0.062797 88.780% 5780
attn_gate 6.2% 0.064263 88.873% 5795
attn_q 4.1% 0.063950 88.867% 5907
attn_output 2.1% 0.063779 89.061% 6004
attn_k 0.3% 0.064468 88.833% 6070
attn_v 0.3% 0.064579 88.804% 6073
ssm_alpha 0.05% 0.064505 88.710% 6063
ssm_beta 0.05% 0.065078 88.620% 6078

(KLD against the bf16 base on 200 chunks of wikitext-2 at -c 512 -b 512 -ub 512, f16 KV cache)

@sanmai

sanmai commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

The same thing for nvidia/Qwen3.8-27B-NVFP4 checkpoint, converted with --outtype bf16 --fp8-as-q8 to 21.47 GB

group %% in W4A8 KLD mean same top p pp4096
all W4A4 0% 0.044101 89.931% 4959
all W4A8 100% 0.021883 92.749% 3346
ffn_up 31.0% 0.037943 90.373% 4326
ffn_down 31.0% 0.038591 90.327% 4280
ffn_gate 31.0% 0.039179 90.384% 4345
output 6.9% 0.038861 90.792% 5031
float activations (-b 8 -ub 8) - 0.021740 92.698% -
Q4_0 with imatrix - 0.029215 92.671% 4030
UD-Q6_K - 0.002071 98.024% 3354

UD-Q6_K is only slightly larger (21.98 GB) and like 10 times less divergent, although W4A8 13% ahead by tg128 (74 vs 65 t/s)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion CUDA Related to the CUDA backend documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning model Model specific mtmd Related to multimodal functionality (video/image/audio) Nvidia GPU Issues specific to Nvidia GPUs python python script changes testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants