Skip to content

hexagon: add q5_k quant type support - #29123

Merged
max-krasnyansky merged 1 commit into
ggml-org:masterfrom
jhen0409:jhen/hexagon-q5k
Sep 25, 2026
Merged

max-krasnyansky merged 1 commit into
ggml-org:masterfrom
jhen0409:jhen/hexagon-q5k

Conversation

@jhen0409

@jhen0409 jhen0409 commented Sep 19, 2026 •

Copy link
Copy Markdown
Member

Overview

Add k-quants q5_k support in ggml-hexagon.

My purpose is to speedup model like unsloth/Qwen3.5-4B-GGUF that included Q5_K (ssm_out) in Q4_0 model. One point worth noting is: If we want npl > 1 cases works well, need to handle #26113 first. (UPDATE: After #29197, the issue should be resolved)

Additional information

Setup: IQ-9075 (Hexagon v73), single NPU HTP0:0, llama-bench -p 512 -n 64 -fa 1 -t 6 --cpu-mask 0xfc -ub 256 -b 512 -r 3. Numbers are pp512 / tg64 t/s. master is 5836771.

Model types master (t/s) + Q5_K (t/s)
Llama-3.2-1B --pure Q4_1 (reference) all Q4_1 1531.0 / 38.0 same
Llama-3.2-1B --pure Q5_K all Q5_K 141.6 / 24.6 (Q5_K on CPU) 1527.2 / 36.3
Llama-3.2-1B Q5_K_M Q5_K + Q6_K 160.7 / 24.0 (Q5_K on CPU) 1478.8 / 37.3
Qwen3.5-4B Q4_0 (unsloth) Q4_0 + Q5_K ssm_out + Q6_K output 212.0 / 11.2 375.9 / 12.1
Old (master b23701f):
Model types master (t/s) + Q5_K (t/s)
Llama-3.2-1B --pure Q4_1 (reference) all Q4_1 1494.4 / 36.5 same
Llama-3.2-1B --pure Q5_K all Q5_K 141.7 / 24.8 (Q5_K on CPU) 1479.6 / 35.0
Llama-3.2-1B Q5_K_M Q5_K + Q6_K 160.2 / 23.8 (Q5_K on CPU) 1433.3 / 36.1
Qwen3.5-4B Q4_0 (unsloth) Q4_0 + Q5_K ssm_out + Q6_K output 123.5 / 10.5 189.8 / 11.3

Got failed to map buffer on -ub 512 randomly with Qwen3.5-4B Q4_0 (master branch, it's fine on the Q5_K impl), so I use 256 as baseline for stability.

Correctness (test-backend-ops test -b HTP0:0, master + Q5_K): MUL_MAT 654/654 (q5_K 29/29), MUL_MAT_ID 399/399 (q5_K 3/3).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, Claude Code, and I did the validation and review.

@jhen0409
jhen0409 requested a review from a team as a code owner September 19, 2026 09:10
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Hexagon labels Sep 19, 2026
@jhen0409
jhen0409 force-pushed the jhen/hexagon-q5k branch 2 times, most recently from 392aef1 to c0c3838 Compare September 21, 2026 12:32
@jhen0409

Copy link
Copy Markdown
Member Author

| Qwen3.5-4B Q4_0 (unsloth) | Q4_0 + Q5_K ssm_out + Q6_K output | 129.8 / 11.2 | 202.3 / 12.2 |

It is actually failed (error: HTP0:0 buffer size 556236800 exceeds max_bufsize 268435456) on master and this PR after #29197, so I use the following workaround to get the test result:

Patch
diff --git a/ggml/src/ggml-hexagon/ggml-hexagon.cpp b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
index ec5a4aeb6..d56fc3dc5 100644
--- a/ggml/src/ggml-hexagon/ggml-hexagon.cpp
+++ b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
@@ -2076,11 +2076,6 @@ static const char * ggml_backend_hexagon_buffer_type_name(ggml_backend_buffer_ty
 static ggml_backend_buffer_t ggml_backend_hexagon_buffer_type_alloc_buffer(
             ggml_backend_buffer_type_t buffer_type, size_t size) {
     auto dev_ctx = static_cast<ggml_backend_hexagon_buffer_type_context *>(buffer_type->context)->dev_ctx;
-    if (size > dev_ctx->max_bufsize) {
-        GGML_LOG_ERROR("ggml-hex: %s buffer size %zu exceeds max_bufsize %zu\n",
-                       dev_ctx->c_name(), size, dev_ctx->max_bufsize);
-        return nullptr;
-    }
     auto sess    = dev_ctx->session();
     if (sess && sess->max_vmem && size > sess->max_vmem) {
         GGML_LOG_ERROR("ggml-hex: %s buffer size %zu exceeds max_vmem %zu\n",
@@ -2088,8 +2083,12 @@ static ggml_backend_buffer_t ggml_backend_hexagon_buffer_type_alloc_buffer(
         return nullptr;
     }
     try {
-        ggml_hexagon_shared_buffer * sbuf = new ggml_hexagon_shared_buffer(sess, size, false);
-        return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_buffer_interface, sbuf, size);
+        auto sbuf = std::make_unique<ggml_hexagon_shared_buffer>(sess, size, false);
+        if (!opt_dma64) {
+            // no extended mappings: map now, large contiguous ranges get scarce later
+            sbuf->mmap();
+        }
+        return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_buffer_interface, sbuf.release(), size);
     } catch (const std::exception & exc) {
         GGML_LOG_ERROR("ggml-hex: %s failed to allocate device buffer context: %s\n", dev_ctx->c_name(), exc.what());
         return nullptr;
@@ -2099,11 +2098,6 @@ static ggml_backend_buffer_t ggml_backend_hexagon_buffer_type_alloc_buffer(
 static ggml_backend_buffer_t ggml_backend_hexagon_host_buffer_type_alloc_buffer(
             ggml_backend_buffer_type_t buffer_type, size_t size) {
     auto dev_ctx = static_cast<ggml_backend_hexagon_buffer_type_context *>(buffer_type->context)->dev_ctx;
-    if (size > dev_ctx->max_bufsize) {
-        GGML_LOG_ERROR("ggml-hex: %s host buffer size %zu exceeds max_bufsize %zu\n",
-                       dev_ctx->c_name(), size, dev_ctx->max_bufsize);
-        return nullptr;
-    }
     auto sess    = dev_ctx->session();
     if (sess && sess->max_vmem && size > sess->max_vmem) {
         GGML_LOG_ERROR("ggml-hex: %s host buffer size %zu exceeds max_vmem %zu\n",
@@ -2111,8 +2105,12 @@ static ggml_backend_buffer_t ggml_backend_hexagon_host_buffer_type_alloc_buffer(
         return nullptr;
     }
     try {
-        ggml_hexagon_shared_buffer * sbuf = new ggml_hexagon_shared_buffer(sess, size, false);
-        return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_host_buffer_interface, sbuf, size);
+        auto sbuf = std::make_unique<ggml_hexagon_shared_buffer>(sess, size, false);
+        if (!opt_dma64) {
+            // no extended mappings: map now, large contiguous ranges get scarce later
+            sbuf->mmap();
+        }
+        return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_host_buffer_interface, sbuf.release(), size);
     } catch (const std::exception & exc) {
         GGML_LOG_ERROR("ggml-hex: %s failed to allocate host buffer context: %s\n", dev_ctx->c_name(), exc.what());
         return nullptr;

CC @max-krasnyansky

@max-krasnyansky

Copy link
Copy Markdown
Member

| Qwen3.5-4B Q4_0 (unsloth) | Q4_0 + Q5_K ssm_out + Q6_K output | 129.8 / 11.2 | 202.3 / 12.2 |

It is actually failed (error: HTP0:0 buffer size 556236800 exceeds max_bufsize 268435456) on master and this PR after #29197, so I use the following workaround to get the test result:
[SNIP]

Hmm. That error looks odd. Are you setting GGML_HEXAGON_MBUF=256 ?
(ie 256MB buffer size). If so that's why you're getting that error.

btw I did realize that the new checks are unnecessary after all.
I'll clean that up in #29199

@jhen0409

jhen0409 commented Sep 22, 2026 •

Copy link
Copy Markdown
Member Author

Hmm. That error looks odd. Are you setting GGML_HEXAGON_MBUF=256 ? (ie 256MB buffer size). If so that's why you're getting that error.

btw I did realize that the new checks are unnecessary after all. I'll clean that up in #29199

After second try, I got fastrpc_mmap failed on default MBUF, the eager mmap patch still useful for Qwen3.5-4B Q4_0 (unsloth). One assumption is that the DSP has already been allocated 2GB+ of space (in fragments) before the output.weight (556MB in the model) load. It should be a separate issue.

I've rebased master again (updated PR description), and confirm that the concerns in #26113 are no longer an issue in this PR (after #29197).

@max-krasnyansky max-krasnyansky left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very clean!

@max-krasnyansky max-krasnyansky added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 25, 2026
@max-krasnyansky

Copy link
Copy Markdown
Member

@lhez second ack please

@max-krasnyansky
max-krasnyansky merged commit 4de0926 into ggml-org:master Sep 25, 2026
19 checks passed
rjtokenring added a commit to rjtokenring/llama.cpp that referenced this pull request Sep 25, 2026
Conflicts with the Q5_K support (ggml-org#29123) and the dynamic quantizer rework (ggml-org#29395):

- matmul-ops.h/.c: keep htp_mm_is_tiled_type and htp_mm_get_tiled_act_row_size; the q8_1 choice
  now uses htp_mm_weight_has_offset, so Q5_K gets q8_1 activations.
- ggml-hexagon.cpp: keep both the F16/BF16/F32 and the Q5_K repack functions.
- QUANTIZE_IMPL now reads the activations from vtcm_act_raw and passes no scratch buffer: the
  fp16-pairs / fp32-splat kernels read the VTCM rows directly (f32 -> f16 through a 64-value
  stack buffer).
- chunked quantization: the F16/F32 tiled weights never take the q8 block quantizer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
feal87 added a commit to feal87/myllama.cpp that referenced this pull request Sep 25, 2026
Merge upstream commits:
  - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364)
  - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294)
  - metal: split fa kernels into per-dtype libraries (ggml-org#29329)
  - metal: FWHT kernels for block widths above 512 (ggml-org#29095)
  - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393)
  - common: extract shared unicode path/string helpers (ggml-org#29415)
  - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432)
  - rpc: include nb in the get_alloc_size cache key (ggml-org#29283)
  - [SYCL] support sparse FA (ggml-org#28796)
  - musa: fix PH1 operator failures and build issues (ggml-org#29193)
  - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231)
  - opencl: add q5_k bin kernel (ggml-org#29401)
  - hexagon: add q5_k quant type support (ggml-org#29123)
  - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404)
  - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403)
  - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409)
  - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422)
  - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417)

Assisted-by: Pi
sky-mighty pushed a commit to sky-mighty/llama.cpp that referenced this pull request Sep 26, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Hexagon merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants