Repository navigation
hexagon: add q5_k quant type support - #29123
Conversation
392aef1 to
c0c3838
Compare
It is actually failed (error: Patchdiff --git a/ggml/src/ggml-hexagon/ggml-hexagon.cpp b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
index ec5a4aeb6..d56fc3dc5 100644
--- a/ggml/src/ggml-hexagon/ggml-hexagon.cpp
+++ b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
@@ -2076,11 +2076,6 @@ static const char * ggml_backend_hexagon_buffer_type_name(ggml_backend_buffer_ty
static ggml_backend_buffer_t ggml_backend_hexagon_buffer_type_alloc_buffer(
ggml_backend_buffer_type_t buffer_type, size_t size) {
auto dev_ctx = static_cast<ggml_backend_hexagon_buffer_type_context *>(buffer_type->context)->dev_ctx;
- if (size > dev_ctx->max_bufsize) {
- GGML_LOG_ERROR("ggml-hex: %s buffer size %zu exceeds max_bufsize %zu\n",
- dev_ctx->c_name(), size, dev_ctx->max_bufsize);
- return nullptr;
- }
auto sess = dev_ctx->session();
if (sess && sess->max_vmem && size > sess->max_vmem) {
GGML_LOG_ERROR("ggml-hex: %s buffer size %zu exceeds max_vmem %zu\n",
@@ -2088,8 +2083,12 @@ static ggml_backend_buffer_t ggml_backend_hexagon_buffer_type_alloc_buffer(
return nullptr;
}
try {
- ggml_hexagon_shared_buffer * sbuf = new ggml_hexagon_shared_buffer(sess, size, false);
- return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_buffer_interface, sbuf, size);
+ auto sbuf = std::make_unique<ggml_hexagon_shared_buffer>(sess, size, false);
+ if (!opt_dma64) {
+ // no extended mappings: map now, large contiguous ranges get scarce later
+ sbuf->mmap();
+ }
+ return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_buffer_interface, sbuf.release(), size);
} catch (const std::exception & exc) {
GGML_LOG_ERROR("ggml-hex: %s failed to allocate device buffer context: %s\n", dev_ctx->c_name(), exc.what());
return nullptr;
@@ -2099,11 +2098,6 @@ static ggml_backend_buffer_t ggml_backend_hexagon_buffer_type_alloc_buffer(
static ggml_backend_buffer_t ggml_backend_hexagon_host_buffer_type_alloc_buffer(
ggml_backend_buffer_type_t buffer_type, size_t size) {
auto dev_ctx = static_cast<ggml_backend_hexagon_buffer_type_context *>(buffer_type->context)->dev_ctx;
- if (size > dev_ctx->max_bufsize) {
- GGML_LOG_ERROR("ggml-hex: %s host buffer size %zu exceeds max_bufsize %zu\n",
- dev_ctx->c_name(), size, dev_ctx->max_bufsize);
- return nullptr;
- }
auto sess = dev_ctx->session();
if (sess && sess->max_vmem && size > sess->max_vmem) {
GGML_LOG_ERROR("ggml-hex: %s host buffer size %zu exceeds max_vmem %zu\n",
@@ -2111,8 +2105,12 @@ static ggml_backend_buffer_t ggml_backend_hexagon_host_buffer_type_alloc_buffer(
return nullptr;
}
try {
- ggml_hexagon_shared_buffer * sbuf = new ggml_hexagon_shared_buffer(sess, size, false);
- return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_host_buffer_interface, sbuf, size);
+ auto sbuf = std::make_unique<ggml_hexagon_shared_buffer>(sess, size, false);
+ if (!opt_dma64) {
+ // no extended mappings: map now, large contiguous ranges get scarce later
+ sbuf->mmap();
+ }
+ return ggml_backend_buffer_init(buffer_type, ggml_backend_hexagon_host_buffer_interface, sbuf.release(), size);
} catch (const std::exception & exc) {
GGML_LOG_ERROR("ggml-hex: %s failed to allocate host buffer context: %s\n", dev_ctx->c_name(), exc.what());
return nullptr; |
Hmm. That error looks odd. Are you setting btw I did realize that the new checks are unnecessary after all. |
c0c3838 to
574f89c
Compare
After second try, I got I've rebased master again (updated PR description), and confirm that the concerns in #26113 are no longer an issue in this PR (after #29197). |
574f89c to
4d6a710
Compare
|
@lhez second ack please |
Conflicts with the Q5_K support (ggml-org#29123) and the dynamic quantizer rework (ggml-org#29395): - matmul-ops.h/.c: keep htp_mm_is_tiled_type and htp_mm_get_tiled_act_row_size; the q8_1 choice now uses htp_mm_weight_has_offset, so Q5_K gets q8_1 activations. - ggml-hexagon.cpp: keep both the F16/BF16/F32 and the Q5_K repack functions. - QUANTIZE_IMPL now reads the activations from vtcm_act_raw and passes no scratch buffer: the fp16-pairs / fp32-splat kernels read the VTCM rows directly (f32 -> f16 through a 64-value stack buffer). - chunked quantization: the F16/F32 tiled weights never take the q8 block quantizer. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge upstream commits: - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364) - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294) - metal: split fa kernels into per-dtype libraries (ggml-org#29329) - metal: FWHT kernels for block widths above 512 (ggml-org#29095) - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393) - common: extract shared unicode path/string helpers (ggml-org#29415) - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432) - rpc: include nb in the get_alloc_size cache key (ggml-org#29283) - [SYCL] support sparse FA (ggml-org#28796) - musa: fix PH1 operator failures and build issues (ggml-org#29193) - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231) - opencl: add q5_k bin kernel (ggml-org#29401) - hexagon: add q5_k quant type support (ggml-org#29123) - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404) - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403) - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409) - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422) - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417) Assisted-by: Pi
(cherry picked from commit 4de0926)
Overview
Add k-quants q5_k support in ggml-hexagon.
My purpose is to speedup model like unsloth/Qwen3.5-4B-GGUF that included Q5_K (
ssm_out) in Q4_0 model. One point worth noting is: If we want npl > 1 cases works well, need to handle #26113 first. (UPDATE: After #29197, the issue should be resolved)Additional information
Setup: IQ-9075 (Hexagon v73), single NPU
HTP0:0,llama-bench -p 512 -n 64 -fa 1 -t 6 --cpu-mask 0xfc -ub 256 -b 512 -r 3. Numbers are pp512 / tg64 t/s. master is 5836771.--pureQ4_1 (reference)--pureQ5_Kssm_out+ Q6_KoutputOld (master b23701f):
--pureQ4_1 (reference)--pureQ5_Kssm_out+ Q6_KoutputCorrectness (
test-backend-ops test -b HTP0:0, master + Q5_K): MUL_MAT 654/654 (q5_K 29/29), MUL_MAT_ID 399/399 (q5_K 3/3).Requirements