Repository navigation
vulkan: never hand back an empty f16-B matmul pipeline set (fixes prefill SIGSEGV with IQ4_XS Flash-Next files) - #12
Closed
NocFlame wants to merge 1 commit into
Conversation
GGML_VK_DENSE_F16B instantiates the f16-B GEMM pipelines for Q4_0..Q6_K and IQ4_NL only, but ggml_vk_get_mul_mat_mat_pipeline returned pipeline_dequant_mul_mat_mat_f16[src0_type] for every quantized type once the conversion was enabled (the default since 4b3f6da / e68dd31). For a type without f16-B kernels that vk_matmul_pipeline2 is a live object whose l/m/s are null; the caller saw mmp != nullptr, skipped its dequant fallback and dereferenced the null pipeline in ggml_vk_guess_matmul_pipeline_align (SIGSEGV reading ->align, fault address 0x4c). The f16 B operand comes from the graph: at prefill (nt >= 32) the qwen4exp hyper-connection mix writes its stream as f16 and every consumer of it is a matmul B operand. With julianmb's Qwen3.8-Flash-Next IQ4_XS files (592 of 1224 tensors are IQ4_XS, all dense projections) llama-server crashed on any prompt that produced a micro-batch of 32 tokens or more, while -ub 8 and 16 worked. ggml_vk_get_mul_mat_mat_id_pipeline has the same hole behind GGML_ASSERT(support_fp32acc), the abort seen with GGML_VK_DISABLE_COOPMAT=1 and reported for a non-coopmat eGPU in strix-halo-llamacpp issue 20. Return nullptr when the selected pipeline set has no kernels, so the caller takes its dequant + f16 x f16 fallback (what GGML_VK_DENSE_F16B=0 did for every matmul), and pick whichever accumulator variant was actually built. Verified on gfx1151 / RADV (Mesa 26) with the IQ4_XS-PLE file: no crash at -ub 2048 with the conversion on, 411-419 tok/s prefill on 4751 tokens (372-381 with it off, 255-290 on ad914eb), 366 tok/s at 32k depth in the 256k configuration, tool calls intact. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Owner
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
llama-serverbuilt fromstrix-halo-vulkanat b02cb35 segfaults during prefill with the publishedQwen3.8-Flash-Next IQ4_XS files (julianmb/haloq38flash:
Qwen3.8-Flash-Next-IQ4_XS-PLE.ggufplus themtp-Qwen3.8-Flash-Next-Q8_0.ggufsidecar) on any prompt that produces a micro-batch of 32 tokens or more.-ub 8and-ub 16work;-ub 32and the default 2048 crash:With
GGML_VK_DISABLE_COOPMAT=1the same prompts abort instead inggml_vk_get_mul_mat_mat_id_pipeline:GGML_ASSERT(support_fp32acc).GGML_VK_DENSE_F16B=0avoids the crash. The README-verified ad914eb does not crash.Root cause
GGML_VK_DENSE_F16Binstantiates the f16-B GEMM pipelines only for Q4_0, Q4_1, Q5_0, Q5_1, Q8_0, Q2_K, Q3_K,Q4_K, Q5_K, Q6_K and IQ4_NL (
ggml_vk_load_shaders, theif (ggml_vk_dense_f16b_enabled())block). But with theconversion enabled (the default since 4b3f6da, "all quantized dense matmuls" since e68dd31)
ggml_vk_get_mul_mat_mat_pipelinereturnspipeline_dequant_mul_mat_mat_f16[src0_type]for every quantizedtype when
src1_type == F16. For a type without f16-B kernels thatvk_matmul_pipeline2is a live object whosel/m/sare null. The caller seesmmp != nullptr, skips its dequant fallback and dereferences the null pipelinein
ggml_vk_guess_matmul_pipeline_align(->alignis at offset 0x4c ofvk_pipeline_struct).The f16 B operand comes from the graph itself: at prefill (
nt >= 32) the qwen4exp hyper-connection mix writesits stream as F16 (
mix_typeinsrc/models/qwen4exp.cpp, gated onqwen4exp_takes_f16_b(), which checks forbf16 weights only), and every consumer of that stream is a matmul B operand: attention and GDN in-projections,
indexer projections, MoE router and experts. In the IQ4_XS files those weights are IQ4_XS (592 of 1224 tensors;
all dense projections). That is the 16-works / 32-crashes threshold. The fork's own recipes use K-quant,
Q4_0/Q5_0/Q8_0 and IQ4_NL tensors, which all have f16-B kernels, so the hole never showed there.
ggml_vk_get_mul_mat_mat_id_pipelinehas the same hole behind theGGML_ASSERT(support_fp32acc).Fix
nullptrsoggml_vk_mul_mat_q_f16takes its dequant + f16 x f16 fallback (exactly whatGGML_VK_DENSE_F16B=0does forevery matmul); if only one accumulator variant was built, use it.
nullptrinstead of asserting when neither accumulator variant exists.A follow-up could instantiate the f16-B pipelines for the remaining types (
matmul_iq4_xs_f16*is alreadycompiled by vulkan-shaders-gen; only the CREATE_MM2 list omits them), but the fork's own note says the IQ4_NL
f32-B kernel beats convert + f16-B, so that wants a perplexity and speed check first.
Verification
Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), Mesa 26 RADV, kernel 6.18, KHR coopmat1,
GGML_VK_MAX_MB_PER_SUBMIT=2048 -fa on -ctk f16 -ctv f16 -np 1, MTP sidecar on.GGML_VK_DENSE_F16B=aGGML_VK_DENSE_F16B=0GGML_VK_DENSE_F16B=0The two scripts used for the table are below. The first starts
llama-serverwith the 32k flags above andsends a 170-token and a 4700-token prompt; the second is the tool-call round trip.
repro-prefill.sh
repro-toolcall.py
Related
Nathanw1014/strix-halo-llamacpp#20 is the mul_mat_id side of the same hole on a GPU without cooperative
matrices (the
GGML_ASSERT(support_fp32acc)abort); the maintainer's fix in testing there addresses themissing kernels on non-coopmat devices. This PR closes the dense path on coopmat devices for types that have
no f16-B kernel, and makes both getters degrade to the fallback instead of asserting or crashing.
Workaround for users on b02cb35
GGML_VK_DENSE_F16B=0in the environment ofllama-server.🤖 Generated with Claude Code