Skip to content

vulkan: never hand back an empty f16-B matmul pipeline set (fixes prefill SIGSEGV with IQ4_XS Flash-Next files) - #12

Closed
NocFlame wants to merge 1 commit into
Nathanw1014:strix-halo-vulkanfrom
NocFlame:fix/vulkan-f16b-empty-pipeline-set
Closed

NocFlame wants to merge 1 commit into
Nathanw1014:strix-halo-vulkanfrom
NocFlame:fix/vulkan-f16b-empty-pipeline-set

Conversation

@NocFlame

@NocFlame NocFlame commented Oct 6, 2026

Copy link
Copy Markdown

Summary

llama-server built from strix-halo-vulkan at b02cb35 segfaults during prefill with the published
Qwen3.8-Flash-Next IQ4_XS files (julianmb/haloq38flash: Qwen3.8-Flash-Next-IQ4_XS-PLE.gguf plus the
mtp-Qwen3.8-Flash-Next-Q8_0.gguf sidecar) on any prompt that produces a micro-batch of 32 tokens or more.
-ub 8 and -ub 16 work; -ub 32 and the default 2048 crash:

SIGSEGV, fault address 0x4c, in ggml_vk_mul_mat_q_f16 <- ggml_vk_mul_mat <- ggml_vk_build_graph

With GGML_VK_DISABLE_COOPMAT=1 the same prompts abort instead in
ggml_vk_get_mul_mat_mat_id_pipeline: GGML_ASSERT(support_fp32acc).
GGML_VK_DENSE_F16B=0 avoids the crash. The README-verified ad914eb does not crash.

Root cause

GGML_VK_DENSE_F16B instantiates the f16-B GEMM pipelines only for Q4_0, Q4_1, Q5_0, Q5_1, Q8_0, Q2_K, Q3_K,
Q4_K, Q5_K, Q6_K and IQ4_NL (ggml_vk_load_shaders, the if (ggml_vk_dense_f16b_enabled()) block). But with the
conversion enabled (the default since 4b3f6da, "all quantized dense matmuls" since e68dd31)
ggml_vk_get_mul_mat_mat_pipeline returns pipeline_dequant_mul_mat_mat_f16[src0_type] for every quantized
type when src1_type == F16. For a type without f16-B kernels that vk_matmul_pipeline2 is a live object whose
l/m/s are null. The caller sees mmp != nullptr, skips its dequant fallback and dereferences the null pipeline
in ggml_vk_guess_matmul_pipeline_align (->align is at offset 0x4c of vk_pipeline_struct).

The f16 B operand comes from the graph itself: at prefill (nt >= 32) the qwen4exp hyper-connection mix writes
its stream as F16 (mix_type in src/models/qwen4exp.cpp, gated on qwen4exp_takes_f16_b(), which checks for
bf16 weights only), and every consumer of that stream is a matmul B operand: attention and GDN in-projections,
indexer projections, MoE router and experts. In the IQ4_XS files those weights are IQ4_XS (592 of 1224 tensors;
all dense projections). That is the 16-works / 32-crashes threshold. The fork's own recipes use K-quant,
Q4_0/Q5_0/Q8_0 and IQ4_NL tensors, which all have f16-B kernels, so the hole never showed there.

ggml_vk_get_mul_mat_mat_id_pipeline has the same hole behind the GGML_ASSERT(support_fp32acc).

Fix

  • dense getter (KHR coopmat branch): if the selected pipeline set has no kernels, return nullptr so
    ggml_vk_mul_mat_q_f16 takes its dequant + f16 x f16 fallback (exactly what GGML_VK_DENSE_F16B=0 does for
    every matmul); if only one accumulator variant was built, use it.
  • mul_mat_id getter: return nullptr instead of asserting when neither accumulator variant exists.

A follow-up could instantiate the f16-B pipelines for the remaining types (matmul_iq4_xs_f16* is already
compiled by vulkan-shaders-gen; only the CREATE_MM2 list omits them), but the fork's own note says the IQ4_NL
f32-B kernel beats convert + f16-B, so that wants a perplexity and speed check first.

Verification

Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), Mesa 26 RADV, kernel 6.18, KHR coopmat1,
GGML_VK_MAX_MB_PER_SUBMIT=2048 -fa on -ctk f16 -ctv f16 -np 1, MTP sidecar on.

build config result
b02cb35 32k, -ub 2048, default env SIGSEGV on a 169-token prompt
b02cb35 same, GGML_VK_DENSE_F16B=a SIGSEGV
b02cb35 same, GGML_VK_DENSE_F16B=0 OK, 390 tok/s prefill on 4751 tokens
b02cb35 + this patch 32k, -ub 2048, default env OK, 411-419 tok/s prefill on 4751 tokens, 22-24 tok/s generation
b02cb35 + this patch same, GGML_VK_DENSE_F16B=0 OK, 372-381 tok/s
b02cb35 + this patch 256k config (-ub 1024, mmap, lazy) loads at 82.7 GiB GTT; 4751 tok 378 tok/s; 32 401 tok 366 tok/s; 200-tok prose 24-25 tok/s, 78% MTP acceptance
b02cb35 + this patch two-turn tool call via /v1/chat/completions identical to ad914eb
ad914eb (reference) 32k, -ub 2048 255-290 tok/s on 4751 tokens; 205 tok/s at 32k depth

The two scripts used for the table are below. The first starts llama-server with the 32k flags above and
sends a 170-token and a 4700-token prompt; the second is the tool-call round trip.

repro-prefill.sh
#!/bin/sh
# Prefill repro for the empty f16-B pipeline crash. Starts llama-server on port 8089 with the 32k config
# from the PR, sends a ~170-token and a ~4700-token prompt through /completion and prints the timings.
# On unpatched b02cb35 the 170-token prompt already kills the server (SIGSEGV in ggml_vk_mul_mat_q_f16)
# unless GGML_VK_DENSE_F16B=0 is exported; with the patch both prompts complete.
# Usage: MODEL=.../Qwen3.8-Flash-Next-IQ4_XS-PLE.gguf MTP=.../mtp-Qwen3.8-Flash-Next-Q8_0.gguf \
#        BIN=./build/bin/llama-server sh pr-repro-prefill.sh [UB]      (UB defaults to 2048; 16 does not crash)
set -u
BIN=${BIN:-./build/bin/llama-server}; MODEL=${MODEL:?set MODEL}; MTP=${MTP:?set MTP}; UB=${1:-2048}
W=$(mktemp -d); LOG=$W/server.log
python3 - "$W" <<'PY'
import json, sys
W = sys.argv[1]
para = ("The Strix Halo platform combines a 16-core Zen 5 CPU with a 40-compute-unit RDNA 3.5 GPU on a shared 256-bit LPDDR5X memory bus. "
        "Because the memory is unified, the GPU can address large pools of system memory through the GTT, which is what makes running very large language models possible on a laptop. ")
tail = "\n\nSummarize the above in one sentence.\n\n"
for name, reps in (("p170", 2), ("p4700", 60)):
    json.dump({"prompt": para * reps + tail, "n_predict": 24, "temperature": 0, "cache_prompt": False}, open(f"{W}/{name}.json", "w"))
PY
GGML_VK_MAX_MB_PER_SUBMIT=${GGML_VK_MAX_MB_PER_SUBMIT:-2048} "$BIN" -m "$MODEL" -md "$MTP" \
  --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.75 \
  -ngl 999 -fa on -ctk f16 -ctv f16 -c 32768 -np 1 -ub "$UB" -b 2048 -t 16 \
  --cache-ram 8192 --ctx-checkpoints 32 --jinja --alias qwen3.8-flash-next \
  --host 127.0.0.1 --port 8089 > "$LOG" 2>&1 &
P=$!
i=0; while [ $i -lt 60 ]; do i=$((i+1)); sleep 5; curl -s -m 3 http://127.0.0.1:8089/health | grep -q '"ok"' && break; kill -0 $P 2>/dev/null || break; done
for pf in p170 p4700; do
  R=$(curl -s -m 900 -X POST http://127.0.0.1:8089/completion -H 'Content-Type: application/json' --data-binary @"$W/$pf.json")
  if [ -z "$R" ]; then echo "$pf: NO RESPONSE, server alive: $(kill -0 $P 2>/dev/null && echo yes || echo no)"; break; fi
  printf '%s' "$R" | python3 -c 'import sys,json; d=json.load(sys.stdin); t=d["timings"]; print("%s: prompt %d tok at %.0f tok/s, gen %.1f tok/s" % (sys.argv[1], t["prompt_n"], t["prompt_n"]/t["prompt_ms"]*1000, t["predicted_n"]/t["predicted_ms"]*1000))' "$pf"
done
grep -E 'f16-B path engaged|Segmentation|GGML_ASSERT|signal' "$LOG" | head -3
kill $P 2>/dev/null; wait $P 2>/dev/null; echo "server log: $LOG"
repro-toolcall.py
#!/usr/bin/env python3
"""Two-turn tool-call round trip against a llama-server started with --jinja: the model must call read_file,
we answer with the file content, and the model must report the secret word.
Usage: python3 pr-repro-toolcall.py [base_url] [model_alias]"""
import json, sys, time, urllib.request
BASE = sys.argv[1] if len(sys.argv) > 1 else "http://127.0.0.1:8089"
MODEL = sys.argv[2] if len(sys.argv) > 2 else "qwen3.8-flash-next"
TOOLS = [{"type": "function", "function": {"name": "read_file", "description": "Read a text file and return its content",
          "parameters": {"type": "object", "properties": {"path": {"type": "string", "description": "file path"}}, "required": ["path"]}}}]
def chat(messages, max_tokens=2048):
    body = {"model": MODEL, "messages": messages, "tools": TOOLS, "tool_choice": "auto", "temperature": 0, "max_tokens": max_tokens}
    req = urllib.request.Request(BASE + "/v1/chat/completions", data=json.dumps(body).encode(), headers={"Content-Type": "application/json"})
    t0 = time.time(); d = json.load(urllib.request.urlopen(req, timeout=900)); dt = time.time() - t0
    return d["choices"][0], d.get("usage", {}), dt
msgs = [{"role": "user", "content": "Find the secret word in the file notes.txt. Use the read_file tool, then answer with only the secret word."}]
c, u, dt = chat(msgs)
m = c["message"]; tcs = m.get("tool_calls") or []
print("turn 1: finish=%s, %.1f s, tokens in/out %s/%s, tool_calls=%d" % (c.get("finish_reason"), dt, u.get("prompt_tokens"), u.get("completion_tokens"), len(tcs)))
if not tcs:
    print("   NO TOOL CALL; content:", repr((m.get("content") or "")[:200]), "reasoning:", repr((m.get("reasoning_content") or "")[:120])); sys.exit(1)
tc = tcs[0]; print("   call:", tc["function"]["name"], tc["function"]["arguments"])
args = json.loads(tc["function"]["arguments"]); assert "notes" in json.dumps(args), "wrong path"
msgs.append({"role": "assistant", "content": m.get("content") or "", "tool_calls": tcs})
msgs.append({"role": "tool", "tool_call_id": tc.get("id", "call_0"), "name": "read_file", "content": "Project notes.\nThe secret word is PELICAN.\n"})
c, u, dt = chat(msgs)
m = c["message"]; txt = (m.get("content") or "").strip()
print("turn 2: finish=%s, %.1f s, tokens in/out %s/%s, answer=%r" % (c.get("finish_reason"), dt, u.get("prompt_tokens"), u.get("completion_tokens"), txt[:120]))
print("RESULT:", "PASS" if "PELICAN" in txt.upper() else "FAIL")

Related

Nathanw1014/strix-halo-llamacpp#20 is the mul_mat_id side of the same hole on a GPU without cooperative
matrices (the GGML_ASSERT(support_fp32acc) abort); the maintainer's fix in testing there addresses the
missing kernels on non-coopmat devices. This PR closes the dense path on coopmat devices for types that have
no f16-B kernel, and makes both getters degrade to the fallback instead of asserting or crashing.

Workaround for users on b02cb35

GGML_VK_DENSE_F16B=0 in the environment of llama-server.

🤖 Generated with Claude Code

GGML_VK_DENSE_F16B instantiates the f16-B GEMM pipelines for Q4_0..Q6_K and
IQ4_NL only, but ggml_vk_get_mul_mat_mat_pipeline returned
pipeline_dequant_mul_mat_mat_f16[src0_type] for every quantized type once the
conversion was enabled (the default since 4b3f6da / e68dd31). For a type
without f16-B kernels that vk_matmul_pipeline2 is a live object whose l/m/s
are null; the caller saw mmp != nullptr, skipped its dequant fallback and
dereferenced the null pipeline in ggml_vk_guess_matmul_pipeline_align
(SIGSEGV reading ->align, fault address 0x4c).

The f16 B operand comes from the graph: at prefill (nt >= 32) the qwen4exp
hyper-connection mix writes its stream as f16 and every consumer of it is a
matmul B operand. With julianmb's Qwen3.8-Flash-Next IQ4_XS files (592 of
1224 tensors are IQ4_XS, all dense projections) llama-server crashed on any
prompt that produced a micro-batch of 32 tokens or more, while -ub 8 and 16
worked. ggml_vk_get_mul_mat_mat_id_pipeline has the same hole behind
GGML_ASSERT(support_fp32acc), the abort seen with GGML_VK_DISABLE_COOPMAT=1
and reported for a non-coopmat eGPU in strix-halo-llamacpp issue 20.

Return nullptr when the selected pipeline set has no kernels, so the caller
takes its dequant + f16 x f16 fallback (what GGML_VK_DENSE_F16B=0 did for
every matmul), and pick whichever accumulator variant was actually built.

Verified on gfx1151 / RADV (Mesa 26) with the IQ4_XS-PLE file: no crash at
-ub 2048 with the conversion on, 411-419 tok/s prefill on 4751 tokens
(372-381 with it off, 255-290 on ad914eb), 366 tok/s at 32k depth in the
256k configuration, tool calls intact.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Nathanw1014

Copy link
Copy Markdown
Owner

Merged into strix-halo-vulkan as ca00832 and shipped in v0.7.8, with an IQ4_XS f16-B test added on top (4daa536). The release notes credit you for the fallback. Thanks for tracking down the root cause so precisely, that made it quick to land. Closing since it's merged.

@Nathanw1014 Nathanw1014 closed this Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants