The code is in our llama.cpp fork (1bit-MONSTER/llama.cpp, 1bit/hrx-vulkan-patched, ggml/src/ggml-hrx), which has issues disabled.
Symptom. gpt-oss-20b (MXFP4) on HRX0 with the expert MUL_MAT_IDs on the CPU while SWIGLU_OAI (fork #67) runs on HRX:
- greedy text is garbage ("The first part (the first 2? maybe 2^? ...");
- wikitext KLD vs CPU is 1.65, same top 12% (8 chunks, c512, -b 512).
Repro: GGML_HRX_DISABLE_DISPATCH=mul_mat_id.f32_f32_wmma. The user-facing trigger is any MUL_MAT_ID shape HRX declines, e.g. an ubatch over 2048 tokens.
| config (all with mul_mat_id.f32_f32_wmma disabled) |
result |
| ADD_ID and SWIGLU_OAI on HRX |
KLD 1.65, garbage |
+ swiglu_oai disabled (SWIGLU_OAI on CPU, ADD_ID on HRX) |
KLD 0.029, correct |
+ add_id disabled |
correct text (KLD run killed by the thermal guard) |
Default gpt-oss (MUL_MAT_ID on HRX) is fine: KLD 0.028, correct text.
Ruled out
GGML_HRX_DEBUG_SERIAL_EXECUTION=1: still wrong.
- Graph program cache (
GGML_HRX_GRAPH_PROGRAM_CACHE=0): still wrong.
- Another dispatch claiming the GLU: the dispatch log shows
extra.swiglu_oai_f32 in both configs.
- The kernel math: per-op dumps with
llama-eval-callback show SWIGLU_OAI-0's values identical whether it runs on HRX or the CPU, and layers 0-11 match to < 1e-5 between the bad and good configs. The eval callback makes the scheduler compute op by op, though, so it hides whatever goes wrong across multi-op HRX splits.
- Not reproducible in isolation:
tests/test-hrx-moe-split.cpp builds gpt-oss's MoE block under ggml_backend_sched (HRX + CPU) with CPU MXFP4 experts, HRX ADD_ID/SWIGLU_OAI and HRX-resident argsort-view ids. It matches the CPU at 2880/32 experts/4 used with 12/128/512 tokens (NMSE 1.6e-4) with SWIGLU_OAI on HRX or not.
Working hypothesis: a lifetime or synchronization problem for tensors crossing several CPU/HRX split boundaries in the full model. In the bad config, SWIGLU_OAI reads the gate ADD_ID output that an earlier HRX split produced, across a CPU split (the up MUL_MAT_ID), and its own output feeds a CPU split. This could hit any model whose graph alternates CPU and HRX splits.
Mitigation: a placement guard in the fork (moe-placement-guard.cpp, PR linked below) claims ADD_ID / SWIGLU_OAI only when their MUL_MAT_ID is HRX-supported, so the failing placement can't happen.
Next probe: force a stream sync at every split boundary (with serial execution). If that fixes it, it's a sync bug; otherwise compare buffer lifetimes across splits.
🤖 Generated with Claude Code
The code is in our llama.cpp fork (1bit-MONSTER/llama.cpp,
1bit/hrx-vulkan-patched,ggml/src/ggml-hrx), which has issues disabled.Symptom. gpt-oss-20b (MXFP4) on HRX0 with the expert MUL_MAT_IDs on the CPU while SWIGLU_OAI (fork #67) runs on HRX:
Repro:
GGML_HRX_DISABLE_DISPATCH=mul_mat_id.f32_f32_wmma. The user-facing trigger is any MUL_MAT_ID shape HRX declines, e.g. an ubatch over 2048 tokens.swiglu_oaidisabled (SWIGLU_OAI on CPU, ADD_ID on HRX)add_iddisabledDefault gpt-oss (MUL_MAT_ID on HRX) is fine: KLD 0.028, correct text.
Ruled out
GGML_HRX_DEBUG_SERIAL_EXECUTION=1: still wrong.GGML_HRX_GRAPH_PROGRAM_CACHE=0): still wrong.extra.swiglu_oai_f32in both configs.llama-eval-callbackshow SWIGLU_OAI-0's values identical whether it runs on HRX or the CPU, and layers 0-11 match to < 1e-5 between the bad and good configs. The eval callback makes the scheduler compute op by op, though, so it hides whatever goes wrong across multi-op HRX splits.tests/test-hrx-moe-split.cppbuilds gpt-oss's MoE block underggml_backend_sched(HRX + CPU) with CPU MXFP4 experts, HRX ADD_ID/SWIGLU_OAI and HRX-resident argsort-view ids. It matches the CPU at 2880/32 experts/4 used with 12/128/512 tokens (NMSE 1.6e-4) with SWIGLU_OAI on HRX or not.Working hypothesis: a lifetime or synchronization problem for tensors crossing several CPU/HRX split boundaries in the full model. In the bad config, SWIGLU_OAI reads the gate ADD_ID output that an earlier HRX split produced, across a CPU split (the up MUL_MAT_ID), and its own output feeds a CPU split. This could hit any model whose graph alternates CPU and HRX splits.
Mitigation: a placement guard in the fork (
moe-placement-guard.cpp, PR linked below) claims ADD_ID / SWIGLU_OAI only when their MUL_MAT_ID is HRX-supported, so the failing placement can't happen.Next probe: force a stream sync at every split boundary (with serial execution). If that fixes it, it's a sync bug; otherwise compare buffer lifetimes across splits.
🤖 Generated with Claude Code