You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
HRX: dense on HRX0 with experts on another device fails for Qwen MoE models (fused router) #108
GLM-4.7-Flash Q4_K_M (sigmoid router, no Qwen fusion)
runs, 25.2 tok/s
Binding length. The Qwen router dispatch (dispatch-moe-router.cpp, two places) binds token_count * route_stride route ids. When the experts are on another device, the ids view becomes a graph output whose buffer ends at (token_count - 1) * route_stride + route_count ids. Binding that length removes the error. This is the same bound HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95's MUL_MAT_ID fix uses.
Wrong output. With (1) fixed, the answer is wrong ("Paris 8T inalytics oneer..."). With GGML_HRX_DISABLE_DISPATCH=moe_router, it goes back to "Compute error.", because fused_context_claim claims the router chain only for the fused dispatch. The fused Qwen router seems to assume its consumers (MUL_MAT_ID and the weighted sum) are on HRX0, so the tensors the other device reads are not all materialized.
Not blocking. All-Vulkan beats every split measured (Coder 92.4 tok/s vs 27.0 with the experts on HRX0), so nothing in the engine uses this split today.
Found while measuring the MoE three-way split (docs/moe-streaming.md), llama.cpp pin 96f6b89 / 5556bf2, Strix Halo:
llama-bench -dev HRX0/Vulkan0 -ts 1/0 -ot exps=Vulkan0(dense on HRX0, routed experts on Vulkan0):binding route_ids ... range=[0, 1024) is outside runtime binding length 544dispatch-moe-router.cpp, two places) bindstoken_count * route_strideroute ids. When the experts are on another device, the ids view becomes a graph output whose buffer ends at(token_count - 1) * route_stride + route_countids. Binding that length removes the error. This is the same bound HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95's MUL_MAT_ID fix uses.GGML_HRX_DISABLE_DISPATCH=moe_router, it goes back to "Compute error.", becausefused_context_claimclaims the router chain only for the fused dispatch. The fused Qwen router seems to assume its consumers (MUL_MAT_ID and the weighted sum) are on HRX0, so the tensors the other device reads are not all materialized.Not blocking. All-Vulkan beats every split measured (Coder 92.4 tok/s vs 27.0 with the experts on HRX0), so nothing in the engine uses this split today.
🤖 Generated with Claude Code