Skip to content

HRX: chat on Qwen3-Coder-30B-A3B and GLM-4.7-Flash fails with 500 "Compute error." #95

Description

@bong-water-water-bong

Found by the registry checks (#94, tools/registry_check.py, engine 8c2805d, Strix Halo). It reproduced on a second run.

tests/serve_e2e.sh build/1bit <model> hrx: the server comes up healthy, but the chat request returns HTTP 500 {"error":{"code":500,"message":"Compute error."}}.

Model GGUF arch hrx vulkan
Qwen3-Coder-30B-A3B-Instruct Q4_K_M qwen3moe fails pass
GLM-4.7-Flash Q4_K_M deepseek2 fails pass
Qwen3.6-35B-A3B Q8_0 qwen35moe pass pass

The wiki records Qwen3-Coder-30B-A3B decoding on HRX at 89.4 tok/s in llama-bench after #84. The serve chat path fails anyway. Both failing models are MoE and Q4_K_M, while the passing 35B is Q8_0.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions