Skip to content

HRX: wrong values when SWIGLU_OAI runs on HRX after a CPU-placed MUL_MAT_ID in a multi-split graph; hidden by per-op sync #286

Description

@bong-water-water-bong

The code is in our llama.cpp fork (1bit-MONSTER/llama.cpp, 1bit/hrx-vulkan-patched, ggml/src/ggml-hrx), which has issues disabled.

Symptom. gpt-oss-20b (MXFP4) on HRX0 with the expert MUL_MAT_IDs on the CPU while SWIGLU_OAI (fork #67) runs on HRX:

  • greedy text is garbage ("The first part (the first 2? maybe 2^? ...");
  • wikitext KLD vs CPU is 1.65, same top 12% (8 chunks, c512, -b 512).

Repro: GGML_HRX_DISABLE_DISPATCH=mul_mat_id.f32_f32_wmma. The user-facing trigger is any MUL_MAT_ID shape HRX declines, e.g. an ubatch over 2048 tokens.

config (all with mul_mat_id.f32_f32_wmma disabled) result
ADD_ID and SWIGLU_OAI on HRX KLD 1.65, garbage
+ swiglu_oai disabled (SWIGLU_OAI on CPU, ADD_ID on HRX) KLD 0.029, correct
+ add_id disabled correct text (KLD run killed by the thermal guard)

Default gpt-oss (MUL_MAT_ID on HRX) is fine: KLD 0.028, correct text.

Ruled out

  • GGML_HRX_DEBUG_SERIAL_EXECUTION=1: still wrong.
  • Graph program cache (GGML_HRX_GRAPH_PROGRAM_CACHE=0): still wrong.
  • Another dispatch claiming the GLU: the dispatch log shows extra.swiglu_oai_f32 in both configs.
  • The kernel math: per-op dumps with llama-eval-callback show SWIGLU_OAI-0's values identical whether it runs on HRX or the CPU, and layers 0-11 match to < 1e-5 between the bad and good configs. The eval callback makes the scheduler compute op by op, though, so it hides whatever goes wrong across multi-op HRX splits.
  • Not reproducible in isolation: tests/test-hrx-moe-split.cpp builds gpt-oss's MoE block under ggml_backend_sched (HRX + CPU) with CPU MXFP4 experts, HRX ADD_ID/SWIGLU_OAI and HRX-resident argsort-view ids. It matches the CPU at 2880/32 experts/4 used with 12/128/512 tokens (NMSE 1.6e-4) with SWIGLU_OAI on HRX or not.

Working hypothesis: a lifetime or synchronization problem for tensors crossing several CPU/HRX split boundaries in the full model. In the bad config, SWIGLU_OAI reads the gate ADD_ID output that an earlier HRX split produced, across a CPU split (the up MUL_MAT_ID), and its own output feeds a CPU split. This could hit any model whose graph alternates CPU and HRX splits.

Mitigation: a placement guard in the fork (moe-placement-guard.cpp, PR linked below) claims ADD_ID / SWIGLU_OAI only when their MUL_MAT_ID is HRX-supported, so the failing placement can't happen.

Next probe: force a stream sync at every split boundary (with serial execution). If that fixes it, it's a sync bug; otherwise compare buffer lifetimes across splits.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions