Skip to content

GPU: pin AMD's HRX pair and upstream ROCmFPX; our GPU kernel work behind -DONEBIT_GPU_PRIVATE - #87

Merged
bong-water-water-bong merged 2 commits into
mainfrom
gpu/private-pins
Sep 25, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
gpu/private-pins

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Our GPU kernel work (ggml-hrx fixes, the HRX prefill split, ROCmI4 on Vulkan) becomes closed source, like the NPU kernels (#85). The engine stops pinning our public forks.

Public pins

Submodule was now
third_party/llama.cpp 1bit-MONSTER/llama.cpp 1bit/hrx-vulkan-patched c075cc1 AMD-Ecosystem/llama.cpp hrx-graph-develop-v2 f1a0aca (AMD's pin, unchanged)
third_party/hrx-system ROCm/hrx-system 51b1739 a351789 (AMD's current pair)
third_party/llama.cpp-rocmfpx 1bit-MONSTER/ROCmFPX 1bit/vulkan-rocmi4 b44bf74 charlie12345/ROCmFPX main fb08d7c

Hook: -DONEBIT_GPU_PRIVATE=<gpu-kernels checkout> includes addons/gpu/gpu.cmake from the private repo, which names the llama.cpp / hrx-system / ROCmFPX trees ONEBIT_HRX and ONEBIT_LEAN build instead. Without it, 1bit serve --prefill-device hrx exits with a message and serve_e2e_vulkan_prefill_hrx is not added.

Workflows: bump-hrx.yml and bump-rocmfpx.yml now only move the public pins (no fork sync, no rebase; HRX_BUMP_TOKEN needs only engine access). The private repo keeps its own pins; a bump flow there is still to do.

Docs: hrx.md (public vs private table, outcome level only; the "Fixed"/"Our patches" internals removed), lean.md, serve.md, PORTING.md, README status, NOTICE.

Verified on Strix Halo (tests/serve_e2e.sh, builds with -DONEBIT_HRX=ON -DONEBIT_LEAN=ON):

Build Qwen3-0.6B Q4_K_M hrx vulkan vulkan --prefill-device hrx Qwen3.6-35B-A3B Q8_0 hrx
public PASS PASS refused (message) FAIL (compute error, known)
-DONEBIT_GPU_PRIVATE PASS PASS PASS PASS

After merge, the fork branches 1bit-MONSTER/llama.cpp 1bit/hrx-vulkan-patched / 1bit/hrx-vulkan and 1bit-MONSTER/ROCmFPX 1bit/vulkan-rocmi4 are no longer used by the engine (private copies are in gpu-kernels).

🤖 Generated with Claude Code

…ind -DONEBIT_GPU_PRIVATE

Our GPU kernel work (ggml-hrx fixes, the HRX prefill split, ROCmI4 on Vulkan) is
closed source, like the NPU kernels. The public build now pins:
- third_party/llama.cpp: AMD-Ecosystem/llama.cpp hrx-graph-develop-v2 f1a0aca and
  third_party/hrx-system a351789, the pair ROCm/ggml-staging-automation tests, unchanged;
- third_party/llama.cpp-rocmfpx: charlie12345/ROCmFPX main fb08d7c.

-DONEBIT_GPU_PRIVATE=<gpu-kernels checkout> includes its addons/gpu/gpu.cmake, which
names the llama.cpp, hrx-system and ROCmFPX trees ONEBIT_HRX and ONEBIT_LEAN build
instead. Without it `1bit serve --prefill-device hrx` exits with a message, and the
serve_e2e_vulkan_prefill_hrx test is not added.

bump-hrx.yml and bump-rocmfpx.yml now only move the public pins (no fork sync or
rebase). Docs keep outcome-level statements of what each build can do.

Verified on Strix Halo (serve_e2e): public: Qwen3-0.6B hrx PASS, vulkan PASS,
--prefill-device hrx refused; Qwen3.6-35B-A3B Q8_0 hrx fails (compute error).
Private: Qwen3-0.6B hrx, vulkan, vulkan --prefill-device hrx PASS; Qwen3.6-35B-A3B
Q8_0 hrx PASS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 5f070bc

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 5f070bc)

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

85 - Partially compliant

Compliant requirements:

  • The Qwen3.6-35B-A3B NPU route is now closed source and moved to a private add-on
  • The engine keeps only a small hook for it
  • Remove the third_party/OpenFlowLM-Next submodule and its NOTICE entry
  • Remove scripts/build-lax.sh
  • Remove npu/lax*
  • Remove app/npu_lax* and the 1bit npu-lax command
  • Remove tests/npu_lax* and tests/golden/npu_lax/
  • Remove docs/npu-lax.md
  • Remove the ONEBIT_NPU_LAX* CMake options
  • Remove the --npu-kernels, --npu-transport and --npu-snapshots flags
  • Update docs to mention the route now has a short note instead
  • Add npu/private_route.{h,cpp}: a registry of private NPU routes keyed by model_type, plus optional 1bit subcommands
  • Add CMake option -DONEBIT_NPU_PRIVATE=<npu-kernels checkout> (needs -DONEBIT_NPU=ON)
  • Add 1bit serve --npu-opt KEY=VALUE (in unified: --opt)
  • Without the add-on, 1bit serve -m <Qwen3.6-35B-A3B dir> --device npu exits with a specific message
  • Add tests/npu_private_route_test.cpp (ctest npu_private_route), added to the CI target list
  • Update docs/npu.md, "Private routes"

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 4 🔵🔵🔵🔵⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Runtime check for private GPU build

The runtime check for the private GPU build (ONEBIT_GPU_PRIVATE) in app/serve.cpp is correctly implemented to prevent usage of --prefill-device hrx without the private build. However, the error message could be more specific about the expected format of the private checkout path to aid in troubleshooting.

#ifndef ONEBIT_GPU_PRIVATE
        if (split)
            throw std::runtime_error("--prefill-device hrx is not part of this build; build with "
                                     "-DONEBIT_GPU_PRIVATE=<gpu-kernels checkout> (docs/hrx.md)");
#endif
CMake logic for private sources

The CMake logic in cmake/hrx.cmake now uses ONEBIT_GPU_LLAMA_SOURCE and ONEBIT_GPU_HRX_SYSTEM_SOURCE variables to determine which source trees to use for building. This change correctly supports the private build, but it's important to ensure that these variables are properly initialized and validated to prevent build errors.

foreach(_src ${ONEBIT_GPU_HRX_SYSTEM_SOURCE} ${ONEBIT_GPU_LLAMA_SOURCE})
    if(NOT EXISTS "${_src}/CMakeLists.txt")
        message(FATAL_ERROR "ONEBIT_HRX needs ${_src}: git submodule update --init --depth 1 in its repository")
    endif()
endforeach()
CMake logic for ROCmFPX source

Similar to cmake/hrx.cmake, cmake/lean.cmake now uses ONEBIT_GPU_ROCMFPX_SOURCE to determine the ROCmFPX source tree. This change correctly supports the private build, but it's important to ensure that this variable is properly initialized and validated to prevent build errors.

# ONEBIT_GPU_ROCMFPX_SOURCE: third_party/llama.cpp-rocmfpx, or the private tree with
# -DONEBIT_GPU_PRIVATE (CMakeLists.txt).
if(NOT EXISTS "${ONEBIT_GPU_ROCMFPX_SOURCE}/CMakeLists.txt")
    message(FATAL_ERROR "ONEBIT_LEAN needs ${ONEBIT_GPU_ROCMFPX_SOURCE}: git submodule update --init --depth 1 in its repository")
endif()
Private GPU build setup

The CMakeLists.txt file correctly sets up the private GPU build by defining ONEBIT_GPU_PRIVATE and including the private gpu.cmake file. However, it's important to ensure that the private build is thoroughly tested to prevent regressions in the public build.

# ── Private GPU trees (docs/hrx.md, "Private GPU build") ──────────────────
# The public pins are AMD's ggml-hrx pair (ONEBIT_HRX) and upstream ROCmFPX (ONEBIT_LEAN).
# ONEBIT_GPU_PRIVATE points at a checkout of the private gpu-kernels repository (branch
# addons/gpu, submodules initialised); its addons/gpu/gpu.cmake names the llama.cpp,
# hrx-system and ROCmFPX trees those builds use instead.
set(ONEBIT_GPU_PRIVATE "" CACHE PATH "A gpu-kernels checkout whose trees ONEBIT_HRX and ONEBIT_LEAN build")
set(ONEBIT_GPU_LLAMA_SOURCE "${CMAKE_SOURCE_DIR}/third_party/llama.cpp")
set(ONEBIT_GPU_HRX_SYSTEM_SOURCE "${CMAKE_SOURCE_DIR}/third_party/hrx-system")
set(ONEBIT_GPU_ROCMFPX_SOURCE "${CMAKE_SOURCE_DIR}/third_party/llama.cpp-rocmfpx")
if(ONEBIT_GPU_PRIVATE)
    if(NOT EXISTS "${ONEBIT_GPU_PRIVATE}/addons/gpu/gpu.cmake")
        message(FATAL_ERROR "ONEBIT_GPU_PRIVATE=${ONEBIT_GPU_PRIVATE} has no addons/gpu/gpu.cmake")
    endif()
    include(${ONEBIT_GPU_PRIVATE}/addons/gpu/gpu.cmake)
    message(STATUS "ONEBIT_GPU_PRIVATE: llama.cpp ${ONEBIT_GPU_LLAMA_SOURCE}, hrx-system ${ONEBIT_GPU_HRX_SYSTEM_SOURCE}, ROCmFPX ${ONEBIT_GPU_ROCMFPX_SOURCE}")
    target_compile_definitions(onebit PRIVATE ONEBIT_GPU_PRIVATE)
endif()

#86's unlinked 1bit-MONSTER)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 5f070bc

@bong-water-water-bong
bong-water-water-bong marked this pull request as ready for review September 25, 2026 16:56
@bong-water-water-bong
bong-water-water-bong merged commit 81f042c into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the gpu/private-pins branch September 25, 2026 16:56
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 5f070bc

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant