Skip to content

serve: --embed and --rerank run on HRX where the build has it - #281

Merged
bong-water-water-bong merged 1 commit into
mainfrom
rag-on-hrx
Oct 2, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
rag-on-hrx

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

The RAG servers started a Vulkan llama-server whatever the build. With HRX (ONEBIT_HRX_SERVER)
they now run on HRX0 with one slot and 2048-token inputs: HRX runs one sequence per batch
(docs/hrx.md), and 2048 is the largest prompt chunk its matmul kernels take in one pass. A build
without HRX keeps Vulkan, 4 slots, 8192 tokens. /v1/models reports the real device.

Checked on Strix Halo against Vulkan0 (HRX build of fork d60cc4f + #62), two repeats each:

  • Qwen3-Embedding-0.6B Q8_0: HRX repeats are bit-identical; cosine to Vulkan 0.99961 / 0.99988 /
    0.99987 on three inputs.
  • bge-reranker-v2-m3 Q8_0: HRX repeats identical; scores 8.609 / -6.756 / -0.401 / -11.020 vs
    Vulkan 8.614 / -6.757 / -0.361 / -11.019, same order.
    jina-reranker-v1-tiny still fails on HRX (docs/serve.md). ctest 19/19.

🤖 Generated with Claude Code

The RAG servers started a Vulkan llama-server whatever the build. With HRX (ONEBIT_HRX_SERVER)
they now run on HRX0 with one slot and 2048-token inputs: HRX runs one sequence per batch
(docs/hrx.md), and 2048 is the largest prompt chunk its matmul kernels take in one pass. A build
without HRX keeps Vulkan, 4 slots, 8192 tokens. /v1/models reports the real device.

Checked on Strix Halo against Vulkan0 (HRX build of fork d60cc4f + #62), two repeats each:
- Qwen3-Embedding-0.6B Q8_0: HRX repeats are bit-identical; cosine to Vulkan 0.99961 / 0.99988 /
  0.99987 on three inputs.
- bge-reranker-v2-m3 Q8_0: HRX repeats identical; scores 8.609 / -6.756 / -0.401 / -11.020 vs
  Vulkan 8.614 / -6.757 / -0.361 / -11.019, same order.
jina-reranker-v1-tiny still fails on HRX (docs/serve.md). ctest 19/19.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Oct 2, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit c5ebb71

@bong-water-water-bong
bong-water-water-bong merged commit ab8691d into main Oct 2, 2026
10 of 11 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the rag-on-hrx branch October 2, 2026 09:10
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
…hat server (#289)

rag_launch (engine #281) started the --embed/--rerank llama-server on HRX0 without the ROCr path the
chat backend gets, so they exited with "invalid device: HRX0" and took 1bit serve down unless the
user had exported IREE_HAL_AMDGPU_LIBHSA_PATH. Found by the power-engineering RAG demo.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant