Repository navigation
serve: --embed and --rerank run on HRX where the build has it - #281
Merged
Merged
Conversation
The RAG servers started a Vulkan llama-server whatever the build. With HRX (ONEBIT_HRX_SERVER) they now run on HRX0 with one slot and 2048-token inputs: HRX runs one sequence per batch (docs/hrx.md), and 2048 is the largest prompt chunk its matmul kernels take in one pass. A build without HRX keeps Vulkan, 4 slots, 8192 tokens. /v1/models reports the real device. Checked on Strix Halo against Vulkan0 (HRX build of fork d60cc4f + #62), two repeats each: - Qwen3-Embedding-0.6B Q8_0: HRX repeats are bit-identical; cosine to Vulkan 0.99961 / 0.99988 / 0.99987 on three inputs. - bge-reranker-v2-m3 Q8_0: HRX repeats identical; scores 8.609 / -6.756 / -0.401 / -11.020 vs Vulkan 8.614 / -6.757 / -0.361 / -11.019, same order. jina-reranker-v1-tiny still fails on HRX (docs/serve.md). ctest 19/19. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Docs7 for 1bit-monster/engine
Commit |
bong-water-water-bong
added a commit
that referenced
this pull request
Oct 2, 2026
…hat server (#289) rag_launch (engine #281) started the --embed/--rerank llama-server on HRX0 without the ROCr path the chat backend gets, so they exited with "invalid device: HRX0" and took 1bit serve down unless the user had exported IREE_HAL_AMDGPU_LIBHSA_PATH. Found by the power-engineering RAG demo. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The RAG servers started a Vulkan llama-server whatever the build. With HRX (ONEBIT_HRX_SERVER)
they now run on HRX0 with one slot and 2048-token inputs: HRX runs one sequence per batch
(docs/hrx.md), and 2048 is the largest prompt chunk its matmul kernels take in one pass. A build
without HRX keeps Vulkan, 4 slots, 8192 tokens. /v1/models reports the real device.
Checked on Strix Halo against Vulkan0 (HRX build of fork d60cc4f + #62), two repeats each:
0.99987 on three inputs.
Vulkan 8.614 / -6.757 / -0.361 / -11.019, same order.
jina-reranker-v1-tiny still fails on HRX (docs/serve.md). ctest 19/19.
🤖 Generated with Claude Code