Repository navigation
1bit serve: one model on any device behind one OpenAI-compatible API - #21
Merged
Merged
Conversation
The engine as a backend a host launches (geramyL's point: embed the engine into Lemonade, not Lemonade into the engine). One model per process: /health (+/v1/health) 503 -> 200, /v1/models, /v1/chat/completions (SSE or not), /v1/completions. - NPU model directory -> the NPU fast lane in process (unified.cpp, which now also answers /health). - .gguf -> this build's llama-server (Vulkan0 or HRX0) or ZINC as a private loopback child; OpenAI routes forwarded, streaming relayed as it arrives, replies carry the served name, and ZINC gets requests without 'model'. - HRX: sets IREE_HAL_AMDGPU_LIBHSA_PATH to TheRock's libhsa itself (--hrx-libhsa, the build's copy, or /opt/rocm-therock). - tests/serve_e2e.sh + ctest serve_e2e_<device> (-DONEBIT_SERVE_TEST_GGUF). Verified on Strix Halo with Qwen3-0.6B Q4_K_M: vulkan, hrx and zinc all PASS (health, models, 'Paris.' under the served name, SSE streaming). The NPU route is not yet run through serve_e2e. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The e2e test left llama-server/zinc children behind: SIGTERM ended serve without its destructor, and nothing could clean up after SIGKILL. The child is now forked with PR_SET_PDEATHSIG (the kernel ends it when serve dies, SIGKILL included), and SIGTERM/SIGINT stop the server and the child cleanly. Verified on Strix Halo: after SIGTERM and after SIGKILL of serve, the child is gone; serve_e2e vulkan still passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on the NPU serve passes --alias through to the in-process NPU server (unified now takes --alias), and the NPU route honours chat_template_kwargs.enable_thinking=false the way Qwen3's template does (an empty think block after the assistant prefix), so every device answers alike. serve_e2e npu PASS on Strix Halo (Qwen3-0.6B NPU model dir: 'Paris' under the served name, 33 SSE chunks). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
enabled auto-merge (squash)
September 23, 2026 22:34
bong-water-water-bong
added a commit
that referenced
this pull request
Sep 26, 2026
…pp-vulkan 546800c) (#161) Fork PR #21: ONEBIT_MOE_SUBST=r lets a resident expert among the router's next choices stand in for a missing one scoring at most 1/r times higher, in the same host sync as the remap. On Qwen3-Coder-30B, r 0.5 cut drive reads 20-27% for a KL divergence of 0.008-0.012, and decode gained 10-35% at 1536-3072 slots. Off by default. moe/ExpertCache gains resident(), as in the fork's copy. Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First step of the inversion geramyL proposed: the engine is a backend Lemonade launches, not a host that embeds Lemonade. It exposes only an OpenAI-compatible API.
It serves one model per process, which is the contract Lemonade launches backends with:
/healthand/v1/health: 503 while loading, then 200/v1/models/v1/chat/completions, streamed (SSE) or not/v1/completions.ggufwith--device vulkanorhrx.ggufwith--device zincmodel.servepointsIREE_HAL_AMDGPU_LIBHSA_PATHat TheRock's libhsa itself.auto= Vulkan for GGUF until Laya routes per request.Verified on Strix Halo (Qwen3-0.6B Q4_K_M,
tests/serve_e2e.sh)All four devices pass, the NPU included. The NPU route now also takes
--aliasand honourschat_template_kwargs.enable_thinking=false, like the GPU backends. After a SIGTERM or SIGKILL ofserve, no backend is left running (the child gets PR_SET_PDEATHSIG).Next, in separate PRs: remove the embedded Lemonade (
third_party/lemonadeand its deltas) from the engine, and open the upstream Lemonade PR, a recipe that launches1bit serve.🤖 Generated with Claude Code