Skip to content

1bit serve: one model on any device behind one OpenAI-compatible API - #21

Merged
bong-water-water-bong merged 3 commits into
mainfrom
serve
Sep 23, 2026
Merged

bong-water-water-bong merged 3 commits into
mainfrom
serve

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

First step of the inversion geramyL proposed: the engine is a backend Lemonade launches, not a host that embeds Lemonade. It exposes only an OpenAI-compatible API.

1bit serve -m <model> [--port 8000] [--device auto|npu|vulkan|hrx|zinc] [--ctx-size N] [--alias NAME]

It serves one model per process, which is the contract Lemonade launches backends with:

  • /health and /v1/health: 503 while loading, then 200
  • /v1/models
  • /v1/chat/completions, streamed (SSE) or not
  • /v1/completions
Model Runs on
NPU model directory the NPU fast lane, in process
.gguf with --device vulkan or hrx this build's llama-server (Vulkan0 or HRX0), as a private child
.gguf with --device zinc this build's ZINC (Vulkan, ROCm or CUDA), as a private child
  • For a child, the OpenAI routes are forwarded and streaming is relayed as it arrives. Replies carry the served name, and ZINC gets requests without model.
  • HRX needs no setup. serve points IREE_HAL_AMDGPU_LIBHSA_PATH at TheRock's libhsa itself.
  • auto = Vulkan for GGUF until Laya routes per request.

Verified on Strix Halo (Qwen3-0.6B Q4_K_M, tests/serve_e2e.sh)

Device health models "Paris." under served name streaming
npu ok ok ok 33 chunks
vulkan ok ok ok 32 chunks
hrx ok ok ok 32 chunks
zinc ok ok ok 6 chunks

All four devices pass, the NPU included. The NPU route now also takes --alias and honours chat_template_kwargs.enable_thinking=false, like the GPU backends. After a SIGTERM or SIGKILL of serve, no backend is left running (the child gets PR_SET_PDEATHSIG).

Next, in separate PRs: remove the embedded Lemonade (third_party/lemonade and its deltas) from the engine, and open the upstream Lemonade PR, a recipe that launches 1bit serve.

🤖 Generated with Claude Code

bong-water-water-bong and others added 3 commits September 23, 2026 18:53
The engine as a backend a host launches (geramyL's point: embed the engine
into Lemonade, not Lemonade into the engine). One model per process:
/health (+/v1/health) 503 -> 200, /v1/models, /v1/chat/completions (SSE or
not), /v1/completions.

- NPU model directory -> the NPU fast lane in process (unified.cpp, which
  now also answers /health).
- .gguf -> this build's llama-server (Vulkan0 or HRX0) or ZINC as a private
  loopback child; OpenAI routes forwarded, streaming relayed as it arrives,
  replies carry the served name, and ZINC gets requests without 'model'.
- HRX: sets IREE_HAL_AMDGPU_LIBHSA_PATH to TheRock's libhsa itself
  (--hrx-libhsa, the build's copy, or /opt/rocm-therock).
- tests/serve_e2e.sh + ctest serve_e2e_<device> (-DONEBIT_SERVE_TEST_GGUF).

Verified on Strix Halo with Qwen3-0.6B Q4_K_M: vulkan, hrx and zinc all PASS
(health, models, 'Paris.' under the served name, SSE streaming). The NPU
route is not yet run through serve_e2e.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The e2e test left llama-server/zinc children behind: SIGTERM ended serve
without its destructor, and nothing could clean up after SIGKILL. The child
is now forked with PR_SET_PDEATHSIG (the kernel ends it when serve dies,
SIGKILL included), and SIGTERM/SIGINT stop the server and the child cleanly.
Verified on Strix Halo: after SIGTERM and after SIGKILL of serve, the child is
gone; serve_e2e vulkan still passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on the NPU

serve passes --alias through to the in-process NPU server (unified now takes
--alias), and the NPU route honours chat_template_kwargs.enable_thinking=false
the way Qwen3's template does (an empty think block after the assistant
prefix), so every device answers alike. serve_e2e npu PASS on Strix Halo
(Qwen3-0.6B NPU model dir: 'Paris' under the served name, 33 SSE chunks).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 23, 2026 22:34
@bong-water-water-bong
bong-water-water-bong merged commit a51749f into main Sep 23, 2026
1 check passed
bong-water-water-bong added a commit that referenced this pull request Sep 26, 2026
…pp-vulkan 546800c) (#161)

Fork PR #21: ONEBIT_MOE_SUBST=r lets a resident expert among the router's next choices stand in
for a missing one scoring at most 1/r times higher, in the same host sync as the remap. On
Qwen3-Coder-30B, r 0.5 cut drive reads 20-27% for a KL divergence of 0.008-0.012, and decode
gained 10-35% at 1536-3072 slots. Off by default. moe/ExpertCache gains resident(), as in the
fork's copy.

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant