Server: 1bit-server speaks Lemonade's backend protocol - #4
Closed
bong-water-water-bong wants to merge 1 commit into
Closed
bong-water-water-bong wants to merge 1 commit into
bong-water-water-bong wants to merge 1 commit into
Conversation
- engine/server/main.cpp: 1bit-server. It takes Lemonade's launch flags (-m --ctx-size --device --port --jinja --metrics --parallel 1) and serves: /health (503 while loading), /v1/models, /v1/chat/completions and /v1/completions, streaming (SSE, [DONE]) and not. Responses carry OpenAI usage and llama-server timings. - generator: prefix-cached generation loop (cache_n), end-of-turn tokens. - text_stream: UTF-8-safe deltas, stop strings across tokens, <think> routed to reasoning_content. - chat_template: the GGUF's Jinja template via minja. - openai: request parsing (rejects n>1, non-text parts, over-long prompts) and response shapes. - core/sampler: temperature/top-k/top-p/min-p/seed, llama.cpp chain order. - CpuModel: configurable n_ctx and truncate() for prefix reuse. - Pinned deps (cmake/deps.cmake, sha256): nlohmann/json 3.12.0, cpp-httplib 0.57.1, minja 021c2293. All MIT, credited in NOTICE. - Tests: unit tests for sampler, text stream and openai (CI); a chat-template golden against HF apply_chat_template (12 cases, CI); server_e2e.py over HTTP with Lemonade's flags. Qwen3-0.6B: greedy /v1/completions over HTTP is byte-identical to HF transformers' fp32 greedy continuation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator
Author
|
Closing: wrong direction. The repo ports the existing 1bit-MONSTER engine instead (embedded Lemonade, HRX with Vulkan, the NPU engine, Laya); see the port plan. |
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
1bit-serverlets Lemonade run the engine through the same backend protocol it uses for its GGUF recipes (llamacpp,llamacpp-hrx). Only the protocol is shared. The model code, tokenizer, sampler and template rendering are all this repo's.What
1bit-server:-m --ctx-size --device --port --jinja --metrics --parallel 1)./health: 503Loading modelwhile loading, then 200./v1/models,/v1/chat/completionsand/v1/completions, streaming (SSE,[DONE]) and not.usageplus llama-servertimings, the fields Lemonade's telemetry reads.chat_template_kwargs(e.g.enable_thinking)<think>routed toreasoning_content, which Lemonade's chat, MCP and Ollama paths readcache_n), so multi-turn chat evaluates only the new turn.cmake/deps.cmake, each with a sha256: nlohmann/json 3.12.0, cpp-httplib 0.57.1 and minja021c2293. They're credited inNOTICE.docs/server.mdcovers the contract, the verification, and what's not done yet.Verification
server_e2e.py: launches the binary with Lemonade's flags and checks over HTTP (Qwen3-0.6B, strixhalo):/v1/completionstext byte-identical to HF transformers' fp32 greedy continuationcache_n = n_prompt - 1)enable_thinking:false, and the thinking splitgolden_template_qwen3: 12 conversations byte-identical to HFapply_chat_template, including tools, tool calls, tool responses and the thinking toggle. Runs in CI.test_text_streamcovers UTF-8, stops and reasoning byte by byte;test_openaichecks response shapes structurally; plustest_sampler. All run in CI.ctest10/10 on strixhalo, 8/8 locally; zero warnings.Not yet (in docs/server.md)
Tool-call output parsing,
/v1/responses, embeddings,/metrics, >1 concurrent sequence, and non-CPU devices. The NPU and GPU backends plug in behind this server next.🤖 Generated with Claude Code