Skip to content

Server: 1bit-server speaks Lemonade's backend protocol - #4

Closed
bong-water-water-bong wants to merge 1 commit into
mainfrom
server
Closed

bong-water-water-bong wants to merge 1 commit into
mainfrom
server

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

1bit-server lets Lemonade run the engine through the same backend protocol it uses for its GGUF recipes (llamacpp, llamacpp-hrx). Only the protocol is shared. The model code, tokenizer, sampler and template rendering are all this repo's.

What

  • 1bit-server:
    • Flags: takes Lemonade's launch flags (-m --ctx-size --device --port --jinja --metrics --parallel 1).
    • /health: 503 Loading model while loading, then 200.
    • Endpoints: /v1/models, /v1/chat/completions and /v1/completions, streaming (SSE, [DONE]) and not.
    • Telemetry: OpenAI usage plus llama-server timings, the fields Lemonade's telemetry reads.
  • Chat:
    • the GGUF's own Jinja template, with chat_template_kwargs (e.g. enable_thinking)
    • <think> routed to reasoning_content, which Lemonade's chat, MCP and Ollama paths read
    • stop strings, and end-of-turn detection
  • Prompt cache: the shared prefix with the previous request is reused (cache_n), so multi-turn chat evaluates only the new turn.
  • Pinned deps: three header-only MIT libraries in cmake/deps.cmake, each with a sha256: nlohmann/json 3.12.0, cpp-httplib 0.57.1 and minja 021c2293. They're credited in NOTICE.
  • Docs: docs/server.md covers the contract, the verification, and what's not done yet.

Verification

  • server_e2e.py: launches the binary with Lemonade's flags and checks over HTTP (Qwen3-0.6B, strixhalo):
    • greedy /v1/completions text byte-identical to HF transformers' fp32 greedy continuation
    • streamed == non-streamed, for completions and chat (content and reasoning)
    • prefix reuse (cache_n = n_prompt - 1)
    • stop strings, enable_thinking:false, and the thinking split
    • the 400/404 error codes
  • golden_template_qwen3: 12 conversations byte-identical to HF apply_chat_template, including tools, tool calls, tool responses and the thinking toggle. Runs in CI.
  • Unit tests: test_text_stream covers UTF-8, stops and reasoning byte by byte; test_openai checks response shapes structurally; plus test_sampler. All run in CI.
  • Totals: ctest 10/10 on strixhalo, 8/8 locally; zero warnings.

Not yet (in docs/server.md)

Tool-call output parsing, /v1/responses, embeddings, /metrics, >1 concurrent sequence, and non-CPU devices. The NPU and GPU backends plug in behind this server next.

🤖 Generated with Claude Code

- engine/server/main.cpp: 1bit-server. It takes Lemonade's launch flags
  (-m --ctx-size --device --port --jinja --metrics --parallel 1) and serves:
  /health (503 while loading), /v1/models, /v1/chat/completions and
  /v1/completions, streaming (SSE, [DONE]) and not. Responses carry
  OpenAI usage and llama-server timings.
- generator: prefix-cached generation loop (cache_n), end-of-turn tokens.
- text_stream: UTF-8-safe deltas, stop strings across tokens, <think>
  routed to reasoning_content.
- chat_template: the GGUF's Jinja template via minja.
- openai: request parsing (rejects n>1, non-text parts, over-long
  prompts) and response shapes.
- core/sampler: temperature/top-k/top-p/min-p/seed, llama.cpp chain order.
- CpuModel: configurable n_ctx and truncate() for prefix reuse.
- Pinned deps (cmake/deps.cmake, sha256): nlohmann/json 3.12.0,
  cpp-httplib 0.57.1, minja 021c2293. All MIT, credited in NOTICE.
- Tests: unit tests for sampler, text stream and openai (CI); a
  chat-template golden against HF apply_chat_template (12 cases, CI);
  server_e2e.py over HTTP with Lemonade's flags.

Qwen3-0.6B: greedy /v1/completions over HTTP is byte-identical to HF
transformers' fp32 greedy continuation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong

Copy link
Copy Markdown
Collaborator Author

Closing: wrong direction. The repo ports the existing 1bit-MONSTER engine instead (embedded Lemonade, HRX with Vulkan, the NPU engine, Laya); see the port plan.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant