Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions docs/lemonade.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,6 +76,39 @@ in the fork that lists the conflicts. It also keeps the fork's own copies of
upstream's workflows disabled. Lemonade asks for an RFC
before a new backend.

## Embedding checklist

What "fully embedded" still needs, measured on Strix Halo on 2026-09-25. The run used the
pinned fork (`third_party/lemonade`, `7650b4f`), the engine from `main`, and Lemonade's LLM suite
(`test/server_llm.py --wrapped-server onebit --backend <device>`). An audit copy of the suite
claimed every feature for `onebit`, so each test ran instead of being skipped.

| device | pass | fail |
|---|---|---|
| Vulkan | 19 / 31 | Responses API (2), embeddings (3), reranking (3), slots, tokenize, echo, generation parameters |
| HRX | 19 / 31 | the same 12 |
| NPU | 3 / 31 | everything that loads a model: Lemonade hands the engine a GGUF, and the NPU runs only 1bit NPU model directories |

1. **Declare what already works.** Tool calls (plain and streaming), `stop`, and async
chat/completions pass on Vulkan and HRX, but the recipe declares them unsupported. (fork)
2. **Test on the device asked for.** Lemonade's test harness passes `--backend` to llamacpp,
sd-cpp and the others, but not to `onebit`, so every `onebit` run used the default device. (fork)
3. **Forward the rest of llama-server's API.** `1bit serve` forwards only chat and completions,
so `/v1/responses`, `/slots` and `/tokenize` return 404. (engine)
4. **Embedding and reranking models.** The recipe declares only `chat`, so Lemonade refuses
embedding and reranking models. `1bit serve` serves them only as companions of a chat model
(`--embed`, `--rerank`). (engine + fork)
5. **The NPU from Lemonade.** Lemonade needs to download and hand over an NPU model
directory, and `1bit serve --device npu` needs to take it. (engine + fork)
6. **Beyond the LLM suite:**
- image generation: `1bit comfy` behind Lemonade's image endpoint;
- the Laya router;
- MLX;
- CUDA through ZINC, which needs an NVIDIA box.

`echo` and generation parameters also fail, but Lemonade's own llamacpp recipe declares
neither, so they aren't gaps.

## On its own

`1bit serve` also works without Lemonade, for any OpenAI client:
Expand Down
Loading