Skip to content

docs/lemonade.md: the embedding checklist (measured) - #68

Merged
bong-water-water-bong merged 1 commit into
mainfrom
lemonade-audit
Sep 25, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
lemonade-audit

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

The gap audit for embedding the engine fully in Lemonade. Lemonade's LLM suite ran through onebit on the pinned fork (7650b4f), with an audit copy that claims every feature so no test is skipped:

  • Vulkan and HRX: 19 of 31 pass.
  • NPU: 3 of 31. Lemonade hands the engine a GGUF, which the NPU route can't run.

Six gaps are listed, each marked with where its fix goes (engine or fork). Fixes follow as separate PRs.

🤖 Generated with Claude Code

Lemonade's LLM suite through onebit on the pinned fork, every feature claimed:
Vulkan 19/31, HRX 19/31, NPU 3/31; six gaps listed with where each fix goes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit f1ade64

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Unmeasured Claim

The PR claims that the gaps were measured on Strix Halo on 2026-09-25, but does not provide specific performance metrics or test results to back up the claim. The document states that the run used a pinned fork and the engine from main, but does not include concrete data such as throughput, latency, or accuracy numbers that would validate the measurements.

## Embedding checklist

What "fully embedded" still needs, measured on Strix Halo on 2026-09-25. The run used the
pinned fork (`third_party/lemonade`, `7650b4f`), the engine from `main`, and Lemonade's LLM suite
(`test/server_llm.py --wrapped-server onebit --backend <device>`). An audit copy of the suite
claimed every feature for `onebit`, so each test ran instead of being skipped.

| device | pass | fail |
|---|---|---|
| Vulkan | 19 / 31 | Responses API (2), embeddings (3), reranking (3), slots, tokenize, echo, generation parameters |
| HRX | 19 / 31 | the same 12 |
| NPU | 3 / 31 | everything that loads a model: Lemonade hands the engine a GGUF, and the NPU runs only 1bit NPU model directories |

1. **Declare what already works.** Tool calls (plain and streaming), `stop`, and async
   chat/completions pass on Vulkan and HRX, but the recipe declares them unsupported. (fork)
2. **Test on the device asked for.** Lemonade's test harness passes `--backend` to llamacpp,
   sd-cpp and the others, but not to `onebit`, so every `onebit` run used the default device. (fork)
3. **Forward the rest of llama-server's API.** `1bit serve` forwards only chat and completions,
   so `/v1/responses`, `/slots` and `/tokenize` return 404. (engine)
4. **Embedding and reranking models.** The recipe declares only `chat`, so Lemonade refuses
   embedding and reranking models. `1bit serve` serves them only as companions of a chat model
   (`--embed`, `--rerank`). (engine + fork)
5. **The NPU from Lemonade.** Lemonade needs to download and hand over an NPU model
   directory, and `1bit serve --device npu` needs to take it. (engine + fork)
6. **Beyond the LLM suite:**
   - image generation: `1bit comfy` behind Lemonade's image endpoint;
   - the Laya router;
   - MLX;
   - CUDA through ZINC, which needs an NVIDIA box.

`echo` and generation parameters also fail, but Lemonade's own llamacpp recipe declares
neither, so they aren't gaps.

@bong-water-water-bong
bong-water-water-bong merged commit e1d8c64 into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the lemonade-audit branch September 25, 2026 10:30
bong-water-water-bong pushed a commit that referenced this pull request Oct 2, 2026
#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
…tic FA, attention sinks) (#295)

* Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss)

#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode)

#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant