Skip to content

Pin Lemonade: third_party/lemonade at the fork's 7650b4f - #67

Merged
bong-water-water-bong merged 1 commit into
mainfrom
pin-lemonade
Sep 25, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
pin-lemonade

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

The engine runs inside Lemonade, so it pins the Lemonade it's tested with.

  • Pin: third_party/lemonade is our fork 1bit-MONSTER/lemonade (upstream plus the onebit recipe) at 7650b4f, which carries the upstream sync of 2026-09-24.
  • Never linked: the engine doesn't compile or link any Lemonade code, so geramyL's point still holds: the engine lives inside Lemonade, not the reverse. scripts/build-lemonade.sh <prefix> builds lemond and the lemonade CLI from the pin; it fetches the submodule if it's missing.
  • Kept current: .github/workflows/bump-lemonade.yml runs daily after the fork sync and opens a PR here when the fork's main moves, the same way bump-zinc does.
  • Docs: NOTICE (Lemonade, Apache-2.0, AMD), docs/lemonade.md (the pin, the build, and the plan to embed fully before going upstream).

Tested on strixhalo: scripts/build-lemonade.sh built lemond from 7650b4f (exit 0). The same build is running Lemonade's LLM suite through onebit for the embedding gap audit.

🤖 Generated with Claude Code

The engine is embedded in Lemonade, so it pins the Lemonade it is tested with: our fork
(upstream + the onebit recipe). Never linked; scripts/build-lemonade.sh builds lemond from
the pin, bump-lemonade.yml opens a PR when the fork's main moves. NOTICE, docs/lemonade.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit eca6366

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Potential Submodule Fetch Failure

The script attempts to fetch the Lemonade submodule only if the CMakeLists.txt file is missing and git is initialized. However, if the submodule is initialized but the CMakeLists.txt file is missing for some reason, the script may not correctly detect this and proceed to build, potentially leading to a build failure. The check should ensure that the submodule is properly initialized and contains the expected files.

if [ ! -f "$src/CMakeLists.txt" ] && git -C "$root" rev-parse --git-dir > /dev/null 2>&1; then
    echo "fetching third_party/lemonade"
    git -C "$root" submodule update --init third_party/lemonade
fi
[ -f "$src/CMakeLists.txt" ] || { echo "third_party/lemonade is empty and could not be fetched: git submodule update --init third_party/lemonade"; exit 1; }
PR Body Formatting

The PR body includes a multi-line string that is not properly formatted for markdown rendering. Specifically, the commit log section uses a raw string that should be formatted as a code block or list for better readability.

  --body "Moves third_party/lemonade from \`${OURS:0:12}\` to the fork's main \`${UPSTREAM:0:12}\`.

Fork commits (last 30):
${log}

CI here has no GPU. Before merging, on Strix Halo:

@bong-water-water-bong
bong-water-water-bong merged commit 5874e8e into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the pin-lemonade branch September 25, 2026 10:10
bong-water-water-bong pushed a commit that referenced this pull request Oct 2, 2026
…512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
…tic FA, attention sinks) (#295)

* Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss)

#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode)

#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant