Skip to content

README, NOTICE: acknowledge FastFlowLM - #74

Merged
bong-water-water-bong merged 1 commit into
mainfrom
ack-fastflowlm
Sep 25, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
ack-fastflowlm

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

The NPU route runs FastFlowLM's Q4NX models, and Lemonade's onebit recipe now downloads FastFlowLM/*-NPU2 (#72, fork #7). FastFlowLM's README asks projects to acknowledge it with Powered by [FastFlowLM](https://github.com/ROCm/FastFlowLM). This adds:

  • that line and a row in the README's thanks table;
  • a NOTICE entry: the code is MIT; the engine reads the Q4NX format and runs the models on its own full-ELF kernels, without FastFlowLM's runtime or kernel binaries.

🤖 Generated with Claude Code

…t asks

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit f12b5df

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

72 - Partially compliant

Compliant requirements:

  • The build's own lax kernels are the default for NPU inference
  • 1bit serve finds the lax kernels without a flag
  • The lookup order is implemented
  • -DONEBIT_NPU_LAX=ON builds the lax kernels into the build tree
  • -DONEBIT_NPU_LAX_KERNELS=<dir> points at an existing build
  • The change is measured on strixhalo with specific md5 hashes and performance metrics
  • Documentation in docs/npu-lax.md and docs/lemonade.md is updated

Non-compliant requirements:

  • The small NPU models are not yet supported
  • reasoning_content on the NPU route is not yet supported
  • Our own HF repos are not yet supported

Requires further human verification:

  • The small NPU models support
  • The reasoning_content on the NPU route
  • Our own HF repos support

7 - Partially compliant

Compliant requirements:

  • Pinned sources for ROCm/hrx, ROCm/hrx-system, and llama.cpp fork
  • Patches 0001–0003 for build plumbing and runtime fix
  • Build with -DONEBIT_HRX=ON builds all three from the build tree
  • Lemonade wiring for 1bit lemonade to run llamacpp-hrx on HRX20 and llamacpp/Vulkan on Vulkan0
  • New hrx_device option in Lemonade
  • Verified on Strix Halo with reproducible kernels and speed parity
  • Default build with ONEBIT_HRX=OFF is unchanged

Non-compliant requirements:

  • zaya Q4NX through Lemonade is not yet supported
  • A relocatable package for HRX libraries is not yet supported
  • HRX in CI is not yet supported

Requires further human verification:

  • zaya Q4NX through Lemonade support
  • A relocatable package for HRX libraries
  • HRX in CI support
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Missing Documentation Update

The PR adds acknowledgment for FastFlowLM in the README but does not update the documentation files (docs/npu-lax.md and docs/lemonade.md) to reflect the new default behavior of using the build's own lax kernels. This is a requirement from the ticket description.

| 11 | [ROCm/FastFlowLM](https://github.com/ROCm/FastFlowLM) | The Q4NX NPU model format and its models on Hugging Face (`FastFlowLM/*-NPU2`), which the engine's NPU route runs on its own kernels | MIT |

Powered by [FastFlowLM](https://github.com/ROCm/FastFlowLM)
Incomplete License Information

The NOTICE file includes a note about FastFlowLM's MIT license for code, but it does not fully reflect the license information for the models themselves, which are under different licenses (e.g., Qwen's Apache-2.0). The description should be more precise about the licensing of the models and their weights.

License: MIT (code). The engine's NPU route reads FastFlowLM's Q4NX model format,
and Lemonade's onebit recipe downloads FastFlowLM's models from Hugging Face
(FastFlowLM/*-NPU2; the weights keep their own licenses, e.g. Qwen's Apache-2.0).
The engine runs them on its own kernels as full ELFs; it does not use
FastFlowLM's runtime or ship its kernel binaries.
Powered by FastFlowLM, as FastFlowLM asks.

@bong-water-water-bong
bong-water-water-bong merged commit de1ef13 into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the ack-fastflowlm branch September 25, 2026 11:18
bong-water-water-bong pushed a commit that referenced this pull request Oct 2, 2026
#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Oct 2, 2026
…tic FA, attention sinks) (#295)

* Pin llama.cpp 2bd7f58: gpt-oss on HRX, TQ1_0/TQ2_0, MLA prompts past 512 tokens, deterministic flash attention

llama.cpp fork since cebcd70 (balanced mode, figures from each PR):
- #66 MUL_MAT_ID at multiples of 32 plus a decode-loader stride fix; #67 ADD_ID and SWIGLU_OAI on HRX;
  #73 a placement guard for the CPU/HRX split bug (engine #286). gpt-oss-20b MXFP4: pp512 25.8 -> ~1000,
  tg128 12.6 -> ~35 tok/s, text correct, KLD vs CPU 0.029.
- #69 TQ1_0/TQ2_0 on HRX: Ternary-Bonsai-1.7B KLD vs CPU 0.000523; pp512/tg128 3542/113 and 4100/156.
- #70 MLA V transpose: GLM-4.7-Flash prompts of 512+ tokens gave garbage (PPL 315,664), now 5.916 (CPU 5.959).
- #71 llama-hadamard folds qwen3next's ssm_ba and refuses unfoldable stamped files.
- #72 masked flash-attention keys reach P*V as V = +0: identical requests now give identical logits
  (Qwen3-0.6B and Qwen3.8-27B bit-identical repeats); pp512 -4.7% on Qwen3-0.6B.

Docs: docs/hrx.md "Our patches". Registry regenerated (no mapping changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp 4485916: attention sinks on HRX (gpt-oss)

#68 runs gpt-oss's sink logits on HRX as an exact post-correction of the flash-attention output.
Docs: docs/hrx.md "Our patches". Registry pin updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp f5b7f4a: PrismML tile bytes (PQ2_0 small-model decode)

#74: the low-token SwiGLU read every PQ2_0 / PTQ1_0 row from row 0; Ternary-Bonsai-1.7B PQ2_0 now
matches the CPU. Found by the release format matrix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant