Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
51 changes: 2 additions & 49 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -13,64 +13,17 @@
# See the License for the specific language governing permissions and
# limitations under the License.
cmake_minimum_required(VERSION 3.24)
project(onebit_engine VERSION 0.0.1 LANGUAGES CXX)
project(onebit_engine VERSION 0.0.1 LANGUAGES C CXX)

set(CMAKE_CXX_STANDARD 23)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)

option(ENGINE_NPU "Build the XDNA 2 NPU backend (needs XRT)" OFF)
option(ENGINE_GPU "Build the llama.cpp GPU backend (Vulkan + HRX)" OFF)

if(NOT CMAKE_BUILD_TYPE AND NOT CMAKE_CONFIGURATION_TYPES)
set(CMAKE_BUILD_TYPE Release)
endif()

add_compile_options(-Wall -Wextra -Wpedantic)

enable_testing()

add_library(onebit_core STATIC
engine/core/gguf.cpp
engine/core/dequant.cpp
engine/core/model_config.cpp
engine/core/parallel.cpp
engine/core/unicode.cpp
engine/core/tokenizer.cpp)
target_include_directories(onebit_core PUBLIC engine/core)
find_package(Threads REQUIRED)
target_link_libraries(onebit_core PUBLIC Threads::Threads)

add_library(onebit_cpu STATIC engine/backends/cpu/cpu_model.cpp)
target_include_directories(onebit_cpu PUBLIC engine/backends/cpu)
target_link_libraries(onebit_cpu PUBLIC onebit_core)

# Unit tests: synthetic data only, so they run anywhere (including CI).
foreach(t test_gguf test_dequant test_unicode)
add_executable(${t} tests/${t}.cpp)
target_include_directories(${t} PRIVATE tests)
target_link_libraries(${t} PRIVATE onebit_core)
add_test(NAME ${t} COMMAND ${t})
endforeach()

# Golden test: the CPU reference against HF transformers fp32 logits. Needs a
# model and a golden dir (tools/golden/make_golden.py), so it is registered
# only when both are given, e.g.
# -DENGINE_GOLDEN_MODEL=/path/Qwen3-0.6B-BF16.gguf -DENGINE_GOLDEN_DIR=/path/golden
add_executable(golden_cpu tests/golden_cpu.cpp)
target_link_libraries(golden_cpu PRIVATE onebit_cpu)
set(ENGINE_GOLDEN_MODEL "" CACHE FILEPATH "GGUF model for the golden test")
set(ENGINE_GOLDEN_DIR "" CACHE PATH "Golden logits directory for the golden test")
if(ENGINE_GOLDEN_MODEL AND ENGINE_GOLDEN_DIR)
add_test(NAME golden_cpu COMMAND golden_cpu ${ENGINE_GOLDEN_MODEL} ${ENGINE_GOLDEN_DIR})
endif()

# Tokenizer golden test: encode/decode/pre-token split against HF tokenizers.
# Runs everywhere: the vocab-only GGUF and the cases are committed
# (tools/golden/make_tokenizer_golden.py, docs/tokenizer.md).
add_executable(golden_tokenizer tests/golden_tokenizer.cpp)
target_link_libraries(golden_tokenizer PRIVATE onebit_core)
add_test(NAME golden_tokenizer_qwen3
COMMAND golden_tokenizer ${CMAKE_SOURCE_DIR}/tests/golden/qwen3-tokenizer/vocab.gguf
${CMAKE_SOURCE_DIR}/tests/golden/qwen3-tokenizer/cases.tsv)
# Components are ported in the order listed in docs/PORTING.md.
16 changes: 9 additions & 7 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,21 +18,23 @@ limitations under the License.

This repository exists to be reviewable. Five rules keep it that way.

1. **Nothing lands without a test.** A backend or kernel change comes with a
golden-logit test against the CPU reference: same model, same prompt, token
agreement and per-step KL within a stated tolerance.
1. **Nothing lands without a test that runs it.** Every ported component comes
with a check that exercises it on real hardware or in CI, and accuracy claims
state the reference they were measured against.
2. **Numbers come from committed files.** Every benchmark or accuracy claim in a
doc names the command, the model hash and the binary it came from, and anyone
can re-run it from a clean checkout.
3. **Recognized is not the same as verified.** The arch registry reports which HF
architectures it *maps* and, separately, which ones have *passed* the golden
test on each backend. Docs quote the second number.
architectures it *maps* and, separately, which ones have *run* and been
checked on each backend. Docs quote the second number.
4. **No binaries without source.** NPU kernels are built from source in this
repository (or a pinned submodule) into full ELFs. No vendored xclbins.
5. **Every file carries the copyright and Apache-2.0 notice.** Run
`python3 tools/copyright.py --fix` before committing; CI runs `--check`.
Files that cannot hold a comment (binaries, JSON, test data) are exempt,
listed in the tool.

Porting from 1bit-MONSTER: port the smallest piece that can be tested, not whole
files. [docs/PORTING.md](docs/PORTING.md) lists the source of each component.
This repository is a port of the working code in
[1bit-MONSTER](https://github.com/1bit-MONSTER/1bit-MONSTER), without its history.
[docs/PORTING.md](docs/PORTING.md) lists the source of each component and the
order it lands in.
4 changes: 0 additions & 4 deletions NOTICE
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,3 @@
Copyright 2026 bong-water-water-bong

This product is licensed under the Apache License, Version 2.0 (see LICENSE).

engine/core/unicode_tables.inc is generated from the Unicode Character
Database (via Python's unicodedata module), Copyright Unicode, Inc., used
under the Unicode License v3 (https://www.unicode.org/license.txt).
45 changes: 11 additions & 34 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,40 +16,17 @@ limitations under the License.
-->
# 1bit engine

A model-agnostic LLM inference backend in C++ for AMD Ryzen AI (Strix Halo):
XDNA 2 NPU, Radeon iGPU (Vulkan and HRX), and a CPU reference path. It serves the
llama-server-compatible OpenAI HTTP API, so [Lemonade](https://github.com/lemonade-sdk/lemonade)
and any OpenAI client can drive it as a drop-in backend.

> **Status: scaffold.** Nothing runs yet. Components land one at a time, each with a
> test against the CPU reference (see [CONTRIBUTING.md](CONTRIBUTING.md)). The
> development history lives in [1bit-MONSTER](https://github.com/1bit-MONSTER/1bit-MONSTER);
> this repository contains only code that has been verified.

## Layout

| Path | What |
|---|---|
| `engine/core/` | GGUF / safetensors / Q4NX loading, tokenizer, arch registry, sampler, KV cache |
| `engine/route/` | Backend selection: a hard capability check, then the Laya scorer |
| `engine/backends/npu/` | XDNA 2 via XRT, full-ELF kernels only (no xclbins) |
| `engine/backends/gpu/` | llama.cpp (AMD-Ecosystem fork) with the Vulkan and HRX20 devices |
| `engine/backends/cpu/` | fp32 reference forward pass; the correctness oracle |
| `engine/server/` | llama-server-compatible HTTP: `/v1/chat/completions`, `/v1/completions`, `/v1/models`, `/health` |
| `tests/` | Golden-logit tests per backend against the CPU reference |
| `lemonade/` | Upstream patch: `BackendDescriptor` + `WrappedServer` for this engine |

## Milestone 1

Qwen3-0.6B end to end on NPU, GPU and CPU through the server, loaded by Lemonade.

## Build

```bash
cmake -B build -DENGINE_NPU=OFF -DENGINE_GPU=OFF
cmake --build build
ctest --test-dir build
```
One binary that runs [Lemonade](https://github.com/lemonade-sdk/lemonade) completely,
with AMD Ryzen AI hardware behind it:

- the XDNA 2 NPU engine
- HRX and Vulkan on the Radeon iGPU, compiled together in one llama.cpp build
- Laya, which decides where each request runs

> **Status:** porting the working engine from
> [1bit-MONSTER](https://github.com/1bit-MONSTER/1bit-MONSTER) in four steps; see
> [docs/PORTING.md](docs/PORTING.md). This repository holds the verified code
> without the development history.

## License

Expand Down
48 changes: 28 additions & 20 deletions docs/PORTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,25 +16,33 @@ limitations under the License.
-->
# Porting map

Where each component comes from in [1bit-MONSTER](https://github.com/1bit-MONSTER/1bit-MONSTER),
and what has to be true before it lands here. Refs are branches or commits in that repo.
This repository is the working 1bit-MONSTER engine, ported without its history.
Each step below is one PR that builds and runs on Strix Halo before the next one starts.

| Component | Source | State at source | Gate to land here |
| Step | Component | Source in 1bit-MONSTER | Done when |
|---|---|---|---|
| CPU reference | `src/gguf_reader.cpp`, `src/tokenizer.cpp`, `tools/qwen36_full_ref.py` | **landed** (qwen3; docs/cpu-reference.md). Tokenizer **landed** (docs/tokenizer.md) | matches HF transformers fp32 logits on Qwen3-0.6B |
| Arch registry | `src/model_registry.cpp` + `Testing/census_*.json` | 569 tokens map 2,030 HF arch strings (mapping only) | data file + separate verified list |
| NPU backend | `engine/npu/src/npu_engine_universal.cpp` (`I8Ctx::init_elf`) | ELF-native, matches its own baseline | golden test vs CPU reference |
| NPU ELF dispatch table | branch `backup/iso-build-elf-native-2026-09-22` (`kElfDesigns`) | ELF and xclbin modes match token for token; 0 xclbin opens | same, in this engine |
| NPU 16-tile layer kernel | branch `bench/fastlane-16tile-corrections-2026-09-22` (435bf36e7) | 24/24 tokens, summed KL 0.000466, 0.340 ms/layer | rebuild from source, reproduce |
| GPU backend | AMD-Ecosystem/llama.cpp fork, `GGML_VULKAN` + `GGML_HRX2`; recipe on `fix/zaya-lmhead-evidence` | Vulkan0 74.8 tok/s, HRX20 18.4 on zaya1-8b; Q4NX is HRX20-only | pinned submodule, linked (no dlopen of copied structs) |
| Router | `src/model_router.cpp` (Q4NX rule) + Laya scorer, branch `backup/laya-and-results-2026-09-22` | Laya not yet checked against its Python reference | Laya matches its Python reference within a stated tolerance |
| Server | `src/server/` | works | llama-server flag and endpoint compatibility test |
| Lemonade recipe | lemonade `src/cpp/include/lemon/backends/*` pattern | n/a | loads Qwen3-0.6B through a local Lemonade build |

Known traps carried over:

- NPU concurrency: one device; each engine instance uses 4 hw contexts; throughput
peaks at about 4 concurrent instances.
- Q4NX containers have separate formats per family (unsigned q4_1 for Qwen3;
additive `w = q*scale + min` for the 35B MoE experts). Verify against a reference, never by eye.
- The NPU model containers currently live in `~/.config/flm/models/*-NPU2` on the dev box.
| 1 | **Embedded Lemonade.** Lemonade's server core runs inside the `1bit` binary, with the `onebit` backend added | `third_party/lemonade` (v11.9.0 plus local deltas, see its `UPSTREAM.md`); `tools/unified_server.cpp` `run_embedded_lemonade` | `1bit lemonade` starts and serves `/v1/models` |
| 2 | **HRX with Vulkan.** The AMD-Ecosystem llama.cpp fork built with `GGML_HRX2=ON` and `GGML_VULKAN=ON` in one build | fork build recipe from `fix/zaya-lmhead-evidence` (67503a794); `src/backend_hrx.cpp` | Lemonade loads a GGUF on `Vulkan0` and a Q4NX on `HRX20` |
| 3 | **NPU engine.** Full ELFs only, and the open 16-tile layer kernel built from source | `engine/npu` (`npu_engine_universal.cpp`, `I8Ctx::init_elf`); ELF dispatch table on `backup/iso-build-elf-native-2026-09-22`; kernel on `bench/fastlane-16tile-corrections-2026-09-22` | Lemonade loads Qwen3-0.6B on the NPU through `onebit` |
| 4 | **Laya router.** A non-autoregressive scorer that picks where each request runs | `src/laya_scorer.cpp`, `include/laya_scorer.h` on `backup/laya-and-results-2026-09-22`; model at `~/models/laya` | matches its Python reference; routes requests |

## How the pieces fit

```
1bit lemonade ── Lemonade server core (in-process)
├─ llamacpp-hrx recipe ──> HRX build: llama.cpp + ggml-hrx + ggml-vulkan
│ devices HRX20 and Vulkan0
├─ onebit recipe ────────> 1bit unified -m <artifact> ──> NPU engine (full ELFs)
└─ Laya ─────────────────> scores each request: which model and device
```

## Known facts carried over

- **Vulkan is not inside HRX.** One build compiles both ggml backends.
- On zaya1-8b, `Vulkan0` decodes at 74.8 tok/s and `HRX20` at 18.4, because
HRX2 falls back to the CPU for 2,088 ops.
- Q4NX loads only on `HRX20`.
- **NPU concurrency:**
- There is one device, and each engine instance uses 4 hardware contexts.
- Throughput peaks at about 4 concurrent instances.
- **Where the NPU model containers live:** they are currently in `~/.config/flm/models/*-NPU2` on the dev box.
78 changes: 0 additions & 78 deletions docs/cpu-reference.md

This file was deleted.

85 changes: 0 additions & 85 deletions docs/tokenizer.md

This file was deleted.

17 changes: 0 additions & 17 deletions engine/backends/cpu/README.md

This file was deleted.

Loading
Loading