Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
4 changes: 2 additions & 2 deletions .github/workflows/bump-hrx.yml
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@
# (docs/hrx.md, "Staying current"). When AMD moves its pins, this workflow:
# 1. syncs 1bit-MONSTER/llama.cpp: master to ggml-org, 1bit/hrx-vulkan to AMD's pin;
# 2. opens a PR here moving both submodules.
# GitHub-hosted CI builds that PR without HRX; run tests/hrx_lemonade_e2e.sh on
# GitHub-hosted CI builds that PR without HRX; run tests/serve_e2e.sh on
# Strix Halo before merging it.
#
# Needs the secret HRX_BUMP_TOKEN: a fine-grained token with Contents
Expand Down Expand Up @@ -127,4 +127,4 @@ jobs:
| third_party/hrx-system (ROCm/hrx-system) | \`${OURS_HRX:0:12}\` | \`${AMD_HRX:0:12}\` |

CI here builds without HRX. Before merging, on Strix Halo:
\`cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit && tests/hrx_lemonade_e2e.sh build/1bit build/hrx/llama/bin/llama-server\`"
\`cmake -B build -G Ninja -DONEBIT_HRX=ON && cmake --build build --target onebit && for d in vulkan hrx; do tests/serve_e2e.sh build/1bit <Qwen3-0.6B Q4_K_M .gguf> \$d; done\`"
48 changes: 30 additions & 18 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -28,24 +28,40 @@ enable_testing()

# Components are ported in the order listed in docs/PORTING.md.

# ── Step 1: Lemonade's server core, embedded ───────────────────────────────
# Vendored at third_party/lemonade (v11.9.0 plus the local deltas listed in
# its UPSTREAM.md, including the `onebit` backend).
set(BUILD_WEB_APP OFF CACHE BOOL "" FORCE)
set(BUILD_TESTING OFF CACHE BOOL "" FORCE)
add_subdirectory(third_party/lemonade EXCLUDE_FROM_ALL)
# ── The engine binary: `1bit serve`, an OpenAI-compatible API (docs/serve.md)
# The engine is a backend a host such as Lemonade launches; it no longer embeds
# one. JSON and the HTTP server/client are pinned here, at the commits the
# engine used before (nlohmann/json v3.11.3, cpp-httplib v0.47.0).
include(FetchContent)
find_package(Threads REQUIRED)
set(JSON_BuildTests OFF CACHE INTERNAL "")
FetchContent_Declare(json
GIT_REPOSITORY https://github.com/nlohmann/json.git
GIT_TAG 9cca280a4d0ccf0c08f47a99aa71d1b0e52f8d03) # v3.11.3
set(HTTPLIB_USE_OPENSSL_IF_AVAILABLE OFF CACHE BOOL "" FORCE)
set(HTTPLIB_USE_ZLIB_IF_AVAILABLE OFF CACHE BOOL "" FORCE)
set(HTTPLIB_USE_BROTLI_IF_AVAILABLE OFF CACHE BOOL "" FORCE)
set(HTTPLIB_USE_ZSTD_IF_AVAILABLE OFF CACHE BOOL "" FORCE)
set(HTTPLIB_INSTALL OFF CACHE BOOL "" FORCE)
FetchContent_Declare(httplib
GIT_REPOSITORY https://github.com/yhirose/cpp-httplib.git
GIT_TAG fe332fa06bac76a1c6d402c08f414052999347da) # v0.47.0
FetchContent_MakeAvailable(json httplib)

add_executable(onebit app/main.cpp app/serve.cpp)
set_target_properties(onebit PROPERTIES OUTPUT_NAME 1bit)
target_link_libraries(onebit PRIVATE lemonade-server-core)
target_link_libraries(onebit PRIVATE nlohmann_json::nlohmann_json httplib::httplib Threads::Threads)

# Runs everywhere, CI included: serve's proxy against a fake backend.
add_test(NAME smoke_serve COMMAND ${CMAKE_SOURCE_DIR}/tests/smoke_serve.sh $<TARGET_FILE:onebit>)

# ── `1bit serve` end to end (docs/serve.md) ───────────────────────────────
# Needs a GGUF (e.g. Qwen3-0.6B Q4_K_M); tests each device this build has.
set(ONEBIT_SERVE_TEST_GGUF "" CACHE FILEPATH "GGUF the serve_e2e tests load")

# ── Step 2: HRX + Vulkan in one llama.cpp build (docs/hrx.md) ─────────────
# Off by default: it needs TheRock/ROCm and the three HRX submodules, and it
# only runs on AMD GPUs. `1bit lemonade` then uses this llama-server for both
# only runs on AMD GPUs. `1bit serve` then uses this llama-server for both
# the llamacpp-hrx recipe (HRX0) and the llamacpp recipe's Vulkan backend.
option(ONEBIT_HRX "Build HRX + Vulkan llama-server and wire it into Lemonade" OFF)
if(ONEBIT_HRX)
Expand All @@ -58,13 +74,11 @@ if(ONEBIT_HRX)
COMMAND ${CMAKE_SOURCE_DIR}/tests/serve_e2e.sh $<TARGET_FILE:onebit> ${ONEBIT_SERVE_TEST_GGUF} ${_dev})
endforeach()
endif()
add_test(NAME hrx_lemonade_e2e
COMMAND ${CMAKE_SOURCE_DIR}/tests/hrx_lemonade_e2e.sh $<TARGET_FILE:onebit> ${ONEBIT_HRX_SERVER})
endif()

# ── ZINC (docs/zinc.md) ────────────────────────────────────────────────────
# Off by default: it builds third_party/zinc with scripts/build-zinc.sh (which
# fetches its own Zig) for one GPU backend. `1bit lemonade` then serves the zinc
# fetches its own Zig) for one GPU backend. `1bit serve --device zinc` then runs the zinc
# recipe with this binary. ONEBIT_ZINC_BACKEND=cuda is the NVIDIA build.
option(ONEBIT_ZINC "Build ZINC (third_party/zinc) and wire it into Lemonade" OFF)
if(ONEBIT_ZINC)
Expand All @@ -84,8 +98,6 @@ if(ONEBIT_ZINC)
add_test(NAME serve_e2e_zinc
COMMAND ${CMAKE_SOURCE_DIR}/tests/serve_e2e.sh $<TARGET_FILE:onebit> ${ONEBIT_SERVE_TEST_GGUF} zinc)
endif()
add_test(NAME zinc_lemonade_e2e
COMMAND ${CMAKE_SOURCE_DIR}/tests/zinc_lemonade_e2e.sh $<TARGET_FILE:onebit> ${ONEBIT_ZINC_SERVER})
endif()

# ── Hugging Face tokenizers (docs/tokenizers.md) ───────────────────────────
Expand Down Expand Up @@ -183,11 +195,11 @@ if(ONEBIT_NPU)
endif()
endif()

add_test(NAME smoke_lemonade COMMAND ${CMAKE_SOURCE_DIR}/tests/smoke_lemonade.sh $<TARGET_FILE:onebit>)

# MLX on Apple Silicon through Lemonade's mlx backend (docs/apple.md).
set(ONEBIT_MLX_SERVER "" CACHE FILEPATH "lemon-mlx-engine's server binary, for tests/mlx_lemonade_e2e.sh")
# MLX on Apple Silicon: `1bit serve --device mlx` (docs/apple.md).
set(ONEBIT_MLX_SERVER "" CACHE FILEPATH "lemon-mlx-engine's server binary, for serve_e2e_mlx")
if(APPLE AND ONEBIT_MLX_SERVER)
add_test(NAME mlx_lemonade_e2e
COMMAND ${CMAKE_SOURCE_DIR}/tests/mlx_lemonade_e2e.sh $<TARGET_FILE:onebit> ${ONEBIT_MLX_SERVER})
add_test(NAME serve_e2e_mlx
COMMAND ${CMAKE_SOURCE_DIR}/tests/serve_e2e.sh $<TARGET_FILE:onebit> mlx-community/Qwen3-0.6B-4bit mlx
--mlx-server ${ONEBIT_MLX_SERVER})
endif()
47 changes: 1 addition & 46 deletions NOTICE
Original file line number Diff line number Diff line change
Expand Up @@ -13,19 +13,6 @@ third_party/ (or in its upstream repository for components that are fetched at
build time rather than stored here). The copyright lines below are taken from
each project's own LICENSE file or source headers.

-------------------------------------------------------------------------------
Included in this repository, with modifications
-------------------------------------------------------------------------------

Lemonade (third_party/lemonade)
https://github.com/lemonade-sdk/lemonade
Copyright (C) 2022-2024, Advanced Micro Devices, Inc.
License: Apache-2.0
Modified: the local changes (the onebit, mlx and zinc backends, the
hrx_device option, the embeddability patch and model-registry entries) are
listed in third_party/lemonade/UPSTREAM.md, as Apache-2.0 section 4(b)
requires. Lemonade bundles AixLog, Copyright (C) 2017-2021 Johannes Pohl (MIT).

-------------------------------------------------------------------------------
Pinned at exact commits (git submodules) and built from source
-------------------------------------------------------------------------------
Expand Down Expand Up @@ -86,8 +73,7 @@ convaiinnovations/laya, pinned in config/laya.json)
License: Apache-2.0

-------------------------------------------------------------------------------
Fetched by the build at pinned commits (third_party/lemonade/CMakeLists.txt),
unless a suitable system copy is found
Fetched by the build at pinned commits (CMakeLists.txt)
-------------------------------------------------------------------------------

JSON for Modern C++ (nlohmann/json v3.11.3)
Expand All @@ -99,34 +85,3 @@ cpp-httplib (v0.47.0)
https://github.com/yhirose/cpp-httplib
Copyright (c) 2017 yhirose
License: MIT

CLI11 (v2.4.2)
https://github.com/CLIUtils/CLI11
Copyright (c) 2017-2026 University of Cincinnati, developed by Henry Schreiner
under NSF AWARD 1414736
License: BSD-3-Clause

curl (8.5.0)
https://github.com/curl/curl
Copyright (c) 1996-2026, Daniel Stenberg, <daniel@haxx.se>, and many contributors
License: curl

Zstandard (zstd v1.5.7)
https://github.com/facebook/zstd
Copyright (c) Meta Platforms, Inc. and affiliates.
License: BSD-3-Clause OR GPL-2.0-only (used here under BSD-3-Clause)

Mbed TLS (v3.6.2)
https://github.com/Mbed-TLS/mbedtls
Copyright The Mbed TLS Contributors
License: Apache-2.0 OR GPL-2.0-or-later (used here under Apache-2.0)

libwebsockets (v4.3.3)
https://github.com/warmcat/libwebsockets
Copyright (C) 2010-2020 Andy Green <andy@warmcat.com> and contributors
License: MIT (some sources under similar permissive licenses; see its LICENSE)

Brotli (v1.1.0, macOS builds only)
https://github.com/google/brotli
Copyright (c) 2009, 2010, 2013-2016 by the Brotli Authors
License: MIT
30 changes: 17 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,22 +16,27 @@ limitations under the License.
-->
# 1bit engine

One binary that runs [Lemonade](https://github.com/lemonade-sdk/lemonade) completely,
with AMD Ryzen AI hardware behind it:
An inference engine for AMD Ryzen AI, built to run inside
[Lemonade](https://github.com/lemonade-sdk/lemonade): the `1bit` binary serves one model behind an
OpenAI-compatible API, whatever device runs it, and Lemonade (or any OpenAI client) launches it
the way it launches `llama-server`. Behind that one API:

- the XDNA 2 NPU engine
- HRX and Vulkan on the Radeon iGPU, compiled together in one llama.cpp build
- ZINC, which also reaches NVIDIA GPUs (CUDA) and Apple GPUs (Metal)
- MLX on Apple Silicon, through lemon-mlx-engine
- Laya, which decides where each request runs
- every Hugging Face model architecture, kept current by a daily census

> **Status:** steps 1–3 of 5 have landed: embedded Lemonade ([docs/lemonade.md](docs/lemonade.md)),
> HRX + Vulkan in one build on AMD's live ggml-hrx ([docs/hrx.md](docs/hrx.md)), and the NPU engine on
> full ELFs with the upstream XDNA stack pinned ([docs/npu.md](docs/npu.md); its layer kernel is
> not yet built from source). Also: MLX on Apple Silicon through Lemonade's `mlx` recipe
> ([docs/apple.md](docs/apple.md)), and ZINC, which reaches NVIDIA GPUs through its CUDA backend
> ([docs/zinc.md](docs/zinc.md)). Next is step 4, the Laya router. The working engine is being ported from
> [1bit-MONSTER](https://github.com/1bit-MONSTER/1bit-MONSTER) in five steps; see
> **Status:** `1bit serve` is the engine's front door ([docs/serve.md](docs/serve.md)): the NPU,
> Vulkan, HRX and ZINC each pass its end-to-end test on Strix Halo. Following geramyL's review, the
> engine is embedded into Lemonade rather than embedding it: the vendored Lemonade is gone, and a
> Lemonade recipe that launches `1bit serve` is being prepared for upstream ([docs/lemonade.md](docs/lemonade.md)).
> Ported so far: HRX + Vulkan in one build on AMD's live ggml-hrx ([docs/hrx.md](docs/hrx.md)), the NPU
> engine on full ELFs with the upstream XDNA stack pinned ([docs/npu.md](docs/npu.md); its layer kernel
> is not yet built from source), ZINC ([docs/zinc.md](docs/zinc.md)) and MLX
> ([docs/apple.md](docs/apple.md)). Next is step 4, the Laya router. The working engine is being ported
> from [1bit-MONSTER](https://github.com/1bit-MONSTER/1bit-MONSTER); see
> [docs/PORTING.md](docs/PORTING.md). Measured results are on the
> [wiki](https://github.com/1bit-MONSTER/engine/wiki). This repository holds the verified code
> without the development history.
Expand All @@ -46,7 +51,7 @@ Apache-2.0. See [LICENSE](LICENSE).
vibecoding, and that is where all of this started. Without it, this engine would not exist.

**The Lemonade team and AMD's developers.** [Lemonade](https://github.com/lemonade-sdk/lemonade)
runs inside this binary. AMD's developers left breadcrumbs all over the place: the XDNA driver
is the home this engine is built to run in. AMD's developers left breadcrumbs all over the place: the XDNA driver
and XRT, IRON and Peano, HRX, their tested llama.cpp integration, their issues, their examples.
This engine is what following those breadcrumbs built.

Expand All @@ -65,9 +70,8 @@ The repositories this engine is built on, in order of importance:
| 9 | [huggingface/tokenizers](https://github.com/huggingface/tokenizers) | Every model's `tokenizer.json`, byte-exact, behind our C ABI | Apache-2.0 |
| 10 | [NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) | The router that decides where each request runs | Apache-2.0 |

Also built on [nlohmann/json](https://github.com/nlohmann/json) (MIT),
[yhirose/cpp-httplib](https://github.com/yhirose/cpp-httplib) (MIT), and the libraries Lemonade
builds with: curl, zstd, Mbed TLS, libwebsockets, CLI11 and Brotli.
Also built on [nlohmann/json](https://github.com/nlohmann/json) (MIT)
and [yhirose/cpp-httplib](https://github.com/yhirose/cpp-httplib) (MIT).

Every third-party copyright and license is listed in [NOTICE](NOTICE). Each project keeps its
own license; nothing here relicenses anyone's work.
Loading
Loading