Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 32 additions & 1 deletion CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -17,4 +17,35 @@ add_compile_options(-Wall -Wextra -Wpedantic)

enable_testing()

# Components are added here as they land, each with its tests.
add_library(onebit_core STATIC
engine/core/gguf.cpp
engine/core/dequant.cpp
engine/core/model_config.cpp
engine/core/parallel.cpp)
target_include_directories(onebit_core PUBLIC engine/core)
find_package(Threads REQUIRED)
target_link_libraries(onebit_core PUBLIC Threads::Threads)

add_library(onebit_cpu STATIC engine/backends/cpu/cpu_model.cpp)
target_include_directories(onebit_cpu PUBLIC engine/backends/cpu)
target_link_libraries(onebit_cpu PUBLIC onebit_core)

# Unit tests: synthetic data only, so they run anywhere (including CI).
foreach(t test_gguf test_dequant)
add_executable(${t} tests/${t}.cpp)
target_include_directories(${t} PRIVATE tests)
target_link_libraries(${t} PRIVATE onebit_core)
add_test(NAME ${t} COMMAND ${t})
endforeach()

# Golden test: the CPU reference against HF transformers fp32 logits. Needs a
# model and a golden dir (tools/golden/make_golden.py), so it is registered
# only when both are given, e.g.
# -DENGINE_GOLDEN_MODEL=/path/Qwen3-0.6B-BF16.gguf -DENGINE_GOLDEN_DIR=/path/golden
add_executable(golden_cpu tests/golden_cpu.cpp)
target_link_libraries(golden_cpu PRIVATE onebit_cpu)
set(ENGINE_GOLDEN_MODEL "" CACHE FILEPATH "GGUF model for the golden test")
set(ENGINE_GOLDEN_DIR "" CACHE PATH "Golden logits directory for the golden test")
if(ENGINE_GOLDEN_MODEL AND ENGINE_GOLDEN_DIR)
add_test(NAME golden_cpu COMMAND golden_cpu ${ENGINE_GOLDEN_MODEL} ${ENGINE_GOLDEN_DIR})
endif()
2 changes: 1 addition & 1 deletion docs/PORTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ and what has to be true before it lands here. Refs are branches or commits in th

| Component | Source | State at source | Gate to land here |
|---|---|---|---|
| CPU reference | `src/gguf_reader.cpp`, `src/tokenizer.cpp`, `tools/qwen36_full_ref.py` | works | matches HF transformers fp32 logits on Qwen3-0.6B |
| CPU reference | `src/gguf_reader.cpp`, `src/tokenizer.cpp`, `tools/qwen36_full_ref.py` | **landed** (qwen3; see docs/cpu-reference.md). Tokenizer not yet | matches HF transformers fp32 logits on Qwen3-0.6B |
| Arch registry | `src/model_registry.cpp` + `Testing/census_*.json` | 569 tokens map 2,030 HF arch strings (mapping only) | data file + separate verified list |
| NPU backend | `engine/npu/src/npu_engine_universal.cpp` (`I8Ctx::init_elf`) | ELF-native, matches its own baseline | golden test vs CPU reference |
| NPU ELF dispatch table | branch `backup/iso-build-elf-native-2026-09-22` (`kElfDesigns`) | ELF and xclbin modes match token for token; 0 xclbin opens | same, in this engine |
Expand Down
62 changes: 62 additions & 0 deletions docs/cpu-reference.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# CPU reference

`engine/backends/cpu/` is the fp32 oracle the NPU and GPU backends are tested
against. It dequantizes every weight to fp32 at load time and runs one token at
a time in fp32.

It accepts only architectures that have passed the golden test below. Any other
architecture is refused by name. Loading also fails if the GGUF holds a tensor
the forward pass does not use, because an unused tensor means the model has
semantics this code does not implement.

| Architecture | Weight types | Golden test |
|---|---|---|
| `qwen3` | F32, F16, BF16, Q8_0 | Qwen3-0.6B: pass |

## Golden test: Qwen3-0.6B

Reference: HF transformers, fp32, eager attention, on the same checkpoint.
The test feeds the 36-token sequence (12 prompt tokens plus the reference's 24
greedy tokens) teacher-forced and compares the next-token logits at every position.

| Metric | Result | Gate |
|---|---|---|
| Argmax agreement | 36 / 36 | all |
| Worst per-position KL(ref ‖ cpu) | 1.3e-10 | ≤ 1e-6 |
| Worst max \|Δlogit\| | 1.0e-4 | reported only |

To confirm the test can fail, two deliberately broken builds were run:

- **Wrong RoPE layout** (Normal instead of Neox): 17 / 36 argmax mismatches, worst KL 7.9.
- **K-norm skipped**: 34 / 36 mismatches, worst KL 24.9.

Both failed the gate.

Inputs, as recorded in `tests/golden/qwen3-0.6b-capitals/meta.json`:

- Checkpoint `Qwen/Qwen3-0.6B` at revision `c1899de289a04d12100db370d81485cdf75e47ca`.
- The GGUF is converted with llama.cpp `convert_hf_to_gguf.py --outtype bf16` (llama.cpp `e71b805`).
Its sha256 is `33d6f6c9f0dd21dd3d9c99f514498ea594119508e538aa36773993a3ffd39aff`.
- The golden logits are 21.9 MB, so they are not committed. They are regenerated
by the script below; `meta.json` records their sha256.

Measured on Strix Halo (32 threads): 0.6 s load, 24 tok/s teacher-forced.

## Reproduce

```bash
python tools/golden/make_golden.py --model Qwen/Qwen3-0.6B \
--prompt "The capital of France is Paris. The capital of Japan is" \
--n-gen 24 --out golden/qwen3-0.6b-capitals
python llama.cpp/convert_hf_to_gguf.py <hf_snapshot_dir> --outtype bf16 --outfile Qwen3-0.6B-BF16.gguf

cmake -B build -DENGINE_GOLDEN_MODEL=$PWD/Qwen3-0.6B-BF16.gguf \
-DENGINE_GOLDEN_DIR=$PWD/golden/qwen3-0.6b-capitals
cmake --build build && ctest --test-dir build --output-on-failure
```

## Not done yet

- **Tokenizer.** The test uses token ids from the HF tokenizer.
- **Quantized types beyond Q8_0** (Q4_K, Q6_K and others). Each needs its own dequant test first.
- **A longer-context golden** (hundreds of positions) to cover RoPE at larger angles.
219 changes: 219 additions & 0 deletions engine/backends/cpu/cpu_model.cpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,219 @@
#include "cpu_model.h"

#include <algorithm>
#include <cmath>
#include <format>
#include <memory>
#include <set>

#include "dequant.h"

namespace onebit {

namespace {

// KV cache length. Kept well under the training context so the reference
// stays small; raise it when a test needs longer sequences.
constexpr uint32_t kMaxCtx = 4096;

void rms_norm(const float* x, const float* w, float* out, uint32_t n, float eps) {
double ss = 0.0;
for (uint32_t i = 0; i < n; ++i) ss += double(x[i]) * x[i];
const float scale = 1.0f / std::sqrt(float(ss / n) + eps);
for (uint32_t i = 0; i < n; ++i) out[i] = w[i] * (x[i] * scale);
}

void rope(float* x, uint32_t n_heads, uint32_t head_dim, uint32_t pos, float theta, RopeStyle style) {
const uint32_t half = head_dim / 2;
for (uint32_t i = 0; i < half; ++i) {
const float inv_freq = 1.0f / std::pow(theta, float(2 * i) / float(head_dim));
const float angle = float(pos) * inv_freq;
const float c = std::cos(angle), s = std::sin(angle);
for (uint32_t h = 0; h < n_heads; ++h) {
float* v = x + h * head_dim;
const uint32_t a = style == RopeStyle::Neox ? i : 2 * i;
const uint32_t b = style == RopeStyle::Neox ? i + half : 2 * i + 1;
const float x0 = v[a], x1 = v[b];
v[a] = x0 * c - x1 * s;
v[b] = x0 * s + x1 * c;
}
}
}

float silu(float x) { return x / (1.0f + std::exp(-x)); }

class Loader {
public:
explicit Loader(const GgufFile& f) : f_(f) {}

// Dequantizes tensor `name`, checking its shape is exactly `shape`.
std::expected<std::vector<float>, std::string> take(const std::string& name, std::vector<uint64_t> shape) {
const GgufTensor* t = f_.tensor(name);
if (!t) return std::unexpected(std::format("missing tensor {}", name));
if (t->ne != shape) return std::unexpected(std::format("tensor {} has an unexpected shape", name));
std::vector<float> out(t->n_elements());
if (auto r = dequantize(t->type, t->data, out.data(), out.size()); !r)
return std::unexpected(std::format("tensor {}: {}", name, r.error()));
used_.insert(name);
return out;
}

bool has(const std::string& name) const { return f_.tensor(name) != nullptr; }

std::expected<void, std::string> check_all_used() const {
for (const GgufTensor& t : f_.tensors())
if (!used_.contains(t.name)) return std::unexpected(std::format("tensor {} is not used by the reference", t.name));
return {};
}

private:
const GgufFile& f_;
std::set<std::string> used_;
};

} // namespace

std::expected<CpuModel, std::string> CpuModel::load(const std::string& gguf_path, size_t n_threads) {
auto file = GgufFile::open(gguf_path);
if (!file) return std::unexpected(file.error());
auto cfg = read_model_config(*file);
if (!cfg) return std::unexpected(cfg.error());

CpuModel m;
m.cfg_ = *cfg;
const ModelConfig& c = m.cfg_;
const uint64_t E = c.n_embd, D = c.head_dim, Q = uint64_t{c.n_head} * D, KV = uint64_t{c.n_head_kv} * D;
const uint64_t F = c.n_ff, V = c.n_vocab;

Loader ld(*file);
#define TAKE(dst, name, ...) \
do { \
auto r_ = ld.take(name, {__VA_ARGS__}); \
if (!r_) return std::unexpected(std::move(r_.error())); \
dst = std::move(*r_); \
} while (0)

TAKE(m.tok_embd_, "token_embd.weight", E, V);
TAKE(m.output_norm_, "output_norm.weight", E);
if (ld.has("output.weight")) TAKE(m.output_, "output.weight", E, V);

m.layers_.resize(c.n_layer);
for (uint32_t i = 0; i < c.n_layer; ++i) {
Layer& L = m.layers_[i];
const std::string p = std::format("blk.{}.", i);
TAKE(L.attn_norm, p + "attn_norm.weight", E);
TAKE(L.wq, p + "attn_q.weight", E, Q);
TAKE(L.wk, p + "attn_k.weight", E, KV);
TAKE(L.wv, p + "attn_v.weight", E, KV);
TAKE(L.wo, p + "attn_output.weight", Q, E);
TAKE(L.q_norm, p + "attn_q_norm.weight", D);
TAKE(L.k_norm, p + "attn_k_norm.weight", D);
TAKE(L.ffn_norm, p + "ffn_norm.weight", E);
TAKE(L.w_gate, p + "ffn_gate.weight", E, F);
TAKE(L.w_up, p + "ffn_up.weight", E, F);
TAKE(L.w_down, p + "ffn_down.weight", F, E);
}
#undef TAKE
if (auto r = ld.check_all_used(); !r) return std::unexpected(r.error());

m.n_ctx_ = std::min(c.n_ctx_train, kMaxCtx);
const size_t cache = size_t{c.n_layer} * m.n_ctx_ * KV;
m.k_cache_.assign(cache, 0.0f);
m.v_cache_.assign(cache, 0.0f);
m.x_.resize(E);
m.xn_.resize(E);
m.q_.resize(Q);
m.k_.resize(KV);
m.v_.resize(KV);
m.attn_.resize(Q);
m.scores_.resize(m.n_ctx_);
m.gate_.resize(F);
m.up_.resize(F);
m.logits_.resize(V);
m.pool_ = std::make_unique<ThreadPool>(n_threads);
return m;
}

void CpuModel::reset() { n_past_ = 0; }

void CpuModel::matvec(const std::vector<float>& w, const float* x, float* y, uint32_t rows, uint32_t cols) {
pool_->parallel_for(rows, [&](size_t begin, size_t end) {
for (size_t r = begin; r < end; ++r) {
const float* row = w.data() + r * cols;
float acc = 0.0f;
for (uint32_t i = 0; i < cols; ++i) acc += row[i] * x[i];
y[r] = acc;
}
});
}

std::expected<const std::vector<float>*, std::string> CpuModel::forward(int32_t token) {
const ModelConfig& c = cfg_;
if (token < 0 || uint32_t(token) >= c.n_vocab) return std::unexpected(std::format("token {} out of range", token));
if (n_past_ >= n_ctx_) return std::unexpected(std::format("context full ({} tokens)", n_ctx_));

const uint32_t E = c.n_embd, D = c.head_dim, H = c.n_head, HKV = c.n_head_kv;
const uint32_t KV = HKV * D, pos = n_past_, group = H / HKV;
const float att_scale = 1.0f / std::sqrt(float(D));

std::copy_n(tok_embd_.data() + size_t(token) * E, E, x_.data());

for (uint32_t l = 0; l < c.n_layer; ++l) {
const Layer& L = layers_[l];

rms_norm(x_.data(), L.attn_norm.data(), xn_.data(), E, c.rms_eps);
matvec(L.wq, xn_.data(), q_.data(), H * D, E);
matvec(L.wk, xn_.data(), k_.data(), KV, E);
matvec(L.wv, xn_.data(), v_.data(), KV, E);
for (uint32_t h = 0; h < H; ++h) rms_norm(q_.data() + h * D, L.q_norm.data(), q_.data() + h * D, D, c.rms_eps);
for (uint32_t h = 0; h < HKV; ++h) rms_norm(k_.data() + h * D, L.k_norm.data(), k_.data() + h * D, D, c.rms_eps);
rope(q_.data(), H, D, pos, c.rope_theta, c.rope_style);
rope(k_.data(), HKV, D, pos, c.rope_theta, c.rope_style);

float* kc = k_cache_.data() + (size_t(l) * n_ctx_) * KV;
float* vc = v_cache_.data() + (size_t(l) * n_ctx_) * KV;
std::copy_n(k_.data(), KV, kc + size_t(pos) * KV);
std::copy_n(v_.data(), KV, vc + size_t(pos) * KV);

for (uint32_t h = 0; h < H; ++h) {
const float* q = q_.data() + h * D;
const uint32_t kvh = h / group;
float mx = -INFINITY;
for (uint32_t t = 0; t <= pos; ++t) {
const float* k = kc + size_t(t) * KV + kvh * D;
float s = 0.0f;
for (uint32_t i = 0; i < D; ++i) s += q[i] * k[i];
scores_[t] = s * att_scale;
mx = std::max(mx, scores_[t]);
}
double sum = 0.0;
for (uint32_t t = 0; t <= pos; ++t) {
scores_[t] = std::exp(scores_[t] - mx);
sum += scores_[t];
}
float* out = attn_.data() + h * D;
std::fill_n(out, D, 0.0f);
for (uint32_t t = 0; t <= pos; ++t) {
const float p = float(scores_[t] / sum);
const float* v = vc + size_t(t) * KV + kvh * D;
for (uint32_t i = 0; i < D; ++i) out[i] += p * v[i];
}
}
matvec(L.wo, attn_.data(), xn_.data(), E, H * D);
for (uint32_t i = 0; i < E; ++i) x_[i] += xn_[i];

rms_norm(x_.data(), L.ffn_norm.data(), xn_.data(), E, c.rms_eps);
matvec(L.w_gate, xn_.data(), gate_.data(), c.n_ff, E);
matvec(L.w_up, xn_.data(), up_.data(), c.n_ff, E);
for (uint32_t i = 0; i < c.n_ff; ++i) gate_[i] = silu(gate_[i]) * up_[i];
matvec(L.w_down, gate_.data(), xn_.data(), E, c.n_ff);
for (uint32_t i = 0; i < E; ++i) x_[i] += xn_[i];
}

rms_norm(x_.data(), output_norm_.data(), xn_.data(), E, c.rms_eps);
matvec(output_.empty() ? tok_embd_ : output_, xn_.data(), logits_.data(), c.n_vocab, E);
++n_past_;
return &logits_;
}

} // namespace onebit
64 changes: 64 additions & 0 deletions engine/backends/cpu/cpu_model.h
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
// cpu_model.h: fp32 reference decoder on the CPU.
//
// This is the correctness oracle for the NPU and GPU backends, so it favours
// plain, checkable code over speed: every weight is dequantized to fp32 at load
// time and every op runs in fp32, one token at a time.
#pragma once

#include <cstdint>
#include <expected>
#include <memory>
#include <string>
#include <vector>

#include "gguf.h"
#include "model_config.h"
#include "parallel.h"

namespace onebit {

class CpuModel {
public:
// Loads every tensor of the GGUF. Fails if a tensor is missing, has the
// wrong shape, has an unsupported type, or is present but unused (an
// unused tensor means the model has semantics this code does not model).
static std::expected<CpuModel, std::string> load(const std::string& gguf_path, size_t n_threads = 0);

const ModelConfig& config() const { return cfg_; }

// Clears the KV cache.
void reset();

// Number of tokens in the KV cache.
uint32_t n_past() const { return n_past_; }

// Runs one token at position n_past() and returns the next-token logits
// (config().n_vocab floats). Fails when the context is full.
std::expected<const std::vector<float>*, std::string> forward(int32_t token);

CpuModel(CpuModel&&) = default;
CpuModel& operator=(CpuModel&&) = default;

private:
CpuModel() = default;

struct Layer {
std::vector<float> attn_norm, wq, wk, wv, wo, q_norm, k_norm;
std::vector<float> ffn_norm, w_gate, w_up, w_down;
};

void matvec(const std::vector<float>& w, const float* x, float* y, uint32_t rows, uint32_t cols);

ModelConfig cfg_;
uint32_t n_ctx_ = 0;
std::vector<float> tok_embd_, output_norm_, output_; // output_ empty when tied to tok_embd_
std::vector<Layer> layers_;
std::vector<float> k_cache_, v_cache_; // [layer][pos][n_head_kv * head_dim]
uint32_t n_past_ = 0;

// Scratch.
std::vector<float> x_, xn_, q_, k_, v_, attn_, scores_, gate_, up_, logits_;
std::unique_ptr<ThreadPool> pool_;
};

} // namespace onebit
Loading
Loading