Skip to content

Laya router: the step-4 scorer in C++, routing each request to a device - #90

Merged
bong-water-water-bong merged 5 commits into
mainfrom
laya/scorer
Sep 25, 2026
Merged

bong-water-water-bong merged 5 commits into
mainfrom
laya/scorer

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

agent-b72a1b's Laya scorer work, brought onto current main.

What it adds

  • laya/scorer.{h,cpp} and laya/safetensors.{h,cpp}: the non-autoregressive step-4 scorer (ModernBERT-large encoder + RLCD head) in C++. Weights load from model.safetensors, tokens come from npu::Tokenizer, and nothing runs in Python.
  • npu::Tokenizer: gains token_id() and the ByteLevel regex, so ModernBERT's tokenizer.json encodes digit runs correctly.
  • laya/route.{h,cpp}: route_device(scorer, state, devices), which answers the fixed routing question among the candidate devices.
  • 1bit route --laya-model DIR --state TEXT [--devices ...]: prints the scorer's decision.
  • 1bit serve --laya-model DIR (with --device auto):
    • each chat or completion request goes to the device the scorer picks, among the ones this build can run a .gguf on (vulkan, hrx, zinc)
    • a device's backend starts the first time it is picked, so one model is not loaded on every device at once
    • a Q4NX model directory still asks the router, with npu as its only candidate.

Merge notes: the work was based on a pre-#69 main. Its serve.cpp rewrite (Backend/backend_for) is re-applied on today's Launch/launch_for structure, and launch_for is unchanged from main. Laya mode:

  • prepares a Launch per candidate device, and starts it lazily under a lock
  • uses the adaptive router's in-flight accounting and each backend's own drop_model/set_model
  • refuses to combine with --adaptive, --embedding/--reranking, --prefill-device and --lean.

Tested on Strix Halo: ctest 8/8:

  • laya_gate_root and laya_gate_typed_decisions: raw logits ≤ 1e-3, act logits ≤ 1e-3 relative, argmax identical, probabilities ≤ 1e-4 against the Python reference, on the pinned checkpoints.
  • laya_route_e2e: a request reaches the Laya-chosen backend, across vulkan and zinc, plus the npu-only decision.
  • smoke_serve, think_split, npu_*.

🤖 Generated with Claude Code

bong-water-water-bong and others added 3 commits September 25, 2026 15:13
…nd route requests per device

The non-autoregressive System 1 scorer (1bit-MONSTER src/laya_scorer.cpp on
backup/laya-and-results-2026-09-22) lands as laya/scorer.{h,cpp} +
laya/safetensors.{h,cpp} (Apache-2.0, zero Python at runtime): ModernBERT-large
encoder + RLCD head, weights via model.safetensors, tokens via npu::Tokenizer
(tokenizer.json). npu::Tokenizer gains token_id() and the Rust-tokenizers
ByteLevel regex so ModernBERT's tokenizer.json encodes digit runs correctly.

Gated against the Python reference on the pinned root and typed-decisions
checkpoints: tests/laya_gate.cpp compares raw logits (<= 1e-3), act logits
(<= 1e-3 rel), argmax (identical) and probabilities (<= 1e-4). Both pass.

Wired into the router: 1bit serve --laya-model routes each request by the
scorer (replacing the auto->vulkan hardcode) and lazily starts the chosen
device's backend; 1bit route prints the decision. tests/laya_route_e2e.sh
proves a request reaches the Laya-chosen backend.
… dirs, GPU backends for .gguf)

route_device now takes a candidate device list. A .gguf routes among vulkan,
hrx and zinc (NPU is never offered for a .gguf, so a request can no longer
route to a device that cannot run it), and an NPU model directory routes to
npu through the same path. 1bit route gains --devices. The e2e test now
proves routing across more than one device (vulkan and zinc) plus the
npu-only decision.
…ckend lines)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit ad47f86

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit ad47f86)

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

69 - Partially compliant

Compliant requirements:

  • Fix the gap where Lemonade's llamacpp backend forwards certain routes to the backend, but 1bit serve answered 404
  • On vulkan, hrx and rocm, 1bit serve now forwards /v1/responses, /tokenize, /detokenize, /apply-template, GET /slots, POST /slots/?action=…, GET /props, GET /metrics to the llama-server behind the model
  • On npu, zinc and mlx these routes answer 501 with a clear message
  • docs/serve.md lists the routes
  • Measured on Strix Halo, Lemonade's test/server_llm.py --wrapped-server onebit shows improvement from 19/31 to 23/31 passing tests for Vulkan and HRX

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 4 🔵🔵🔵🔵⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Incorrect device routing for NPU model directories

When using --laya-model with a Q4NX model directory, the code attempts to route the request to the NPU device, but the logic in run_serve function incorrectly assumes that the device is always one of the GGUF devices (vulkan, hrx, zinc). This can lead to incorrect routing decisions when the model directory is intended for NPU only.

        // A Q4NX model directory runs on the NPU only; with --laya-model the router is still asked
        // (with npu its only candidate), so every device the engine serves goes through Laya.
        if (!o.laya_model.empty()) {
            onebit::laya::Scorer scorer;
            if (!scorer.load(o.laya_model)) throw std::runtime_error("laya: " + scorer.error());
            const std::string dev = onebit::laya::route_device(scorer, o.model, {"npu"});
            std::fprintf(stderr, "1bit serve: laya routes %s to %s (Q4NX model directory)\n", model_id(o).c_str(), dev.c_str());
        }
#endif
Potential race condition in backend startup

The ensure function in serve_child uses a mutex (start_mu) to protect the started vector, but there's a potential race condition where a backend could be started by one thread while another thread is checking if it's already started. This could lead to multiple processes being started for the same backend.

auto ensure = [&](size_t i) -> bool {
    std::lock_guard<std::mutex> lock(start_mu);
    if (started[i]) return true;
    const Launch& b = backends[i];
    std::fprintf(stderr, "1bit serve: %s on %s (%s), picked by laya\n", id.c_str(), b.device.c_str(), b.argv[0].c_str());
    auto c = std::make_unique<Child>(b.argv, b.env, b.port);
    if (!c->wait_ready(std::chrono::seconds(600), g_stop)) return false;
    started[i] = c.get();
    children.push_back(std::move(c));
    return true;
};
Potential memory allocation in hot path

The score function in Scorer class allocates several vectors (seqs, markers, input_ids, attn, mpos, mmask, h, tmp, qkv, attn_out, q, k, v, scores, hn, qkv, attn_out, tmp, tmp2, ffn, m, mn, m1, logits, act_logits, z, e, gp, rp) on the stack. If the number of questions or the length of the input text is large, this could lead to stack overflow, especially in a high-concurrency environment.

bool Scorer::score(const std::string& state, const std::vector<Question>& questions,
                   std::vector<Answer>& answers, RawOutput* raw) {
    err_.clear();
    std::unique_ptr<onebit::npu::Tokenizer> tok;
    try {
        tok = std::make_unique<onebit::npu::Tokenizer>(tok_path_);
    } catch (const std::exception& e) {
        err_ = std::string("load tokenizer: ") + e.what();
        return false;
    }
    const int N = (int)questions.size();
    if (N == 0) { answers.clear(); return true; }

    // ---- build token sequences + markers (laya/common.py:build_sequence) ----
    std::vector<std::vector<int>> seqs(N), markers(N);
    std::vector<int> qtype(N);
    for (int i = 0; i < N; ++i) {
        const Question& q = questions[i];
        std::vector<std::string> opts;
        int qt;
        if (q.type == "choice") {
            qt = 0;
            for (auto& kv : q.criteria)
                opts.push_back(kv.second.empty() ? kv.first : kv.first + ": " + kv.second);
        } else if (q.type == "score") {
            qt = 1;
            for (size_t j = 0; j < q.criteria.size(); ++j)
                opts.push_back("level " + std::to_string(j) + ": " + q.criteria[j].second);
        } else {
            qt = 2;
            const std::string *fc = nullptr, *tc = nullptr;
            for (auto& kv : q.criteria) { if (kv.first == "false") fc = &kv.second; else if (kv.first == "true") tc = &kv.second; }
            opts.push_back("false: " + (fc && !fc->empty() ? *fc : "no, the statement does not hold"));
            opts.push_back("true: " + (tc && !tc->empty() ? *tc : "yes, the statement holds"));
        }
        qtype[i] = qt;

        std::vector<int> head_ids = encode(*tok, q.type + " question: " + q.instructions);
        std::vector<std::vector<int>> opt_ids(opts.size());
        for (size_t oi = 0; oi < opts.size(); ++oi) {
            std::vector<int> o = encode(*tok, " " + opts[oi]);
            if ((int)o.size() > 48) o.resize(48);
            opt_ids[oi].push_back(mask_id_);
            opt_ids[oi].insert(opt_ids[oi].end(), o.begin(), o.end());
        }
        int opt_budget = head_max_len_;
        for (auto& o : opt_ids) opt_budget -= (int)o.size();
        if (opt_budget < 16) {
            int per = std::max(4, (head_max_len_ - 16) / std::max(1, (int)opt_ids.size()));
            for (auto& o : opt_ids) o.resize(std::min((size_t)per, o.size()));
            opt_budget = head_max_len_;
            for (auto& o : opt_ids) opt_budget -= (int)o.size();
        }
        int hb = std::max(8, opt_budget);
        if ((int)head_ids.size() > hb) head_ids.resize(hb);

        std::vector<int> ids;
        ids.push_back(cls_id_);
        ids.insert(ids.end(), head_ids.begin(), head_ids.end());
        ids.push_back(sep_id_);
        for (auto& o : opt_ids) {
            markers[i].push_back((int)ids.size());
            ids.insert(ids.end(), o.begin(), o.end());
        }
        ids.push_back(sep_id_);
        int room = std::max(0, max_len_ - (int)ids.size() - 1);
        std::vector<int> st = encode(*tok, state);
        if ((int)st.size() > room) st.resize(room);  // truncate_left=false
        ids.insert(ids.end(), st.begin(), st.end());
        ids.push_back(sep_id_);
        if ((int)ids.size() > max_len_) ids.resize(max_len_);
        seqs[i] = std::move(ids);
        std::vector<int> m2;
        for (int mm : markers[i]) if (mm < max_len_) m2.push_back(mm);
        markers[i] = std::move(m2);
    }

    // ---- collate (pad to L, K) ----
    int L = 0, K = 0;
    for (int i = 0; i < N; ++i) { L = std::max(L, (int)seqs[i].size()); K = std::max(K, (int)markers[i].size()); }
    std::vector<int64_t> input_ids((size_t)N * L, pad_id_);
    std::vector<int64_t> attn((size_t)N * L, 0);
    std::vector<int64_t> mpos((size_t)N * K, 0);
    std::vector<int8_t> mmask((size_t)N * K, 0);
    for (int i = 0; i < N; ++i) {
        for (size_t j = 0; j < seqs[i].size(); ++j) { input_ids[i * L + j] = seqs[i][j]; attn[i * L + j] = 1; }
        for (size_t j = 0; j < markers[i].size(); ++j) { mpos[i * K + j] = markers[i][j]; mmask[i * K + j] = 1; }
    }

    // ---- ModernBERT encoder forward ----
    std::vector<float> h((size_t)N * L * D);
    for (int b = 0; b < N; ++b)
        for (int t = 0; t < L; ++t)
            memcpy(h.data() + ((size_t)b * L + t) * D, tok_emb_.data() + (size_t)input_ids[b * L + t] * D, D * 4);
    {
        std::vector<float> tmp((size_t)N * L * D);
        layernorm(h.data(), emb_norm_.data(), nullptr, N * L, D, tmp.data());
        h.swap(tmp);
    }
    // RoPE tables
    auto build_cs = [](float theta, int L) {
        std::vector<float> cos(L * HD), sin(L * HD);
        for (int p = 0; p < L; ++p) for (int i = 0; i < HD / 2; ++i) {
            float inv = 1.0f / powf(theta, (float)(2 * i) / HD);
            float f = (float)p * inv;
            cos[p * HD + i] = cosf(f); cos[p * HD + HD/2 + i] = cosf(f);
            sin[p * HD + i] = sinf(f); sin[p * HD + HD/2 + i] = sinf(f);
        }
        return std::make_pair(cos, sin);
    };
    auto full_cs = build_cs(160000.0f, L);
    auto slid_cs = build_cs(10000.0f, L);
    std::vector<int8_t> key_pad((size_t)N * L);
    for (int b = 0; b < N; ++b) for (int t = 0; t < L; ++t) key_pad[b * L + t] = (attn[b * L + t] == 0);

    std::vector<float> hn((size_t)N * L * D), qkv((size_t)N * L * 3072), attn_out((size_t)N * L * D);
    std::vector<float> q((size_t)N * NHEAD * L * HD), k((size_t)N * NHEAD * L * HD), v((size_t)N * NHEAD * L * HD);
    std::vector<float> scores((size_t)L * L);
    for (int li = 0; li < NLAYER; ++li) {
        LayerW& Lw = layers_[li];
        bool is_full = (li % 3 == 0);
        const auto& cs = is_full ? full_cs : slid_cs;
        if (li > 0) layernorm(h.data(), Lw.attn_norm.data(), nullptr, N * L, D, hn.data());
        else memcpy(hn.data(), h.data(), h.size() * 4);
        for (int b = 0; b < N; ++b) for (int t = 0; t < L; ++t) {
            const float* x = hn.data() + ((size_t)b * L + t) * D;
            float* o = qkv.data() + ((size_t)b * L + t) * 3072;
            for (int j = 0; j < 3072; ++j) {
                const float* wr = Lw.Wqkv.data() + (size_t)j * D;
                float acc = 0; for (int c = 0; c < D; ++c) acc += x[c] * wr[c];
                o[j] = acc;
            }
        }
        for (int b = 0; b < N; ++b) for (int t = 0; t < L; ++t) {
            const float* qkv_t = qkv.data() + ((size_t)b * L + t) * 3072;
            for (int hd = 0; hd < NHEAD; ++hd) {
                float* qo = q.data() + (((size_t)b * NHEAD + hd) * L + t) * HD;
                float* ko = k.data() + (((size_t)b * NHEAD + hd) * L + t) * HD;
                float* vo = v.data() + (((size_t)b * NHEAD + hd) * L + t) * HD;
                for (int d = 0; d < HD; ++d) { qo[d] = qkv_t[hd*HD+d]; ko[d] = qkv_t[1024+hd*HD+d]; vo[d] = qkv_t[2048+hd*HD+d]; }
            }
        }
        // RoPE
        for (int b = 0; b < N; ++b) for (int hd = 0; hd < NHEAD; ++hd) for (int t = 0; t < L; ++t) {
            float* qo = q.data() + (((size_t)b * NHEAD + hd) * L + t) * HD;
            float* ko = k.data() + (((size_t)b * NHEAD + hd) * L + t) * HD;
            const float* c = cs.first.data() + (size_t)t * HD;
            const float* s = cs.second.data() + (size_t)t * HD;
            for (int d = 0; d < HD/2; ++d) {
                float qa = qo[d], qb = qo[d+HD/2];
                qo[d] = qa*c[d] - qb*s[d]; qo[d+HD/2] = qb*c[d+HD/2] + qa*s[d+HD/2];
                float ka = ko[d], kb = ko[d+HD/2];
                ko[d] = ka*c[d] - kb*s[d]; ko[d+HD/2] = kb*c[d+HD/2] + ka*s[d+HD/2];
            }
        }
        // attention
        const float scale = 1.0f / sqrtf((float)HD);
        for (int b = 0; b < N; ++b) for (int hd = 0; hd < NHEAD; ++hd) {
            const float* qb = q.data() + (((size_t)b * NHEAD + hd) * L) * HD;
            const float* kb = k.data() + (((size_t)b * NHEAD + hd) * L) * HD;
            for (int qi = 0; qi < L; ++qi) {
                float* sr = scores.data() + (size_t)qi * L;
                for (int kj = 0; kj < L; ++kj) {
                    if (attn[b * L + kj] == 0) { sr[kj] = -INFINITY; continue; }
                    if (!is_full && abs(qi - kj) > 64) { sr[kj] = -INFINITY; continue; }
                    const float* qr = qb + (size_t)qi * HD;
                    const float* kr = kb + (size_t)kj * HD;
                    float acc = 0; for (int d = 0; d < HD; ++d) acc += qr[d]*kr[d];
                    sr[kj] = acc * scale;
                }
                float mx = -INFINITY; for (int kj = 0; kj < L; ++kj) mx = fmaxf(mx, sr[kj]);
                float sum = 0; for (int kj = 0; kj < L; ++kj) { sr[kj] = expf(sr[kj]-mx); sum += sr[kj]; }
                for (int kj = 0; kj < L; ++kj) sr[kj] /= sum;
                float* orow = attn_out.data() + ((size_t)b * L + qi) * D + (size_t)hd * HD;
                for (int d = 0; d < HD; ++d) orow[d] = 0;
                for (int kj = 0; kj < L; ++kj) {
                    const float* vr = v.data() + (((size_t)b * NHEAD + hd) * L + kj) * HD;
                    float s = sr[kj];
                    for (int d = 0; d < HD; ++d) orow[d] += s * vr[d];
                }
            }
        }
        // Wo projection + residual
        for (int b = 0; b < N; ++b) for (int t = 0; t < L; ++t) {
            const float* x = attn_out.data() + ((size_t)b * L + t) * D;
            float* o = hn.data() + ((size_t)b * L + t) * D;
            for (int j = 0; j < D; ++j) {
                const float* wr = Lw.Wo.data() + (size_t)j * D;
                float acc = 0; for (int c = 0; c < D; ++c) acc += x[c] * wr[c];
                o[j] = acc;
            }
        }
        for (size_t z = 0; z < h.size(); ++z) h[z] += hn[z];
        // mlp_norm + geglu + residual
        layernorm(h.data(), Lw.mlp_norm.data(), nullptr, N * L, D, hn.data());
        geglu(hn.data(), Lw.Wi.data(), Lw.Wo_mlp.data(), N * L, attn_out.data());
        for (size_t z = 0; z < h.size(); ++z) h[z] += attn_out[z];
    }
    {
        std::vector<float> tmp((size_t)N * L * D);
        layernorm(h.data(), final_norm_.data(), nullptr, N * L, D, tmp.data());
        h.swap(tmp);
    }

    // ---- decision head ----
    for (int b = 0; b < N; ++b) {
        const float* te = type_emb_.data() + (size_t)qtype[b] * D;
        for (int t = 0; t < L; ++t) {
            float* o = h.data() + ((size_t)b * L + t) * D;
            for (int d = 0; d < D; ++d) o[d] += te[d];
        }
    }
    std::vector<float> tmp((size_t)N * L * D), tmp2((size_t)N * L * D), ffn((size_t)N * L * 4096);
    for (int layer = 0; layer < (int)head_layers_.size(); ++layer) {
        HeadLayerW& H = head_layers_[layer];
        layernorm(h.data(), H.n1w.data(), H.n1b.data(), N * L, D, tmp.data());
        mha(tmp.data(), H.in_w.data(), H.in_b.data(), H.out_w.data(), H.out_b.data(), N, L, key_pad.data(), tmp2.data());
        for (size_t z = 0; z < h.size(); ++z) h[z] += tmp2[z];
        layernorm(h.data(), H.n2w.data(), H.n2b.data(), N * L, D, tmp.data());
        linear(tmp.data(), H.l1w.data(), H.l1b.data(), N * L, D, 4096, ffn.data());
        for (size_t z = 0; z < (size_t)N * L * 4096; ++z) ffn[z] = fmaxf(0.0f, ffn[z]);  // relu
        linear(ffn.data(), H.l2w.data(), H.l2b.data(), N * L, 4096, D, tmp.data());
        for (size_t z = 0; z < h.size(); ++z) h[z] += tmp[z];
    }
    // gather markers
    std::vector<float> m((size_t)N * K * D);
    for (int b = 0; b < N; ++b) for (int kk = 0; kk < K; ++kk) {
        long long p = mpos[b * K + kk]; if (p < 0) p = 0;
        memcpy(m.data() + ((size_t)b * K + kk) * D, h.data() + ((size_t)b * L + p) * D, D * 4);
    }
    // scorer: LayerNorm -> Linear(D->D) -> GELU -> Linear(D->1)
    std::vector<float> mn((size_t)N * K * D), m1((size_t)N * K * D);
    layernorm(m.data(), sc0w.data(), sc0b.data(), N * K, D, mn.data());
    linear(mn.data(), sc1w.data(), sc1b.data(), N * K, D, D, m1.data());
    for (auto& x : m1) x = gelu(x);
    std::vector<float> logits((size_t)N * K);
    linear(m1.data(), sc3w.data(), sc3b.data(), N * K, D, 1, logits.data());
    for (int b = 0; b < N; ++b) for (int kk = 0; kk < K; ++kk)
        if (!mmask[b * K + kk]) logits[b * K + kk] = -1e4f;
    // act_head
    std::vector<float> act_logits((size_t)N * 2);
    for (int b = 0; b < N; ++b) {
        const float* lrow = logits.data() + (size_t)b * K;
        int kvalid = 0; for (int kk = 0; kk < K; ++kk) if (mmask[b * K + kk]) kvalid++;
        float keff = kvalid < 2 ? 2.0f : (float)kvalid;
        float p[64], mx = -INFINITY, sum = 0;
        for (int kk = 0; kk < K; ++kk) mx = fmaxf(mx, lrow[kk]);
        for (int kk = 0; kk < K; ++kk) { p[kk] = expf(lrow[kk] - mx); sum += p[kk]; }
        for (int kk = 0; kk < K; ++kk) p[kk] /= sum;
        float t1 = -INFINITY, t2 = -INFINITY;
        for (int kk = 0; kk < K; ++kk) { if (p[kk] > t1) { t2 = t1; t1 = p[kk]; } else if (p[kk] > t2) t2 = p[kk]; }
        if (t2 == -INFINITY) t2 = t1;
        float ent = 0; for (int kk = 0; kk < K; ++kk) ent += -p[kk] * logf(fmaxf(p[kk], 1e-9f));
        ent /= logf(keff);
        float feats[4] = { t1, t1 - t2, ent, keff / 255.0f };
        const float* pooled = h.data() + (size_t)b * L * D;
        float cat[1028];
        memcpy(cat, pooled, D * 4); memcpy(cat + D, feats, 4 * 4);
        float h256[256];
        linear(cat, act0w.data(), act0b.data(), 1, 1028, 256, h256);
        for (auto& x : h256) x = gelu(x);
        float a2[2];
        linear(h256, act2w.data(), act2b.data(), 1, 256, 2, a2);
        act_logits[b * 2] = a2[0]; act_logits[b * 2 + 1] = a2[1];
    }

    // raw pre-softmax outputs, for gating
    if (raw) {
        raw->logits.clear();
        raw->act_logits.clear();
        raw->logits.reserve(N);
        raw->act_logits.reserve(N);
        for (int b = 0; b < N; ++b) {
            raw->logits.emplace_back(logits.begin() + (size_t)b * K, logits.begin() + (size_t)(b + 1) * K);
            raw->act_logits.emplace_back(act_logits.begin() + (size_t)b * 2, act_logits.begin() + (size_t)(b + 1) * 2);
        }
    }

    // ---- temperature + softmax + confidence -> answers ----
    answers.clear();
    answers.reserve(N);
    for (int b = 0; b < N; ++b) {
        const Question& q = questions[b];
        int kvalid = 0; for (int kk = 0; kk < K; ++kk) if (mmask[b * K + kk]) kvalid++;
        // temperature bucket
        float tscale = temperature_[qtype[b]];
        std::string bucket;
        int sz = kvalid <= 2 ? 2 : kvalid <= 5 ? 5 : kvalid <= 10 ? 10 : 11;
        bucket = std::string(qtype[b] == 0 ? "choice:" : qtype[b] == 1 ? "score:" : "noul:") +
                 (sz == 2 ? "2" : sz == 5 ? "3-5" : sz == 10 ? "6-10" : "11+");
        for (auto& kv : temperature_by_options_) if (kv.first == bucket) { tscale = kv.second; break; }
        float z[64], mx = -INFINITY, sum = 0;
        for (int kk = 0; kk < kvalid; ++kk) { z[kk] = logits[b * K + kk] / tscale; mx = fmaxf(mx, z[kk]); }
        for (int kk = 0; kk < kvalid; ++kk) { z[kk] = expf(z[kk] - mx); sum += z[kk]; }
        for (int kk = 0; kk < kvalid; ++kk) z[kk] /= sum;
        float ent = 0; for (int kk = 0; kk < kvalid; ++kk) ent += -z[kk] * logf(fmaxf(z[kk], 1e-9f));
        float conf = 1.0f - ent / logf((float)kvalid);
        float actp = 0.0f;
        {
            // stable softmax: the act logits are large (~1e3-1e4), so expf alone overflows.
            float mx = fmaxf(act_logits[b*2], act_logits[b*2+1]);
            float ea0 = expf(act_logits[b*2] - mx), ea1 = expf(act_logits[b*2+1] - mx);
            actp = ea0 / (ea0 + ea1);
        }

        Answer a;
        a.type = q.type;
        a.confidence = std::clamp(conf, 0.0f, 1.0f);
        a.act_probability = actp;
        if (q.type == "choice") {
            int bi = 0; for (int kk = 1; kk < kvalid; ++kk) if (z[kk] > z[bi]) bi = kk;
            a.choice = q.criteria[bi].first;
            for (int kk = 0; kk < kvalid; ++kk) a.probabilities.push_back({q.criteria[kk].first, z[kk]});
        } else if (q.type == "score") {
            float exp_score = 0; for (int kk = 0; kk < kvalid; ++kk) exp_score += (float)kk * z[kk];
            a.score = exp_score;
            for (int kk = 0; kk < kvalid; ++kk) a.probabilities.push_back({std::to_string(kk), z[kk]});
        } else {
            a.noul = z[1];
            a.confidence = fmaxf(z[1], 1.0f - z[1]);
        }
        answers.push_back(std::move(a));
    }
    return true;
}

⚠️ Review coverage: The following files were not included in this review because of the token budget:

  • CMakeLists.txt
  • laya/route.h

… for bit)

One routing decision took about 15 s on one core of Strix Halo's 32.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Preparing review...

@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 25, 2026 18:26
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 6fc774b

@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit ad47f86

@bong-water-water-bong
bong-water-water-bong merged commit 27fefaf into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the laya/scorer branch September 25, 2026 18:29
bong-water-water-bong added a commit that referenced this pull request Sep 25, 2026
Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Oct 5, 2026
…to our fork (Loom NaN backport) (#320)

* Bump HRX: llama.cpp 522dab4789ce (long-prompt fix), hrx-system to our fork 98d05d94a9f9

third_party/llama.cpp e57beb9721af -> 522dab4789ce (fork #89 runtime overhead,
#90: graph inputs never written back to host — every multi-ubatch prompt was
wrong since #82 — and the non-replay upload race).
third_party/hrx-system moves from ROCm/hrx-system 51b1739ae5fd to our fork
1bit-MONSTER/hrx-system 1bit/main 98d05d94a9f9 = the same AMD commit plus the
GFX11 wave64 lane-mask drain backport (#313 NaN) and two libhrx device knobs.
bump-hrx.yml now moves only the llama.cpp pin; hrx-system is rebased by hand.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* registry: regenerated at llama.cpp 522dab4789ce

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant