Skip to content

ZAYA1 (Zyphra) on the GPU: pin llama.cpp 5556bf2, serve routes fork-only architectures - #107

Merged
bong-water-water-bong merged 5 commits into
mainfrom
zaya-vulkan
Sep 26, 2026
Merged

bong-water-water-bong merged 5 commits into
mainfrom
zaya-vulkan

Conversation

@bong-water-water-bong

@bong-water-water-bong bong-water-water-bong commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

ZAYA1-8B (Zyphra) joins Qwen in the showcase, served by 1bit serve on the Strix Halo GPU.

What changes

Checks (strixhalo, Radeon 8060S)

Unit test: gguf_meta (new, hosted CI), plus smoke_serve and think_split.

Routing: with the upstream Vulkan server set to /bin/false, ZAYA still passes serve_e2e (routed to our build) and Qwen3-0.6B fails (routed to upstream). The routing works both ways.

registry_check, vulkan and hrx rows: 16/16 PASS. This ran at engine b9812e0 with the e2e6ca2 HRX build; #6 changes only the -fa off path, and ZAYA and Qwen3-0.6B pass -fa off on HRX0 with it. That's Qwen3-0.6B, Qwen2.5-7B, Qwen3-Coder-30B-A3B, Qwen3.6-35B-A3B, GLM-4.7-Flash, MiniCPM4-8B, MiniCPM5-1B and ZAYA1-8B, each on both devices.

Accuracy against transformers FP32, teacher-forced (96 positions): the ZAYA1 F16 GGUF agrees on the top-1 token at 95/96. The miss is a 0.07-nat tie; transformers' own BF16 run matches FP32 at 91/96.

Speed, ZAYA1-8B Q4_K_M:

Device Prefill (pp512) Decode (tg128) Perplexity
Vulkan0 3,437 tok/s 93.0 tok/s 21.57
ROCm0 ~2,400 tok/s 61.5 tok/s 21.78
HRX0 1,175 tok/s 25.5 tok/s 21.55

The engine's --device rocm server is built from ROCmFPX's tree, which has no ZAYA. The ROCm numbers above come from our llama.cpp built with GGML_HIP=ON.

🤖 Generated with Claude Code

@context7

context7 Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit 6a4d800

bong-water-water-bong and others added 4 commits September 25, 2026 22:01
… llama.cpp

`1bit serve --device vulkan` reads the GGUF's general.architecture (app/gguf_meta.h) and
runs an architecture only our llama.cpp implements (zaya) on the HRX build's Vulkan0, even
when the upstream Vulkan build is present. registry_build counts such architectures for
vulkan the same way, and check_models.tsv checks ZAYA1-8B Q4_K_M on vulkan.

docs/vulkan.md: converting and serving ZAYA1, with the checks against transformers and
the measured speeds on Strix Halo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egistry

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ama.cpp #2, #5, #6)

The pin adds the zaya architecture and its converter (#2, #5) on top of #104's #95 fixes, and
carries #106's scalar-scatter SET_ROWS onto the tracked branch (#6): #106 pinned c358334,
which is on no branch of the fork. registry/architectures.json, regenerated:
ZayaForCausalLM maps to hrx and vulkan (266 architectures, all mapped).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1-8B included)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong bong-water-water-bong changed the title ZAYA1 (Zyphra) on the GPU: pin llama.cpp e2e6ca2, serve routes fork-only architectures ZAYA1 (Zyphra) on the GPU: pin llama.cpp 5556bf2, serve routes fork-only architectures Sep 26, 2026
@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

PR Reviewer Guide 🔍

(Review updated until commit 6a4d800)

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

2 - Partially compliant

Compliant requirements:

  • Tokenizer implementation matches HF tokenizers exactly
  • All verification requirements met with zero mismatches
  • Mutation sweep confirms correctness

Non-compliant requirements:

  • None

Requires further human verification:

  • None

5 - Partially compliant

Compliant requirements:

  • Removed engine/core and engine/backends/cpu
  • Removed tests and golden tools
  • Updated README, PORTING.md and contributing rules
  • Kept license, NOTICE, copyright check and CI

Non-compliant requirements:

  • None

Requires further human verification:

  • None

104 - Partially compliant

Compliant requirements:

  • Pin llama.cpp to 96f6b89
  • Fixed GLM-4.7-Flash and Qwen3-Coder at full context
  • Fixed non-FA MoE graph compile failures
  • Fixed MUL_MAT_ID reading route ids
  • Fixed FA value_stride range
  • Added debug switches
  • Removed 32768 context default for --device hrx
  • Verified on Strix Halo with 7/7 models passing
  • Verified serve_e2e passes
  • Verified test-backend-ops passes
  • Updated docs and registry

Non-compliant requirements:

  • None

Requires further human verification:

  • None
⏱️ Estimated effort to review: 4 🔵🔵🔵🔵⚪
🧪 PR contains tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Architecture routing logic

The new fork_only_arch function in app/serve.cpp introduces logic to route specific architectures (currently only "zaya") to the HRX build's Vulkan0 even when the upstream Vulkan build is present. This could potentially cause issues if the architecture is not properly handled by the HRX build or if there are conflicts in the routing logic. The function is used in the Launch launch_for function to determine which server to use, and it's important that this logic is robust and doesn't introduce unexpected behavior.

// Architectures our llama.cpp (third_party/llama.cpp, the HRX build) implements and upstream's
// does not: a GGUF of one runs on that build's Vulkan0 even when the upstream build is present.
bool fork_only_arch(const std::string& gguf) {
    const std::string arch = gguf_architecture(gguf);
    return arch == "zaya";  // Zyphra ZAYA1 (docs/vulkan.md)
}
GGUF parsing robustness

The gguf_architecture function in app/gguf_meta.h reads GGUF headers to extract the architecture field. While it handles basic GGUF structure parsing, it could be more robust against malformed or malicious GGUF files. Specifically, it doesn't validate the maximum number of key-value pairs beyond 4096, which could lead to excessive memory usage or processing time if a file has many key-value pairs. Additionally, it doesn't check for potential integer overflows in the key length or value size calculations.

inline std::string gguf_architecture(const std::string& path) {
    std::ifstream f(path, std::ios::binary);
    auto rd = [&](void* p, size_t n) { return static_cast<bool>(f.read(static_cast<char*>(p), n)); };
    uint32_t magic = 0, version = 0;
    uint64_t n_tensors = 0, n_kv = 0;
    if (!rd(&magic, 4) || magic != 0x46554747u /* "GGUF" */ || !rd(&version, 4) || version < 2 ||
        !rd(&n_tensors, 8) || !rd(&n_kv, 8))
        return "";
    auto str = [&](std::string* out) {
        uint64_t n = 0;
        if (!rd(&n, 8) || n > (1u << 20)) return false;
        if (!out) return static_cast<bool>(f.seekg(static_cast<std::streamoff>(n), std::ios::cur));
        out->resize(n);
        return rd(out->data(), n);
    };
    // bytes of each scalar value type: u8 i8 u16 i16 u32 i32 f32 bool (string) (array) u64 i64 f64
    auto scalar = [](uint32_t t) -> int {
        static const int size[] = {1, 1, 2, 2, 4, 4, 4, 1, 0, 0, 8, 8, 8};
        return t < 13 ? size[t] : -1;
    };
    for (uint64_t i = 0; i < n_kv && i < 4096; ++i) {
        std::string key;
        uint32_t type = 0;
        if (!str(&key) || !rd(&type, 4)) return "";
        if (type == 8) {  // string
            std::string v;
            if (!str(key == "general.architecture" ? &v : nullptr)) return "";
            if (key == "general.architecture") return v;
        } else if (type == 9) {  // array: element type, count, elements
            uint32_t et = 0;
            uint64_t n = 0;
            if (!rd(&et, 4) || !rd(&n, 8)) return "";
            if (et == 8) {
                for (uint64_t j = 0; j < n; ++j)
                    if (!str(nullptr)) return "";
            } else if (scalar(et) > 0) {
                f.seekg(static_cast<std::streamoff>(n * scalar(et)), std::ios::cur);
            } else {
                return "";  // nested arrays do not occur in GGUF metadata
            }
        } else if (scalar(type) > 0) {
            f.seekg(scalar(type), std::ios::cur);
        } else {
            return "";
        }
        if (!f) return "";
    }
    return "";
}
Test coverage for edge cases

The gguf_meta_test.cpp test file provides good coverage for basic GGUF parsing, but it doesn't test edge cases like malformed GGUF files with invalid magic numbers, negative key lengths, or very large key-value pairs that could potentially cause buffer overflows or memory issues. It would be beneficial to add tests for these edge cases to ensure robustness.

// Copyright 2026 bong-water-water-bong
// SPDX-License-Identifier: Apache-2.0
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
//     http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.

// gguf_architecture (app/gguf_meta.h) on GGUF headers written here: the key found after
// scalars and arrays, and "" for files that are not GGUFs or carry no architecture.
#include "gguf_meta.h"

#include <cstdio>
#include <filesystem>
#include <fstream>
#include <string>
#include <vector>

namespace {

struct Gguf {
    std::string b;
    uint64_t n_kv = 0;
    template <class T> void raw(T v) { b.append(reinterpret_cast<const char*>(&v), sizeof v); }
    void str(const std::string& s) { raw<uint64_t>(s.size()); b += s; }
    void kv_u32(const std::string& k, uint32_t v) { str(k); raw<uint32_t>(4); raw(v); ++n_kv; }
    void kv_str(const std::string& k, const std::string& v) { str(k); raw<uint32_t>(8); str(v); ++n_kv; }
    void kv_strs(const std::string& k, const std::vector<std::string>& v) {
        str(k); raw<uint32_t>(9); raw<uint32_t>(8); raw<uint64_t>(v.size());
        for (const auto& s : v) str(s);
        ++n_kv;
    }
    void kv_f32s(const std::string& k, size_t n) {
        str(k); raw<uint32_t>(9); raw<uint32_t>(6); raw<uint64_t>(n);
        for (size_t i = 0; i < n; ++i) raw<float>(0.5f * i);
        ++n_kv;
    }
    std::string write(const std::string& path) const {
        std::ofstream f(path, std::ios::binary);
        const uint32_t magic = 0x46554747u, version = 3;
        const uint64_t n_tensors = 0;
        f.write(reinterpret_cast<const char*>(&magic), 4);
        f.write(reinterpret_cast<const char*>(&version), 4);
        f.write(reinterpret_cast<const char*>(&n_tensors), 8);
        f.write(reinterpret_cast<const char*>(&n_kv), 8);
        f << b;
        return path;
    }
};

int failures = 0;
void expect(const std::string& what, const std::string& got, const std::string& want) {
    const bool ok = got == want;
    std::printf("%s %s: \"%s\"\n", ok ? "ok  " : "FAIL", what.c_str(), got.c_str());
    failures += !ok;
}

}  // namespace

int main() {
    const auto dir = std::filesystem::temp_directory_path() / "onebit_gguf_meta_test";
    std::filesystem::create_directories(dir);

    Gguf first;
    first.kv_str("general.architecture", "qwen3");
    expect("architecture first", onebit::gguf_architecture(first.write(dir / "first.gguf")), "qwen3");

    Gguf later;
    later.kv_u32("general.quantization_version", 2);
    later.kv_strs("tokenizer.ggml.tokens", {"<bos>", "\n", "hello"});
    later.kv_f32s("tokenizer.ggml.scores", 3);
    later.kv_str("general.name", "ZAYA1 8B");
    later.kv_str("general.architecture", "zaya");
    expect("architecture after arrays", onebit::gguf_architecture(later.write(dir / "later.gguf")), "zaya");

    Gguf none;
    none.kv_u32("general.quantization_version", 2);
    expect("no architecture", onebit::gguf_architecture(none.write(dir / "none.gguf")), "");

    { std::ofstream(dir / "text.gguf") << "not a gguf file at all"; }
    expect("not a GGUF", onebit::gguf_architecture((dir / "text.gguf").string()), "");
    expect("missing file", onebit::gguf_architecture((dir / "missing.gguf").string()), "");

    std::filesystem::remove_all(dir);
    return failures ? 1 : 0;
}

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Persistent review updated to latest commit 6a4d800

@bong-water-water-bong
bong-water-water-bong merged commit 2f1c45f into main Sep 26, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the zaya-vulkan branch September 26, 2026 01:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant