Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
84 changes: 84 additions & 0 deletions .github/workflows/census.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# Copyright 2026 bong-water-water-bong
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#
# The daily HF census (docs/registry.md): rebuild registry/architectures.json from the
# pinned llama.cpp, HRX fork and ZINC, sweep every HF text-generation model, and open a
# PR with the new counts. The mapped and checked counts are reported separately;
# registry/checked.json only changes when tools/registry_check.py runs on Strix Halo.
#
# Uses the same HRX_BUMP_TOKEN as the bump workflows (a PR opened with GITHUB_TOKEN
# would not run CI).

name: census

on:
schedule:
- cron: "17 5 * * *"
workflow_dispatch:

permissions:
contents: read

concurrency:
group: census
cancel-in-progress: false

jobs:
census:
runs-on: ubuntu-latest
timeout-minutes: 120
steps:
- name: Require the token
env:
HRX_BUMP_TOKEN: ${{ secrets.HRX_BUMP_TOKEN }}
run: |
if [ -z "$HRX_BUMP_TOKEN" ]; then
echo "::error::secret HRX_BUMP_TOKEN is not set (see the header of this workflow)"
exit 1
fi

- uses: actions/checkout@v4
with:
token: ${{ secrets.HRX_BUMP_TOKEN }}

- name: Check out the pinned sources the registry reads
run: git submodule update --init --depth 1 third_party/llama.cpp-vulkan third_party/llama.cpp third_party/zinc

- name: Rebuild the registry and sweep HF
run: |
set -euo pipefail
python3 tools/registry_build.py
python3 tools/census.py | tee census.txt

- name: Open the census PR
env:
GH_TOKEN: ${{ secrets.HRX_BUMP_TOKEN }}
run: |
set -euo pipefail
if git diff --quiet -- registry; then echo "no change"; exit 0; fi
day=$(date -u +%F)
branch="census/$day"
if git ls-remote --exit-code origin "refs/heads/$branch" > /dev/null; then
echo "$branch already exists"; exit 0
fi
summary=$(cat census.txt)
python3 tools/census.py --pr-body > body.md
git config user.name "census"
git config user.email "census@users.noreply.github.com"
git switch -c "$branch"
git add registry
git commit -q -m "Census $day: $summary"
git push -q origin "$branch"
gh pr create --base main --head "$branch" --title "Census $day" --body-file body.md
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,10 @@ serves each model behind an OpenAI-compatible API (`1bit serve`), whatever devic
> 16.3-16.5 tok/s decode ([docs/npu.md](docs/npu.md#private-routes)). Step 4, the Laya router,
> has landed as an opt-in: `1bit serve --device auto --laya-model <dir>` picks the device per
> request with the C++ scorer, which matches the Python reference; a decision takes 0.38 s on
> Strix Halo ([docs/laya.md](docs/laya.md)). The working engine is being ported from 1bit-MONSTER,
> Strix Halo ([docs/laya.md](docs/laya.md)). Step 5, the model registry, has landed: of 332,565
> HF text-generation models with an architecture, 93.28% are mapped to a backend and 64.18%
> have an architecture checked end to end on Strix Halo; a daily census keeps the counts
> current ([docs/registry.md](docs/registry.md)). The working engine is being ported from 1bit-MONSTER,
> our private development repository; see [docs/PORTING.md](docs/PORTING.md).
> Measured results are on the [wiki](https://github.com/1bit-MONSTER/engine/wiki). This repository
> holds the verified code without the development history.
Expand Down
2 changes: 1 addition & 1 deletion docs/PORTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ Each step below is one PR (or a short series) that builds and runs on Strix Halo
| 3x | **35B MoE on the NPU (experimental, closed source).** Qwen3.6-35B-A3B on the NPU | the private `1bit-MONSTER/npu-kernels` repository | **landed as a private add-on** ([docs/npu.md](npu.md#private-routes)): built into `1bit` with `-DONEBIT_NPU_PRIVATE`; parity against the fp64 reference passes (3 positions, argmax 846 / 198 / 3710), 16.3-16.5 tok/s decode, and `1bit serve --device npu` answers "The capital of France is Paris." Without the add-on, `serve` says the route is not part of the build |
| + | **Linux kernel.** The kernel that provides `amdxdna` and `amdgpu`, pinned to upstream | `torvalds/linux` release tags; config from the Strix Halo kernel of 2026-09-23 | **pinned** ([docs/kernel.md](kernel.md)): v7.3-rc4 builds into Debian packages with `amdxdna` in-tree; kept current by `bump-linux.yml`. Installing it on Strix Halo is a separate, deliberate step |
| 4 | **Laya router.** A non-autoregressive scorer that picks where each request runs | `src/laya_scorer.cpp`, `include/laya_scorer.h` on `backup/laya-and-results-2026-09-22`; model at `~/models/laya` | **landed, opt-in** ([docs/laya.md](laya.md)): source `NandhaKishorM/laya` + the three HF checkpoints pinned, hash-verified fetch, `bump-laya.yml` (#18); the C++ scorer matches the Python reference on the root and `typed-decisions/` checkpoints (max logit diff 8.6e-6, same argmax) and `1bit serve --device auto --laya-model <dir>` routes each request (#90). Load 2.8 s once, then 0.38 s per decision (was 8.75 s, #91). Next: `multilingual/` (mmBERT-base) |
| 5 | **Every HF model, kept current.** The architecture registry (569 tokens mapping 2,030 HF arch strings) and the daily HF census that finds new architectures and proposes mappings | `src/model_registry*.cpp`, `Testing/census_*.py` and `.json`, `.github/workflows/census-{watch,sweep,autopr}.yml` | the census runs daily in CI; docs report *mapped* and *run and checked* counts separately |
| 5 | **Every HF model, kept current.** The architecture registry and the daily HF census that finds new architectures and ranks the unmapped ones by model count (1bit-MONSTER's registry mapped 2,030 HF arch strings to its own kernels; here the backends' pinned code decides) | `src/model_registry*.cpp`, `Testing/census_*.py` and `.json`, `.github/workflows/census-{watch,sweep,autopr}.yml` | **landed** ([docs/registry.md](registry.md)): `registry/architectures.json` generated from the pinned llama.cpp, HRX fork and ZINC (265 HF architectures); `census.yml` sweeps HF daily. First sweep: 415,414 text-generation models, mapped 93.28%, checked 64.18% of those with an architecture (reported separately). HRX fails Qwen3-Coder-30B-A3B and GLM-4.7-Flash with a compute error |
| 6 | **ZINC (NVIDIA and more).** Upstream `zolotukhin/zinc`, a Zig GGUF engine with Vulkan, ROCm, CUDA and Metal backends; its CUDA backend reaches NVIDIA GPUs (Ada `sm_89`, Blackwell `sm_120`) | not in 1bit-MONSTER; pinned from upstream `main` | **pinned** ([docs/zinc.md](zinc.md)): `scripts/build-zinc.sh` builds it privately; the Vulkan build gives 12095 (" Paris") at 295 tok/s on Strix Halo; the CUDA build answers " Paris." at 167–173 tok/s on an RTX 5090 (Qwen3.5-9B); kept current by `bump-zinc.yml`. `1bit serve --device zinc` runs it (`-DONEBIT_ZINC=ON`, e2e passes) |

## How the pieces fit
Expand Down
101 changes: 101 additions & 0 deletions docs/registry.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
<!--
Copyright 2026 bong-water-water-bong
SPDX-License-Identifier: Apache-2.0

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
-->
# Model registry and HF census

Step 5 of the port ([PORTING.md](PORTING.md)): which Hugging Face models this engine can
run, kept current every day. Two numbers are reported, and they are never added together:

- **Mapped:** a backend's own code accepts the model's architecture.
- **Checked:** a model of that architecture loaded, answered and streamed through
`1bit serve` on Strix Halo (`tests/serve_e2e.sh`).

## The registry

`registry/architectures.json` maps each HF architecture (a config's `architectures[0]`,
such as `Qwen3ForCausalLM`) to its GGUF architecture and the backends that accept it.
`tools/registry_build.py` generates it from the pinned sources. None of it is typed in by
hand:

| Backend | Accepts the architecture when |
|---|---|
| HF -> GGUF | a `@ModelBase.register(...)` class in llama.cpp's converter names it (upstream pin, then the HRX fork) |
| `vulkan` | the GGUF architecture is in the upstream pin's `src/llama-arch.cpp` |
| `hrx` | the GGUF architecture is in our HRX fork's `src/llama-arch.cpp` |
| `zinc` | ZINC's `parseArchitecture` accepts the GGUF architecture |
| `npu` | the fast lane's model type: `qwen3` (the lane kernels are built for Qwen3-0.6B's shapes) |

At the current pins: 265 HF architectures, vulkan 265, hrx 246, zinc 44, npu 1.

## Checked models

`tools/registry_check.py <1bit> --models <dir>` runs `tests/serve_e2e.sh` for every row of
`registry/check_models.tsv` and records the result in `registry/checked.json`, failures
included. The architecture recorded is the one the backend loads: the GGUF
`general.architecture`, or the NPU directory's `model_type`. Checked on Strix Halo
2026-09-25, engine `8c2805d`:

| GGUF architecture | Model | vulkan | hrx | zinc | npu |
|---|---|---|---|---|---|
| `qwen3` | Qwen3-0.6B | pass | pass | pass | pass |
| `qwen2` | Qwen2.5-7B-Instruct | pass | pass | fails: no answer | |
| `qwen3moe` | Qwen3-Coder-30B-A3B | pass | fails: compute error | pass | |
| `qwen35moe` | Qwen3.6-35B-A3B Q8_0 | pass | pass | pass | |
| `deepseek2` | GLM-4.7-Flash | pass | fails: compute error | not mapped | |
| `minicpm` | MiniCPM4-8B | pass | pass | not mapped | |
| `llama` | MiniCPM5-1B | pass | pass | fails: no answer | |

The two HRX compute errors reproduce on a second run. Both are MoE models, and the chat
request itself fails with HTTP 500 "Compute error."

## The census

`tools/census.py` walks every page of the HF API's text-generation listing (with each
model's config inline) and counts models by architecture. `registry/census.json` keeps the
counts and the coverage read against the registry and the checked results. The
`census.yml` workflow runs daily at 05:17 UTC. It rebuilds the registry from the pins,
sweeps HF, and opens a PR when anything changed.

The first full sweep, 2026-09-25:

| | models | share of those with an architecture |
|---|---|---|
| Text-generation models on HF | 415,414 | |
| With an architecture in their config | 332,565 (2,610 architectures) | |
| Mapped | 310,221 | 93.28% |
| Checked | 213,456 | 64.18% |

| Backend | Mapped | Checked |
|---|---|---|
| vulkan | 93.28% | 64.18% |
| hrx | 93.19% | 63.39% |
| zinc | 69.69% | 10.28% |
| npu | 9.29% | 9.29% |

"Checked" counts every model whose architecture has a passing model. It does not mean each
of those models was run. The unmapped architectures with the most models are
`Step1MoEForCausalLM` (2,882), `OPTForCausalLM` (2,097), `ParlerTTSForConditionalGeneration`
(1,586) and `GPTNeoForCausalLM` (1,565). `tools/census.py --pr-body` lists the top ten.

## Commands

```
tools/registry_build.py # rebuild registry/architectures.json from the pins
tools/registry_build.py --check # exit 1 if it is stale
tools/registry_check.py build/1bit --models ~/models # run the checks on Strix Halo
tools/census.py # full HF sweep (about 420 pages)
tools/census.py --report # recompute coverage from the saved counts
```
Loading
Loading