Skip to content

Model registry and the daily HF census (PORTING step 5) - #94

Merged
bong-water-water-bong merged 2 commits into
mainfrom
registry/census
Sep 25, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
registry/census

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Step 5: which HF models this engine can run, kept current daily. The two counts are reported separately and never added together.

Registry. registry/architectures.json is generated by tools/registry_build.py from the pinned sources. Nothing in it is typed in by hand:

  • llama.cpp's converter (HF -> GGUF architecture);
  • src/llama-arch.cpp in the upstream pin (vulkan) and in the HRX fork (hrx);
  • ZINC's parseArchitecture;
  • the NPU fast lane's qwen3.

At the current pins: 265 HF architectures. vulkan maps 265, hrx 246, zinc 44, npu 1.

Checked. tools/registry_check.py runs tests/serve_e2e.sh over registry/check_models.tsv and writes registry/checked.json, failures included. On Strix Halo, 17 of 20 backend-model runs pass across seven GGUF architectures. The three failures:

  • HRX on Qwen3-Coder-30B-A3B and GLM-4.7-Flash: HTTP 500 "Compute error.", reproduced on a second run;
  • ZINC on Qwen2.5-7B and MiniCPM5-1B: no answer.

Census. tools/census.py sweeps every HF text-generation page, and census.yml runs it daily at 05:17 UTC and opens a PR when anything changed. First sweep:

models share of those with an architecture
text-generation models on HF 415,414
with an architecture 332,565 (2,610 architectures)
mapped 310,221 93.28%
checked 213,456 64.18%

"Checked" counts a model when a model of its architecture passed. It does not mean every counted model was run.

Docs: docs/registry.md (in the site nav), plus PORTING row 5 and a README status line.

The workflow has not run in CI yet. Its steps were run locally: registry_build.py, the full census.py sweep, and census.py --pr-body.

🤖 Generated with Claude Code

bong-water-water-bong and others added 2 commits September 25, 2026 16:29
registry/architectures.json maps each HF architecture to its GGUF architecture and the
backends whose pinned code accepts it (tools/registry_build.py, from llama.cpp's
converter and llama-arch.cpp in the upstream and HRX pins, ZINC's parseArchitecture,
and the NPU fast lane's model type). tools/registry_check.py runs serve_e2e per
backend and model and records registry/checked.json, failures included.
tools/census.py sweeps every HF text-generation model; census.yml runs it daily and
opens a PR. Mapped and checked coverage are reported separately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong enabled auto-merge (squash) September 25, 2026 19:42
@context7

context7 Bot commented Sep 25, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit c298263

@bong-water-water-bong
bong-water-water-bong merged commit 67706c8 into main Sep 25, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the registry/census branch September 25, 2026 19:45
@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 4 🔵🔵🔵🔵⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Incorrect GGUF architecture resolution

The hf_to_gguf function attempts to resolve GGUF architectures from class inheritance, but it does not correctly handle cases where a class inherits from multiple base classes. This can lead to incorrect architecture mappings when the resolution logic picks an unintended base class. The function should be more robust in handling complex inheritance hierarchies.

def resolve(name, seen=()):
    if name not in classes or name in seen:
        return None
    bases, arch, _ = classes[name]
    if arch:
        return arch
    for b in bases:
        a = resolve(b, seen + (name,))
        if a:
            return a
    return None

out = {}
for cname, (_, _, hf) in classes.items():
    arch = resolve(cname)
    if not arch or arch == "MMPROJ" or arch not in names:
        continue  # vision/audio projector classes carry no text architecture
    for h in hf:
        out.setdefault(h, names[arch])
return out
Incomplete error handling for API requests

The sweep function retries API requests up to 6 times with exponential backoff, but it does not handle cases where the API returns a non-200 status code that is not a transient error. This could lead to silent failures or incorrect model counts if the API returns an error that should be treated as a permanent failure.

for attempt in range(6):
    try:
        req = urllib.request.Request(url, headers={"User-Agent": "1bit-engine-census"})
        with urllib.request.urlopen(req, timeout=120) as r:
            batch = json.load(r)
            link = r.headers.get("Link", "")
        break
    except Exception as e:  # rate limits and transient errors: back off and retry
        if attempt == 5:
            raise
        print(f"page {pages + 1}: {e}; retrying", file=sys.stderr)
        time.sleep(30 * (attempt + 1))
for m in batch:
Potential race condition in result updates

The registry_check.py script updates registry/checked.json after each model check, but it does not use file locking or atomic operations. If multiple instances of the script run concurrently, it could lead to data corruption or loss of results.

print(f"{'PASS' if ok else 'FAIL'} {backend:6} {arch:10} {model}  said: {said[:60]!r}")
data["results"] = sorted(results.values(), key=lambda r: (r["arch"], r["backend"], r["model"]))
with open(OUT, "w") as f:  # after every run, so an interrupted sweep keeps what finished
    json.dump(data, f, indent=1)
    f.write("\n")

⚠️ Review coverage: The following files were not included in this review because of the token budget:

  • registry/census.json
  • .github/workflows/census.yml
  • README.md

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant