diff --git a/.github/workflows/census.yml b/.github/workflows/census.yml new file mode 100644 index 00000000..1a7e399a --- /dev/null +++ b/.github/workflows/census.yml @@ -0,0 +1,84 @@ +# Copyright 2026 bong-water-water-bong +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. +# +# The daily HF census (docs/registry.md): rebuild registry/architectures.json from the +# pinned llama.cpp, HRX fork and ZINC, sweep every HF text-generation model, and open a +# PR with the new counts. The mapped and checked counts are reported separately; +# registry/checked.json only changes when tools/registry_check.py runs on Strix Halo. +# +# Uses the same HRX_BUMP_TOKEN as the bump workflows (a PR opened with GITHUB_TOKEN +# would not run CI). + +name: census + +on: + schedule: + - cron: "17 5 * * *" + workflow_dispatch: + +permissions: + contents: read + +concurrency: + group: census + cancel-in-progress: false + +jobs: + census: + runs-on: ubuntu-latest + timeout-minutes: 120 + steps: + - name: Require the token + env: + HRX_BUMP_TOKEN: ${{ secrets.HRX_BUMP_TOKEN }} + run: | + if [ -z "$HRX_BUMP_TOKEN" ]; then + echo "::error::secret HRX_BUMP_TOKEN is not set (see the header of this workflow)" + exit 1 + fi + + - uses: actions/checkout@v4 + with: + token: ${{ secrets.HRX_BUMP_TOKEN }} + + - name: Check out the pinned sources the registry reads + run: git submodule update --init --depth 1 third_party/llama.cpp-vulkan third_party/llama.cpp third_party/zinc + + - name: Rebuild the registry and sweep HF + run: | + set -euo pipefail + python3 tools/registry_build.py + python3 tools/census.py | tee census.txt + + - name: Open the census PR + env: + GH_TOKEN: ${{ secrets.HRX_BUMP_TOKEN }} + run: | + set -euo pipefail + if git diff --quiet -- registry; then echo "no change"; exit 0; fi + day=$(date -u +%F) + branch="census/$day" + if git ls-remote --exit-code origin "refs/heads/$branch" > /dev/null; then + echo "$branch already exists"; exit 0 + fi + summary=$(cat census.txt) + python3 tools/census.py --pr-body > body.md + git config user.name "census" + git config user.email "census@users.noreply.github.com" + git switch -c "$branch" + git add registry + git commit -q -m "Census $day: $summary" + git push -q origin "$branch" + gh pr create --base main --head "$branch" --title "Census $day" --body-file body.md diff --git a/README.md b/README.md index 3f6cf705..43b90a07 100644 --- a/README.md +++ b/README.md @@ -46,7 +46,10 @@ serves each model behind an OpenAI-compatible API (`1bit serve`), whatever devic > 16.3-16.5 tok/s decode ([docs/npu.md](docs/npu.md#private-routes)). Step 4, the Laya router, > has landed as an opt-in: `1bit serve --device auto --laya-model ` picks the device per > request with the C++ scorer, which matches the Python reference; a decision takes 0.38 s on -> Strix Halo ([docs/laya.md](docs/laya.md)). The working engine is being ported from 1bit-MONSTER, +> Strix Halo ([docs/laya.md](docs/laya.md)). Step 5, the model registry, has landed: of 332,565 +> HF text-generation models with an architecture, 93.28% are mapped to a backend and 64.18% +> have an architecture checked end to end on Strix Halo; a daily census keeps the counts +> current ([docs/registry.md](docs/registry.md)). The working engine is being ported from 1bit-MONSTER, > our private development repository; see [docs/PORTING.md](docs/PORTING.md). > Measured results are on the [wiki](https://github.com/1bit-MONSTER/engine/wiki). This repository > holds the verified code without the development history. diff --git a/docs/PORTING.md b/docs/PORTING.md index 5830258a..7e193459 100644 --- a/docs/PORTING.md +++ b/docs/PORTING.md @@ -27,7 +27,7 @@ Each step below is one PR (or a short series) that builds and runs on Strix Halo | 3x | **35B MoE on the NPU (experimental, closed source).** Qwen3.6-35B-A3B on the NPU | the private `1bit-MONSTER/npu-kernels` repository | **landed as a private add-on** ([docs/npu.md](npu.md#private-routes)): built into `1bit` with `-DONEBIT_NPU_PRIVATE`; parity against the fp64 reference passes (3 positions, argmax 846 / 198 / 3710), 16.3-16.5 tok/s decode, and `1bit serve --device npu` answers "The capital of France is Paris." Without the add-on, `serve` says the route is not part of the build | | + | **Linux kernel.** The kernel that provides `amdxdna` and `amdgpu`, pinned to upstream | `torvalds/linux` release tags; config from the Strix Halo kernel of 2026-09-23 | **pinned** ([docs/kernel.md](kernel.md)): v7.3-rc4 builds into Debian packages with `amdxdna` in-tree; kept current by `bump-linux.yml`. Installing it on Strix Halo is a separate, deliberate step | | 4 | **Laya router.** A non-autoregressive scorer that picks where each request runs | `src/laya_scorer.cpp`, `include/laya_scorer.h` on `backup/laya-and-results-2026-09-22`; model at `~/models/laya` | **landed, opt-in** ([docs/laya.md](laya.md)): source `NandhaKishorM/laya` + the three HF checkpoints pinned, hash-verified fetch, `bump-laya.yml` (#18); the C++ scorer matches the Python reference on the root and `typed-decisions/` checkpoints (max logit diff 8.6e-6, same argmax) and `1bit serve --device auto --laya-model ` routes each request (#90). Load 2.8 s once, then 0.38 s per decision (was 8.75 s, #91). Next: `multilingual/` (mmBERT-base) | -| 5 | **Every HF model, kept current.** The architecture registry (569 tokens mapping 2,030 HF arch strings) and the daily HF census that finds new architectures and proposes mappings | `src/model_registry*.cpp`, `Testing/census_*.py` and `.json`, `.github/workflows/census-{watch,sweep,autopr}.yml` | the census runs daily in CI; docs report *mapped* and *run and checked* counts separately | +| 5 | **Every HF model, kept current.** The architecture registry and the daily HF census that finds new architectures and ranks the unmapped ones by model count (1bit-MONSTER's registry mapped 2,030 HF arch strings to its own kernels; here the backends' pinned code decides) | `src/model_registry*.cpp`, `Testing/census_*.py` and `.json`, `.github/workflows/census-{watch,sweep,autopr}.yml` | **landed** ([docs/registry.md](registry.md)): `registry/architectures.json` generated from the pinned llama.cpp, HRX fork and ZINC (265 HF architectures); `census.yml` sweeps HF daily. First sweep: 415,414 text-generation models, mapped 93.28%, checked 64.18% of those with an architecture (reported separately). HRX fails Qwen3-Coder-30B-A3B and GLM-4.7-Flash with a compute error | | 6 | **ZINC (NVIDIA and more).** Upstream `zolotukhin/zinc`, a Zig GGUF engine with Vulkan, ROCm, CUDA and Metal backends; its CUDA backend reaches NVIDIA GPUs (Ada `sm_89`, Blackwell `sm_120`) | not in 1bit-MONSTER; pinned from upstream `main` | **pinned** ([docs/zinc.md](zinc.md)): `scripts/build-zinc.sh` builds it privately; the Vulkan build gives 12095 (" Paris") at 295 tok/s on Strix Halo; the CUDA build answers " Paris." at 167–173 tok/s on an RTX 5090 (Qwen3.5-9B); kept current by `bump-zinc.yml`. `1bit serve --device zinc` runs it (`-DONEBIT_ZINC=ON`, e2e passes) | ## How the pieces fit diff --git a/docs/registry.md b/docs/registry.md new file mode 100644 index 00000000..40ee22d9 --- /dev/null +++ b/docs/registry.md @@ -0,0 +1,101 @@ + +# Model registry and HF census + +Step 5 of the port ([PORTING.md](PORTING.md)): which Hugging Face models this engine can +run, kept current every day. Two numbers are reported, and they are never added together: + +- **Mapped:** a backend's own code accepts the model's architecture. +- **Checked:** a model of that architecture loaded, answered and streamed through + `1bit serve` on Strix Halo (`tests/serve_e2e.sh`). + +## The registry + +`registry/architectures.json` maps each HF architecture (a config's `architectures[0]`, +such as `Qwen3ForCausalLM`) to its GGUF architecture and the backends that accept it. +`tools/registry_build.py` generates it from the pinned sources. None of it is typed in by +hand: + +| Backend | Accepts the architecture when | +|---|---| +| HF -> GGUF | a `@ModelBase.register(...)` class in llama.cpp's converter names it (upstream pin, then the HRX fork) | +| `vulkan` | the GGUF architecture is in the upstream pin's `src/llama-arch.cpp` | +| `hrx` | the GGUF architecture is in our HRX fork's `src/llama-arch.cpp` | +| `zinc` | ZINC's `parseArchitecture` accepts the GGUF architecture | +| `npu` | the fast lane's model type: `qwen3` (the lane kernels are built for Qwen3-0.6B's shapes) | + +At the current pins: 265 HF architectures, vulkan 265, hrx 246, zinc 44, npu 1. + +## Checked models + +`tools/registry_check.py <1bit> --models ` runs `tests/serve_e2e.sh` for every row of +`registry/check_models.tsv` and records the result in `registry/checked.json`, failures +included. The architecture recorded is the one the backend loads: the GGUF +`general.architecture`, or the NPU directory's `model_type`. Checked on Strix Halo +2026-09-25, engine `8c2805d`: + +| GGUF architecture | Model | vulkan | hrx | zinc | npu | +|---|---|---|---|---|---| +| `qwen3` | Qwen3-0.6B | pass | pass | pass | pass | +| `qwen2` | Qwen2.5-7B-Instruct | pass | pass | fails: no answer | | +| `qwen3moe` | Qwen3-Coder-30B-A3B | pass | fails: compute error | pass | | +| `qwen35moe` | Qwen3.6-35B-A3B Q8_0 | pass | pass | pass | | +| `deepseek2` | GLM-4.7-Flash | pass | fails: compute error | not mapped | | +| `minicpm` | MiniCPM4-8B | pass | pass | not mapped | | +| `llama` | MiniCPM5-1B | pass | pass | fails: no answer | | + +The two HRX compute errors reproduce on a second run. Both are MoE models, and the chat +request itself fails with HTTP 500 "Compute error." + +## The census + +`tools/census.py` walks every page of the HF API's text-generation listing (with each +model's config inline) and counts models by architecture. `registry/census.json` keeps the +counts and the coverage read against the registry and the checked results. The +`census.yml` workflow runs daily at 05:17 UTC. It rebuilds the registry from the pins, +sweeps HF, and opens a PR when anything changed. + +The first full sweep, 2026-09-25: + +| | models | share of those with an architecture | +|---|---|---| +| Text-generation models on HF | 415,414 | | +| With an architecture in their config | 332,565 (2,610 architectures) | | +| Mapped | 310,221 | 93.28% | +| Checked | 213,456 | 64.18% | + +| Backend | Mapped | Checked | +|---|---|---| +| vulkan | 93.28% | 64.18% | +| hrx | 93.19% | 63.39% | +| zinc | 69.69% | 10.28% | +| npu | 9.29% | 9.29% | + +"Checked" counts every model whose architecture has a passing model. It does not mean each +of those models was run. The unmapped architectures with the most models are +`Step1MoEForCausalLM` (2,882), `OPTForCausalLM` (2,097), `ParlerTTSForConditionalGeneration` +(1,586) and `GPTNeoForCausalLM` (1,565). `tools/census.py --pr-body` lists the top ten. + +## Commands + +``` +tools/registry_build.py # rebuild registry/architectures.json from the pins +tools/registry_build.py --check # exit 1 if it is stale +tools/registry_check.py build/1bit --models ~/models # run the checks on Strix Halo +tools/census.py # full HF sweep (about 420 pages) +tools/census.py --report # recompute coverage from the saved counts +``` diff --git a/registry/architectures.json b/registry/architectures.json new file mode 100644 index 00000000..bcf0c50d --- /dev/null +++ b/registry/architectures.json @@ -0,0 +1,1900 @@ +{ + "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", + "sources": { + "llama.cpp (vulkan)": "7fe450e19305b828c199d602c23a8337aaa1f03b", + "llama.cpp (hrx)": "c075cc1c66525f723de7b7d23ee331a70a732ce8", + "zinc": "3a35e76d64ebb91e2d82e16ddda20ee865ce1d45" + }, + "counts": { + "hrx": 246, + "npu": 1, + "vulkan": 265, + "zinc": 44, + "architectures": 265, + "mapped": 265 + }, + "architectures": { + "AfmoeForCausalLM": { + "gguf": "afmoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ApertusForCausalLM": { + "gguf": "apertus", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ArceeForCausalLM": { + "gguf": "arcee", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ArcticForCausalLM": { + "gguf": "arctic", + "backends": [ + "hrx", + "vulkan" + ] + }, + "AudioFlamingo3ForConditionalGeneration": { + "gguf": "qwen2", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "BaiChuanForCausalLM": { + "gguf": "baichuan", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BaichuanForCausalLM": { + "gguf": "baichuan", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BailingMoeForCausalLM": { + "gguf": "bailingmoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BailingMoeV2ForCausalLM": { + "gguf": "bailingmoe2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BailingMoeV3ForCausalLM": { + "gguf": "bailingmoe3", + "backends": [ + "vulkan" + ] + }, + "BambaForCausalLM": { + "gguf": "granitehybrid", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BertForMaskedLM": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BertForSequenceClassification": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BertModel": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BitNetForCausalLM": { + "gguf": "bitnet", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BitnetForCausalLM": { + "gguf": "bitnet", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BloomForCausalLM": { + "gguf": "bloom", + "backends": [ + "hrx", + "vulkan" + ] + }, + "BloomModel": { + "gguf": "bloom", + "backends": [ + "hrx", + "vulkan" + ] + }, + "CamembertModel": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ChameleonForCausalLM": { + "gguf": "chameleon", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ChameleonForConditionalGeneration": { + "gguf": "chameleon", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ChatGLMForConditionalGeneration": { + "gguf": "chatglm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ChatGLMModel": { + "gguf": "chatglm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "CodeShellForCausalLM": { + "gguf": "codeshell", + "backends": [ + "hrx", + "vulkan" + ] + }, + "CogVLMForCausalLM": { + "gguf": "cogvlm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Cohere2ForCausalLM": { + "gguf": "cohere2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Cohere2MoeForCausalLM": { + "gguf": "cohere2moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "CohereForCausalLM": { + "gguf": "command-r", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DFlash2DraftModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DFlashDraftModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DSparkDraftModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DSparkSpeculator": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DbrxForCausalLM": { + "gguf": "dbrx", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeciLMForCausalLM": { + "gguf": "deci", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeepseekForCausalLM": { + "gguf": "deepseek", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeepseekOCRForCausalLM": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeepseekV2ForCausalLM": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeepseekV32ForCausalLM": { + "gguf": "deepseek32", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeepseekV3ForCausalLM": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeepseekV4DSparkModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DeepseekV4ForCausalLM": { + "gguf": "deepseek4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DistilBertForMaskedLM": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DistilBertForSequenceClassification": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "DistilBertModel": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Dots1ForCausalLM": { + "gguf": "dots1", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Dots3NoteForCausalLM": { + "gguf": "dots3note", + "backends": [ + "vulkan" + ] + }, + "Dots3NoteForConditionalGeneration": { + "gguf": "dots3note", + "backends": [ + "vulkan" + ] + }, + "Dots3NoteTextForCausalLM": { + "gguf": "dots3note", + "backends": [ + "vulkan" + ] + }, + "DotsOCRForCausalLM": { + "gguf": "qwen2", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "DreamModel": { + "gguf": "dream", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Eagle3DraftModel": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Eagle3LlamaForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Eagle3Speculator": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Ernie4_5ForCausalLM": { + "gguf": "ernie4_5", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Ernie4_5_ForCausalLM": { + "gguf": "ernie4_5", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Ernie4_5_MoeForCausalLM": { + "gguf": "ernie4_5-moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "EuroBertModel": { + "gguf": "eurobert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Exaone4ForCausalLM": { + "gguf": "exaone4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Exaone4_5_ForConditionalGeneration": { + "gguf": "exaone4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ExaoneForCausalLM": { + "gguf": "exaone", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ExaoneMoEForCausalLM": { + "gguf": "exaone-moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ExaoneMoeForCausalLM": { + "gguf": "exaone-moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "FalconForCausalLM": { + "gguf": "falcon", + "backends": [ + "hrx", + "vulkan" + ] + }, + "FalconH1ForCausalLM": { + "gguf": "falcon-h1", + "backends": [ + "hrx", + "vulkan" + ] + }, + "FalconMambaForCausalLM": { + "gguf": "mamba", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "GPT2LMHeadModel": { + "gguf": "gpt2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GPTBigCodeForCausalLM": { + "gguf": "starcoder", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GPTNeoXForCausalLM": { + "gguf": "gptneox", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GPTRefactForCausalLM": { + "gguf": "refact", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma2ForCausalLM": { + "gguf": "gemma2", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Gemma3ForCausalLM": { + "gguf": "gemma3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma3ForConditionalGeneration": { + "gguf": "gemma3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma3TextModel": { + "gguf": "gemma-embedding", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma3nForCausalLM": { + "gguf": "gemma3n", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma3nForConditionalGeneration": { + "gguf": "gemma3n", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma4AssistantForCausalLM": { + "gguf": "gemma4-assistant", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma4DSparkModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma4ForCausalLM": { + "gguf": "gemma4", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Gemma4ForConditionalGeneration": { + "gguf": "gemma4", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Gemma4UnifiedAssistantForCausalLM": { + "gguf": "gemma4-assistant", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Gemma4UnifiedForConditionalGeneration": { + "gguf": "gemma4", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "GemmaForCausalLM": { + "gguf": "gemma", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Glm4ForCausalLM": { + "gguf": "glm4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Glm4MoeForCausalLM": { + "gguf": "glm4moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Glm4MoeLiteForCausalLM": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Glm4vForConditionalGeneration": { + "gguf": "glm4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Glm4vMoeForConditionalGeneration": { + "gguf": "glm4moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GlmForCausalLM": { + "gguf": "chatglm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GlmMoeDsaForCausalLM": { + "gguf": "glm-dsa", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GlmOcrForConditionalGeneration": { + "gguf": "glm4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GptOssForCausalLM": { + "gguf": "gpt-oss", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "GraniteForCausalLM": { + "gguf": "granite", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GraniteMoeForCausalLM": { + "gguf": "granitemoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GraniteMoeHybridForCausalLM": { + "gguf": "granitehybrid", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GraniteMoeSWAForCausalLM": { + "gguf": "granite_swa", + "backends": [ + "vulkan" + ] + }, + "GraniteMoeSharedForCausalLM": { + "gguf": "granitemoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GraniteSWAForCausalLM": { + "gguf": "granite_swa", + "backends": [ + "vulkan" + ] + }, + "GraniteSwitchForCausalLM": { + "gguf": "graniteswitch", + "backends": [ + "vulkan" + ] + }, + "Grok1ForCausalLM": { + "gguf": "grok", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GrokForCausalLM": { + "gguf": "grok", + "backends": [ + "hrx", + "vulkan" + ] + }, + "GroveMoeForCausalLM": { + "gguf": "grovemoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "HYV3ForCausalLM": { + "gguf": "hy_v3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "HYV4ForCausalLM": { + "gguf": "hy_v4", + "backends": [ + "vulkan" + ] + }, + "HrmTextForCausalLM": { + "gguf": "hrm_text", + "backends": [ + "vulkan" + ] + }, + "HunYuanDenseV1ForCausalLM": { + "gguf": "hunyuan-dense", + "backends": [ + "hrx", + "vulkan" + ] + }, + "HunYuanMoEV1ForCausalLM": { + "gguf": "hunyuan-moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "HunYuanVLForConditionalGeneration": { + "gguf": "hunyuan_vl", + "backends": [ + "hrx", + "vulkan" + ] + }, + "IQuestCoderForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "InternLM2ForCausalLM": { + "gguf": "internlm2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "InternLM3ForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "JAISLMHeadModel": { + "gguf": "jais", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Jais2ForCausalLM": { + "gguf": "jais2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "JambaForCausalLM": { + "gguf": "jamba", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "JanusForConditionalGeneration": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "JinaBertForMaskedLM": { + "gguf": "jina-bert-v2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "JinaBertModel": { + "gguf": "jina-bert-v2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "JinaEmbeddingsV5Model": { + "gguf": "eurobert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "KORMoForCausalLM": { + "gguf": "qwen2", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "KimiK25ForConditionalGeneration": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "KimiK3ForConditionalGeneration": { + "gguf": "kimi-k3", + "backends": [ + "vulkan" + ] + }, + "KimiLinearForCausalLM": { + "gguf": "kimi-linear", + "backends": [ + "hrx", + "vulkan" + ] + }, + "KimiLinearModel": { + "gguf": "kimi-linear", + "backends": [ + "hrx", + "vulkan" + ] + }, + "KimiVLForConditionalGeneration": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LFM2ForCausalLM": { + "gguf": "lfm2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LLaDAMoEModel": { + "gguf": "llada-moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LLaDAMoEModelLM": { + "gguf": "llada-moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LLaDAModelLM": { + "gguf": "llada", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LLaMAForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "LagunaForCausalLM": { + "gguf": "laguna", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Lfm25AudioTokenizer": { + "gguf": "lfm2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Lfm2BidirectionalModel": { + "gguf": "lfm2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Lfm2DSparkDraftModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Lfm2ForCausalLM": { + "gguf": "lfm2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Lfm2Model": { + "gguf": "lfm2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Lfm2MoeForCausalLM": { + "gguf": "lfm2moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LingDSparkModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Llama4ForCausalLM": { + "gguf": "llama4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Llama4ForConditionalGeneration": { + "gguf": "llama4", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LlamaBidirectionalModel": { + "gguf": "llama-embed", + "backends": [ + "hrx", + "vulkan" + ] + }, + "LlamaForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "LlamaForCausalLMEagle3": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "LlamaModel": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "LlavaForConditionalGeneration": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "LlavaStableLMEpochForCausalLM": { + "gguf": "stablelm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MPTForCausalLM": { + "gguf": "mpt", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MT5ForConditionalGeneration": { + "gguf": "t5", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MaincoderForCausalLM": { + "gguf": "maincoder", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Mamba2ForCausalLM": { + "gguf": "mamba2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MambaForCausalLM": { + "gguf": "mamba", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "MambaLMHeadModel": { + "gguf": "mamba", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "MapleForCausalLM": { + "gguf": "maple", + "backends": [ + "vulkan" + ] + }, + "MellumForCausalLM": { + "gguf": "mellum", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiMoV2FlashForCausalLM": { + "gguf": "mimo2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiMoV2ForCausalLM": { + "gguf": "mimo2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiniCPM3ForCausalLM": { + "gguf": "minicpm3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiniCPMForCausalLM": { + "gguf": "minicpm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiniCPMV4_6ForConditionalGeneration": { + "gguf": "qwen35", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "MiniMaxM1ForCausalLM": { + "gguf": "minimax-01", + "backends": [ + "vulkan" + ] + }, + "MiniMaxM2ForCausalLM": { + "gguf": "minimax-m2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiniMaxM3SparseForCausalLM": { + "gguf": "minimax-m3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiniMaxM3SparseForConditionalGeneration": { + "gguf": "minimax-m3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MiniMaxText01ForCausalLM": { + "gguf": "minimax-01", + "backends": [ + "vulkan" + ] + }, + "Ministral3ForCausalLM": { + "gguf": "mistral3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Mistral3ForConditionalGeneration": { + "gguf": "mistral3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MistralForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "MixtralForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "ModernBertForMaskedLM": { + "gguf": "modern-bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ModernBertForSequenceClassification": { + "gguf": "modern-bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "ModernBertModel": { + "gguf": "modern-bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MuseGlimmerAssistantModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "MuseGlimmerForConditionalGeneration": { + "gguf": "muse-glimmer", + "backends": [ + "vulkan", + "zinc" + ] + }, + "NanbeigeForCausalLM": { + "gguf": "nanbeige", + "backends": [ + "hrx", + "vulkan" + ] + }, + "NemotronForCausalLM": { + "gguf": "nemotron", + "backends": [ + "hrx", + "vulkan" + ] + }, + "NemotronHForCausalLM": { + "gguf": "nemotron_h", + "backends": [ + "hrx", + "vulkan" + ] + }, + "NemotronHPuzzleForCausalLM": { + "gguf": "nemotron_h_moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "NeoBERT": { + "gguf": "neo-bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "NeoBERTForSequenceClassification": { + "gguf": "neo-bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "NeoBERTLMHead": { + "gguf": "neo-bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "NomicBertModel": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "OLMoForCausalLM": { + "gguf": "olmo", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Olmo2ForCausalLM": { + "gguf": "olmo2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Olmo3ForCausalLM": { + "gguf": "olmo2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "OlmoForCausalLM": { + "gguf": "olmo", + "backends": [ + "hrx", + "vulkan" + ] + }, + "OlmoeForCausalLM": { + "gguf": "olmoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "OpenELMForCausalLM": { + "gguf": "openelm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "OrionForCausalLM": { + "gguf": "orion", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PLMForCausalLM": { + "gguf": "plm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PLaMo2ForCausalLM": { + "gguf": "plamo2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PLaMo3ForCausalLM": { + "gguf": "plamo3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PaddleOCRVLForConditionalGeneration": { + "gguf": "paddleocr", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PanguEmbeddedForCausalLM": { + "gguf": "pangu-embedded", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Phi3ForCausalLM": { + "gguf": "phi3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Phi4ForCausalLMV": { + "gguf": "phi3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PhiForCausalLM": { + "gguf": "phi2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PhiMoEForCausalLM": { + "gguf": "phimoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Plamo2ForCausalLM": { + "gguf": "plamo2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Plamo3ForCausalLM": { + "gguf": "plamo3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PlamoForCausalLM": { + "gguf": "plamo", + "backends": [ + "hrx", + "vulkan" + ] + }, + "PocketTTSModel": { + "gguf": "pockettts", + "backends": [ + "vulkan" + ] + }, + "QWenLMHeadModel": { + "gguf": "qwen", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen2AudioForConditionalGeneration": { + "gguf": "qwen2", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen2ForCausalLM": { + "gguf": "qwen2", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen2Model": { + "gguf": "qwen2", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen2MoeForCausalLM": { + "gguf": "qwen2moe", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen2VLForConditionalGeneration": { + "gguf": "qwen2vl", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen2VLModel": { + "gguf": "qwen2vl", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen2_5OmniModel": { + "gguf": "qwen2vl", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen2_5_VLForConditionalGeneration": { + "gguf": "qwen2vl", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen3ASRForConditionalGeneration": { + "gguf": "qwen3vl", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen3DSparkModel": { + "gguf": "dflash", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen3ForCausalLM": { + "gguf": "qwen3", + "backends": [ + "hrx", + "npu", + "vulkan", + "zinc" + ], + "npu_model_type": "qwen3" + }, + "Qwen3Model": { + "gguf": "qwen3", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen3MoeForCausalLM": { + "gguf": "qwen3moe", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen3NextForCausalLM": { + "gguf": "qwen3next", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen3OmniMoeForConditionalGeneration": { + "gguf": "qwen3vlmoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen3TTSForConditionalGeneration": { + "gguf": "qwen3tts", + "backends": [ + "vulkan" + ] + }, + "Qwen3VLForConditionalGeneration": { + "gguf": "qwen3vl", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen3VLMoeForConditionalGeneration": { + "gguf": "qwen3vlmoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Qwen3_5ForCausalLM": { + "gguf": "qwen35", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen3_5ForConditionalGeneration": { + "gguf": "qwen35", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen3_5MoeForCausalLM": { + "gguf": "qwen35moe", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen3_5MoeForConditionalGeneration": { + "gguf": "qwen35moe", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "Qwen4ExpForCausalLM": { + "gguf": "qwen4exp", + "backends": [ + "vulkan" + ] + }, + "Qwen4ExpForConditionalGeneration": { + "gguf": "qwen4exp", + "backends": [ + "vulkan" + ] + }, + "RND1": { + "gguf": "rnd1", + "backends": [ + "hrx", + "vulkan" + ] + }, + "RWForCausalLM": { + "gguf": "falcon", + "backends": [ + "hrx", + "vulkan" + ] + }, + "RWKV6Qwen2ForCausalLM": { + "gguf": "rwkv6qwen2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "RWKV7ForCausalLM": { + "gguf": "rwkv7", + "backends": [ + "hrx", + "vulkan" + ] + }, + "RobertaForSequenceClassification": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "RobertaModel": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "RuGPT3XLForCausalLM": { + "gguf": "gpt2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Rwkv6ForCausalLM": { + "gguf": "rwkv6", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Rwkv7ForCausalLM": { + "gguf": "rwkv7", + "backends": [ + "hrx", + "vulkan" + ] + }, + "RwkvHybridForCausalLM": { + "gguf": "arwkv7", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Sarashina2VisionForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "SarvamMoEForCausalLM": { + "gguf": "bailingmoe2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "SeedOssForCausalLM": { + "gguf": "seed_oss", + "backends": [ + "hrx", + "vulkan" + ] + }, + "SmallThinkerForCausalLM": { + "gguf": "smallthinker", + "backends": [ + "hrx", + "vulkan" + ] + }, + "SmolLM3ForCausalLM": { + "gguf": "smollm3", + "backends": [ + "hrx", + "vulkan" + ] + }, + "SolarOpenForCausalLM": { + "gguf": "glm4moe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Spark2_5ForCausalLM": { + "gguf": "spark2_5", + "backends": [ + "vulkan" + ] + }, + "StableLMEpochForCausalLM": { + "gguf": "stablelm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "StableLmForCausalLM": { + "gguf": "stablelm", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Starcoder2ForCausalLM": { + "gguf": "starcoder2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Step3p5ForCausalLM": { + "gguf": "step35", + "backends": [ + "hrx", + "vulkan" + ] + }, + "Step3p7ForConditionalGeneration": { + "gguf": "step35", + "backends": [ + "hrx", + "vulkan" + ] + }, + "StepVLForConditionalGeneration": { + "gguf": "qwen3", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "T5EncoderModel": { + "gguf": "t5encoder", + "backends": [ + "hrx", + "vulkan" + ] + }, + "T5ForConditionalGeneration": { + "gguf": "t5", + "backends": [ + "hrx", + "vulkan" + ] + }, + "T5WithLMHeadModel": { + "gguf": "t5", + "backends": [ + "hrx", + "vulkan" + ] + }, + "TalkieForCausalLM": { + "gguf": "talkie", + "backends": [ + "hrx", + "vulkan" + ] + }, + "UMT5ForConditionalGeneration": { + "gguf": "t5", + "backends": [ + "hrx", + "vulkan" + ] + }, + "UMT5Model": { + "gguf": "t5", + "backends": [ + "hrx", + "vulkan" + ] + }, + "UltravoxModel": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "UnlimitedOCRForCausalLM": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "VLlama3ForCausalLM": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "VoxtralForConditionalGeneration": { + "gguf": "llama", + "backends": [ + "hrx", + "vulkan", + "zinc" + ] + }, + "WavTokenizerDec": { + "gguf": "wavtokenizer-dec", + "backends": [ + "hrx", + "vulkan" + ] + }, + "XLMRobertaForSequenceClassification": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "XLMRobertaModel": { + "gguf": "bert", + "backends": [ + "hrx", + "vulkan" + ] + }, + "XverseForCausalLM": { + "gguf": "xverse", + "backends": [ + "hrx", + "vulkan" + ] + }, + "YoutuForCausalLM": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "YoutuVLForConditionalGeneration": { + "gguf": "deepseek2", + "backends": [ + "hrx", + "vulkan" + ] + }, + "modeling_grove_moe.GroveMoeForCausalLM": { + "gguf": "grovemoe", + "backends": [ + "hrx", + "vulkan" + ] + }, + "modeling_sarvam_moe.SarvamMoEForCausalLM": { + "gguf": "bailingmoe2", + "backends": [ + "hrx", + "vulkan" + ] + } + } +} diff --git a/registry/census.json b/registry/census.json new file mode 100644 index 00000000..24e7ecdb --- /dev/null +++ b/registry/census.json @@ -0,0 +1,2758 @@ +{ + "date": "2026-09-25", + "coverage": { + "models": 415414, + "with_architecture": 332565, + "architectures_seen": 2610, + "mapped_models": 310221, + "mapped_pct": 93.28, + "checked_models": 213456, + "checked_pct": 64.18, + "per_backend": { + "vulkan": { + "mapped": 310221, + "checked": 213456, + "mapped_pct": 93.28, + "checked_pct": 64.18 + }, + "hrx": { + "mapped": 309917, + "checked": 210820, + "mapped_pct": 93.19, + "checked_pct": 63.39 + }, + "zinc": { + "mapped": 231768, + "checked": 34192, + "mapped_pct": 69.69, + "checked_pct": 10.28 + }, + "npu": { + "mapped": 30879, + "checked": 30879, + "mapped_pct": 9.29, + "checked_pct": 9.29 + } + }, + "top_unmapped": [ + { + "architecture": "Step1MoEForCausalLM", + "models": 2882 + }, + { + "architecture": "OPTForCausalLM", + "models": 2097 + }, + { + "architecture": "ParlerTTSForConditionalGeneration", + "models": 1586 + }, + { + "architecture": "GPTNeoForCausalLM", + "models": 1565 + }, + { + "architecture": "ChessForCausalLM", + "models": 970 + }, + { + "architecture": "GPTJForCausalLM", + "models": 614 + }, + { + "architecture": "LlavaLlamaForCausalLM", + "models": 451 + }, + { + "architecture": "CodeGenForCausalLM", + "models": 386 + }, + { + "architecture": "PicoDecoderHF", + "models": 310 + }, + { + "architecture": "DynamicAlibiForCausalLM", + "models": 245 + }, + { + "architecture": "OpenLMForCausalLM", + "models": 245 + }, + { + "architecture": "DynamicForgettingForCausalLM", + "models": 240 + }, + { + "architecture": "BertLMHeadModel", + "models": 212 + }, + { + "architecture": "MobileLLMForCausalLM", + "models": 176 + }, + { + "architecture": "BartForConditionalGeneration", + "models": 174 + }, + { + "architecture": "MixFormerSequentialForCausalLM", + "models": 156 + }, + { + "architecture": "RobertaForCausalLM", + "models": 139 + }, + { + "architecture": "NanochatGPTForCausalLM", + "models": 121 + }, + { + "architecture": "CambrianQwenForCausalLM", + "models": 113 + }, + { + "architecture": "MBartForConditionalGeneration", + "models": 110 + }, + { + "architecture": "BartForCausalLM", + "models": 100 + }, + { + "architecture": "OpensciForCausalLM", + "models": 98 + }, + { + "architecture": "LlavaQwen2ForCausalLM", + "models": 97 + }, + { + "architecture": "BioGptForCausalLM", + "models": 86 + }, + { + "architecture": "RwkvForCausalLM", + "models": 86 + } + ] + }, + "raw": { + "total": 415414, + "no_arch": 82849, + "pages": 416, + "complete": true, + "counts": { + "LlamaForCausalLM": 100458, + "Qwen2ForCausalLM": 46457, + "GPT2LMHeadModel": 31744, + "Qwen3ForCausalLM": 30879, + "MistralForCausalLM": 27912, + "GPTNeoXForCausalLM": 8049, + "StableLmForCausalLM": 6679, + "Gemma3ForCausalLM": 6074, + "Gemma2ForCausalLM": 5579, + "GemmaForCausalLM": 5518, + "Phi3ForCausalLM": 3217, + "Qwen3_5ForConditionalGeneration": 3046, + "Step1MoEForCausalLM": 2882, + "MixtralForCausalLM": 2829, + "OPTForCausalLM": 2097, + "GptOssForCausalLM": 1920, + "Qwen3MoeForCausalLM": 1715, + "PhiForCausalLM": 1649, + "ParlerTTSForConditionalGeneration": 1586, + "GPTNeoForCausalLM": 1565, + "Qwen3_5MoeForConditionalGeneration": 1388, + "Gemma4ForConditionalGeneration": 1260, + "Qwen3_5ForCausalLM": 1215, + "BloomForCausalLM": 1154, + "Lfm2ForCausalLM": 1133, + "Gemma3ForConditionalGeneration": 1093, + "Olmo3ForCausalLM": 1076, + "T5ForConditionalGeneration": 1032, + "ChessForCausalLM": 970, + "FalconForCausalLM": 919, + "GlmMoeDsaForCausalLM": 616, + "GPTJForCausalLM": 614, + "OlmoForCausalLM": 601, + "GraniteForCausalLM": 589, + "DeepseekV4ForCausalLM": 579, + "Olmo2ForCausalLM": 558, + "CohereForCausalLM": 535, + "GraniteMoeHybridForCausalLM": 532, + "DeepseekV3ForCausalLM": 521, + "NemotronHForCausalLM": 505, + "LlavaLlamaForCausalLM": 451, + "SmolLM3ForCausalLM": 439, + "GPTBigCodeForCausalLM": 433, + "CodeGenForCausalLM": 386, + "MT5ForConditionalGeneration": 363, + "Qwen3NextForCausalLM": 356, + "QWenLMHeadModel": 347, + "Glm4MoeForCausalLM": 341, + "MiniMaxM2ForCausalLM": 322, + "Starcoder2ForCausalLM": 313, + "PicoDecoderHF": 310, + "RWForCausalLM": 300, + "MPTForCausalLM": 279, + "Lfm2MoeForCausalLM": 249, + "DynamicAlibiForCausalLM": 245, + "OpenLMForCausalLM": 245, + "DynamicForgettingForCausalLM": 240, + "Gemma4UnifiedForConditionalGeneration": 235, + "SeedOssForCausalLM": 226, + "Glm4ForCausalLM": 224, + "BertLMHeadModel": 212, + "Mistral3ForConditionalGeneration": 194, + "Qwen3_5MoeForCausalLM": 194, + "OlmoeForCausalLM": 190, + "ExaoneForCausalLM": 189, + "Cohere2ForCausalLM": 183, + "DeepseekV2ForCausalLM": 183, + "MobileLLMForCausalLM": 176, + "BartForConditionalGeneration": 174, + "LagunaForCausalLM": 173, + "StableLMEpochForCausalLM": 171, + "JambaForCausalLM": 169, + "Glm4MoeLiteForCausalLM": 166, + "OpenELMForCausalLM": 166, + "FalconH1ForCausalLM": 163, + "MambaForCausalLM": 158, + "MixFormerSequentialForCausalLM": 156, + "LLaMAForCausalLM": 153, + "RobertaForCausalLM": 139, + "ApertusForCausalLM": 135, + "BaichuanForCausalLM": 131, + "DeciLMForCausalLM": 129, + "DFlashDraftModel": 127, + "LlamaForCausalLMEagle3": 126, + "NanochatGPTForCausalLM": 121, + "InternLM2ForCausalLM": 120, + "Qwen2MoeForCausalLM": 116, + "CambrianQwenForCausalLM": 113, + "MBartForConditionalGeneration": 110, + "MiniCPMForCausalLM": 109, + "HunYuanDenseV1ForCausalLM": 108, + "BartForCausalLM": 100, + "OpensciForCausalLM": 98, + "LlavaQwen2ForCausalLM": 97, + "DeepseekV32ForCausalLM": 89, + "JAISLMHeadModel": 89, + "HYV3ForCausalLM": 88, + "BioGptForCausalLM": 86, + "RwkvForCausalLM": 86, + "GPT2LMHeadCustomModel": 84, + "Gemma4ForCausalLM": 84, + "AfmoeForCausalLM": 83, + "Llama4ForCausalLM": 81, + "DynamicSlidingWindowForCausalLM": 79, + "XGLMForCausalLM": 79, + "Glm5NextForConditionalGeneration": 77, + "K2HorizonForCausalLM": 77, + "MarianForCausalLM": 76, + "MiMoV2ForCausalLM": 75, + "LISAForCausalLM": 74, + "NMMaskMoELLaVAQwen3ForCausalLM": 73, + "KORMoForCausalLM": 72, + "NanoChatForCausalLM": 71, + "TransformerForCausalLM": 68, + "Ernie4_5_MoeForCausalLM": 66, + "Qwen4ExpForConditionalGeneration": 66, + "Exaone4ForCausalLM": 65, + "NemotronForCausalLM": 64, + "LlamaModel": 63, + "T5WithLMHeadModel": 59, + "YiForCausalLM": 57, + "CacaForCausalLM": 55, + "Moondream": 55, + "NanbeigeForCausalLM": 55, + "SparseMistralforCausalLM": 55, + "BitLlamaForCausalLM": 54, + "BitNetForCausalLM": 54, + "Gemma4AssistantForCausalLM": 54, + "LlavaMistralForCausalLM": 53, + "MinistralForCausalLM": 52, + "Phi3VForCausalLM": 52, + "BailingMoeV3ForCausalLM": 51, + "DenseControlLM": 50, + "IQuestCoderForCausalLM": 50, + "Phi4MMForCausalLM": 50, + "RWKV7ForCausalLM": 50, + "SarvamMoEForCausalLM": 50, + "Spark2_5ForCausalLM": 49, + "LLaDAModelLM": 48, + "DbrxForCausalLM": 44, + "GPT": 44, + "GraniteMoeForCausalLM": 44, + "PITForCausalLM": 43, + "RoleSLM": 42, + "SlidingWindowForCausalLM": 42, + "CustomLlamaForCausalLM": 41, + "HrmTextForCausalLM": 41, + "NotaGenLMHeadModel": 41, + "DFlash2DraftModel": 40, + "GPJTGPT2ModelForCausalLM": 40, + "LLaDA2MoeModelLM": 40, + "MPLUGOwl2LlamaForCausalLM": 40, + "MellumForCausalLM": 40, + "MptForCausalLM": 40, + "DreamModel": 39, + "HelixForCausalLM": 39, + "Qwen3VLForConditionalGeneration": 39, + "SDARForCausalLM": 39, + "HyperCLOVAXForCausalLM": 38, + "OLMoForCausalLM": 38, + "Qwen2_5_VLForConditionalGeneration": 38, + "RavenForCausalLM": 37, + "XLNetLMHeadModel": 37, + "BailingMoeV2ForCausalLM": 36, + "T5GemmaForConditionalGeneration": 36, + "InternLM3ForCausalLM": 35, + "Ministral3ForCausalLM": 35, + "OrionForCausalLM": 35, + "Qwen2Model": 35, + "TalkieForCausalLM": 35, + "ChatGLMForConditionalGeneration": 34, + "LoopLMForCausalLM": 34, + "KimiK25ForConditionalGeneration": 33, + "KimiLinearForCausalLM": 33, + "MorphT5AutoForConditionalGeneration": 32, + "MorphT5ConcatForConditionalGeneration": 32, + "MorphT5SumForConditionalGeneration": 32, + "TransformerLM": 32, + "LlavaQwenForCausalLM": 31, + "LongcatCausalLM": 31, + "OuroForCausalLM": 31, + "ZhinaoForCausalLM": 31, + "Gemma3nForConditionalGeneration": 30, + "XLMRobertaForCausalLM": 30, + "DogeForCausalLM": 28, + "HyperLlamaForCausalLM": 28, + "NanoGPTForCausalLM": 28, + "Step3p5ForCausalLM": 28, + "Cohere2MoeForCausalLM": 27, + "DSparkDraftModel": 27, + "DaisyForCausalLM": 27, + "SkyworkForCausalLM": 27, + "AquilaForCausalLM": 26, + "BailingMoeV2_5ForCausalLM": 26, + "DeepseekV41ForCausalLM": 26, + "GatedDeltaNetForCausalLM": 26, + "GlmForCausalLM": 26, + "Llama4ForConditionalGeneration": 26, + "MBartForCausalLM": 26, + "MultiScaleForCausalLM": 26, + "OpenAIGPTLMHeadModel": 26, + "GPTNeoXModel": 25, + "LlavaLlamaModel": 25, + "Phi3SmallForCausalLM": 25, + "PhiMoEForCausalLM": 25, + "SparseLlamaForCausalLM": 25, + "CTRLLMHeadModel": 24, + "CogVLMForCausalLM": 24, + "ConstrainedLlamaForCausalLM": 24, + "DetikzifyForConditionalGeneration": 24, + "FP8Qwen3ForCausalLM": 24, + "MuseGlimmerForConditionalGeneration": 24, + "XverseForCausalLM": 24, + "BaiChuanForCausalLM": 23, + "BloomModel": 23, + "CENOForCausalLM": 23, + "ForgettingTransformerForCausalLM": 23, + "GPT2MoEForCausalLM": 23, + "GrugMoeForCausalLM": 23, + "M2M100ForConditionalGeneration": 23, + "GPT2LMAndValueHeadModel": 22, + "HunYuanMoEV1ForCausalLM": 22, + "MiMoV2FlashForCausalLM": 22, + "DeepseekForCausalLM": 21, + "DistilBertForSequenceClassification": 21, + "FlexOlmoForCausalLM": 21, + "HymbaForCausalLM": 21, + "LlavaPhiForCausalLM": 21, + "ProGenForCausalLM": 21, + "SkipMiddleModel": 21, + "ZayaForCausalLM": 21, + "AutoModelForCausalLM": 20, + "GPT2ALMHeadModel": 20, + "LlavaForConditionalGeneration": 20, + "MotifForCausalLM": 20, + "PoptorchPipelinedGPT2LMHeadModel": 20, + "Qwen3VLMoeForConditionalGeneration": 20, + "BertForMaskedLM": 19, + "GPT2Model": 19, + "MllamaForCausalLM": 19, + "QuasarForCausalLM": 19, + "SolarOpen2ForCausalLM": 19, + "BVVForCausalLM": 18, + "BailingMoeLinearV2ForCausalLM": 18, + "FalconMambaForCausalLM": 18, + "MFuyuForCausalLM": 18, + "Mamba2ForCausalLM": 18, + "RECAST8b_LlamaForCausalLM": 18, + "Rwkv7ForCausalLM": 18, + "StickbreakingForCausalLM": 18, + "AlibiForCausalLM": 17, + "CubeLM": 17, + "GPTForCausalLM": 17, + "HGRNForCausalLM": 17, + "InternLMForCausalLM": 17, + "KimiK3ForConditionalGeneration": 17, + "MiniMaxM3SparseForConditionalGeneration": 17, + "MobilintLlamaForCausalLM": 17, + "NCPOlmo3ForCausalLM": 17, + "PldrllmForCausalLM": 17, + "RecurrentGemmaForCausalLM": 17, + "RetNetForCausalLM": 17, + "SarvamMLAForCausalLM": 17, + "ChatGLMModel": 16, + "KPhi3ForCausalLM": 16, + "LongcatFlashNgramForCausalLM": 16, + "NARBartForConditionalGeneration": 16, + "NandiForCausalLM": 16, + "NemotronLabsDiffusionModel": 16, + "QuietForCausalLM": 16, + "Qwen3DSparkModel": 16, + "SolarOpenForCausalLM": 16, + "YoutuForCausalLM": 16, + "transformerModel": 16, + "ArceeForCausalLM": 15, + "DiffusionGemmaForBlockDiffusion": 15, + "ExaoneMoEForCausalLM": 15, + "NanochronoForCausalLM": 15, + "Qwen4ExpForCausalLM": 15, + "WhisperForCausalLM": 15, + "GPTJiangForCausalLM": 14, + "HunYuanForCausalLM": 14, + "JetMoEForCausalLM": 14, + "PawQwen3ForCausalLM": 14, + "ProgressiveYocoLlamaForCausalLM": 14, + "Qwen2ReasoningForCausalLM": 14, + "Qwen2VLForConditionalGeneration": 14, + "Qwen3Model": 14, + "ReVisionForConditionalGeneration": 14, + "STLForCausalLM": 14, + "SpikeWhaleLM": 14, + "TinyLlavaForConditionalGeneration": 14, + "Xing4_0ForCausalLM": 14, + "XpertGPTForCausalLM": 14, + "AdapterMoELLaVAQwen3ForCausalLM": 13, + "BailingMoeForCausalLM": 13, + "BitnetForCausalLM": 13, + "DiscreteDiffusionModel": 13, + "Eagle3DeepseekV2ForCausalLM": 13, + "EncoderDecoderModel": 13, + "Ernie4_5_ForCausalLM": 13, + "InstellaMoEForCausalLM": 13, + "MobilintQwen3ForCausalLM": 13, + "OpenMoeForCausalLM": 13, + "Qwen3RecoveredForCausalLM": 13, + "RECAST1B_LlamaForCausalLM": 13, + "Rwkv6ForCausalLM": 13, + "SelfDebiasingGPT2LMHeadModel": 13, + "YatGPTForCausalLM": 13, + "AV2TextForConditionalGeneration": 12, + "ArgonneModel": 12, + "BartModel": 12, + "BertModel": 12, + "BlockMTPForCausalLM": 12, + "Ernie4_5ForCausalLM": 12, + "FimmyForCausalLM": 12, + "GPTJXForCausalLM": 12, + "GPTJXMoEForCausalLM": 12, + "IndustrySLM": 12, + "LLamaLongBEL": 12, + "MapleForCausalLM": 12, + "MistralModel": 12, + "MobiLlamaForCausalLM": 12, + "ModernBertDecoderForCausalLM": 12, + "OrkhonForCausalLM": 12, + "PlamoForCausalLM": 12, + "RoFormerForCausalLM": 12, + "StripedHyenaModelForCausalLM": 12, + "T5LaForConditionalGeneration": 12, + "TTTForCausalLM": 12, + "A2DQwen3LMHeadModel": 11, + "CamembertForCausalLM": 11, + "Eagle3DraftModel": 11, + "ElectraForCausalLM": 11, + "GLAForCausalLM": 11, + "HGRNBitForCausalLM": 11, + "IndexForCausalLM": 11, + "JarvisTitanMoEForCausalLM": 11, + "KeuralMoECausalLM": 11, + "MiMoForCausalLM": 11, + "MiniCPM3ForCausalLM": 11, + "MoLM": 11, + "MobilintExaoneForCausalLM": 11, + "MobilintQwen2ForCausalLM": 11, + "NeedleForToolCalling": 11, + "Qwen2BMForCausalLM": 11, + "Qwen2ParScaleForCausalLM": 11, + "TinyGPTForCausalLM": 11, + "VideoChatGPTLlamaForCausalLM": 11, + "YuanForCausalLM": 11, + "Zamba2ForCausalLM": 11, + "AnemoneForCausalLM": 10, + "Avey": 10, + "BabyLlamaForCausalLM": 10, + "BambaForCausalLM": 10, + "EmoForCausalLM": 10, + "FinanceDecoder": 10, + "GPTBERTForCausalLM": 10, + "Gemma4UnifiedAssistantForCausalLM": 10, + "GraniteSwitchForCausalLM": 10, + "IQuestLoopCoderForCausalLM": 10, + "InternVLChatModel": 10, + "JapaneseStableLMAlphaForCausalLM": 10, + "LlamaMoEForCausalLM": 10, + "LlavaNextForConditionalGeneration": 10, + "MarianMTModel": 10, + "MegaForCausalLM": 10, + "MiMoGDNForCausalLM": 10, + "ModernBertForMaskedLM": 10, + "MoshiForConditionalGeneration": 10, + "MossForCausalLM": 10, + "NanochatForCausalLM": 10, + "Olmo3SiameseDepthForCausalLM": 10, + "Ovis": 10, + "Prot2TextModel": 10, + "Qwen2_5OmniForConditionalGeneration": 10, + "Qwen3OmniMoeForConditionalGeneration": 10, + "ReformerModelWithLMHead": 10, + "RingAttentionGPT2LMHeadModel": 10, + "modeling_sparsetral.MistralForCausalLM": 10, + "ACIPModel": 9, + "BananaMind2NanoForCausalLM": 9, + "BananaMind2PicoForCausalLM": 9, + "BlenderbotForCausalLM": 9, + "BlenderbotForConditionalGeneration": 9, + "BunnyPhiForCausalLM": 9, + "ChessGPT": 9, + "CognicaPoEForCausalLM": 9, + "Evo2ForCausalLM": 9, + "GPTPanguForCausalLM": 9, + "HelpingAIForCausalLM": 9, + "HybridQwen3ForCausalLM": 9, + "InstellaForCausalLM": 9, + "Lfm2MoEForCausalLM": 9, + "LightOnOCRForConditionalGeneration": 9, + "LimiteForCausalLM": 9, + "LlavaLlamaAttForCausalLM": 9, + "LongcatFlashForCausalLM": 9, + "MiniMindLM": 9, + "MoELLaVAQwen3ForCausalLM": 9, + "MobileMoEForCausalLM": 9, + "ParamBharatGenForCausalLM": 9, + "Qwen2VLAudioForConditionalGeneration": 9, + "RECAST7b_LlamaForCausalLM": 9, + "RWKV": 9, + "RobertaForMaskedLM": 9, + "Rwkv5ForCausalLM": 9, + "TFGPT2LMHeadModel": 9, + "TPPForCausalLM": 9, + "TrainableM2MForConditionalGeneration": 9, + "WeDLMForCausalLM": 9, + "AliceAIForCausalLM": 8, + "BTLMLMHeadModel": 8, + "BananaMind2ProForCausalLM": 8, + "BartEncodecForConditionalGeneration": 8, + "ColMaskMoELLaVAQwen3ForCausalLM": 8, + "DUO": 8, + "DeepForCausalLM": 8, + "Dots1ForCausalLM": 8, + "DuchifatCore": 8, + "FP8Qwen2ForCausalLM": 8, + "GPTOptim": 8, + "GravityMoEForCausalLM": 8, + "HYV4ForCausalLM": 8, + "HeliumForCausalLM": 8, + "HyenaDNAForCausalLM": 8, + "LlavaMPTForCausalLM": 8, + "LlavaQwen3ForCausalLM": 8, + "MiniGeminiLlamaForCausalLM": 8, + "MllamaForConditionalGeneration": 8, + "MoSMambaForCausalLM": 8, + "MolmoForCausalLM": 8, + "MonetForCausalLM": 8, + "MosaicGPT": 8, + "MoveDecoder": 8, + "MyLlamaForCausalLM": 8, + "NemotronH_Nano_Omni_Reasoning_V3": 8, + "OlmoModelForCausalLM": 8, + "Plamo3ForCausalLM": 8, + "PolyverseForConditionalGeneration": 8, + "Qwen3Mamba3ForCausalLM": 8, + "TelechatForCausalLM": 8, + "TinyGDNForCausalLM": 8, + "TransfoXLLMHeadModel": 8, + "Transformer": 8, + "TransnormerForCausalLM": 8, + "TwinyForCausalLM": 8, + "ZgcmForCausalLM": 8, + "_A2DQwen3LMHeadModel": 8, + "BD3LM": 7, + "BigBirdForCausalLM": 7, + "BlueLMForCausalLM": 7, + "BrujulaForCausalLM": 7, + "BunnyLlamaForCausalLM": 7, + "BunnyPhi3ForCausalLM": 7, + "CanopyForCausalLM": 7, + "DynColMaskMoELLaVAQwen2ForCausalLM": 7, + "EduLLMForCausalLM": 7, + "ExaoneMoeForCausalLM": 7, + "GPT3DevLMHeadModel": 7, + "GigaChat35ForCausalLM": 7, + "Int8OPTForCausalLM": 7, + "K3DSparkModel": 7, + "LLaMAModel": 7, + "LizzyForCausalLM": 7, + "LlavaGemmaForCausalLM": 7, + "LongLlamaForCausalLM": 7, + "LongformerForSequenceClassification": 7, + "MiniMaxM1ForCausalLM": 7, + "MiniMindForCausalLM": 7, + "NMMaskMoELLaVAPhiForCausalLM": 7, + "NemotronHPuzzleForCausalLM": 7, + "NeuronSparkForCausalLM": 7, + "OpenSciForCausalLM": 7, + "PegasusForConditionalGeneration": 7, + "Phi4FlashForCausalLM": 7, + "PhimoeForCausalLM": 7, + "Plamo2ForCausalLM": 7, + "ProGen2ForPreTraining": 7, + "Qwen3AudioWrappedForCausalLM": 7, + "Qwen3TDMoEForCausalLM": 7, + "SDARMoeForCausalLM": 7, + "SpeckForCausalLM": 7, + "T5Model": 7, + "ASVDLlamaForCausalLM": 6, + "AprielForCausalLM": 6, + "BaichuanM1ForCausalLM": 6, + "BlenderbotSmallForCausalLM": 6, + "BolmoForCausalLM": 6, + "CENOPForCausalLM": 6, + "CLIPT5ForConditionalGeneration": 6, + "CogAgentForCausalLM": 6, + "CustomBartForConditionalGeneration": 6, + "DeepQwenVLForCausalLM": 6, + "DetikzifyCambrianForConditionalGeneration": 6, + "Eagle3Speculator": 6, + "EchoForCausalLM": 6, + "FHN_T4Max_150M": 6, + "FunctionaryForCausalLM": 6, + "GDN2ForCausalLM": 6, + "GENERatorForCausalLM": 6, + "GPT2LLMHeadModel": 6, + "GPTNeoForCausalLMTiered": 6, + "GSAForCausalLM": 6, + "GemmoeForCausalLM": 6, + "HybridGPT2LMHeadModel": 6, + "Idefics3ForConditionalGeneration": 6, + "ImpForCausalLM": 6, + "InklingForConditionalGeneration": 6, + "InternLMXComposer2ForCausalLM": 6, + "JugnuVRForCausalLM": 6, + "LSGBartForConditionalGeneration": 6, + "LilleForCausalLM": 6, + "LlavaMambaForCausalLM": 6, + "LoRDCoderForCausalLM": 6, + "MaskMoELLaVAQwen3ForCausalLM": 6, + "MediKoForCausalLM": 6, + "MiniCPMSALAForCausalLM": 6, + "MixtralMoleForCausalLM": 6, + "MoAMetricLM": 6, + "MobileLLMP1ForCausalLM": 6, + "MyT5ForConditionalGeneration": 6, + "NorovoxAlphaMoE": 6, + "NovaForCausalLM": 6, + "OLMo3ForCausalLM": 6, + "PanguEmbeddedForCausalLM": 6, + "PegasusForCausalLM": 6, + "PointLLMLlamaForCausalLM": 6, + "Qwen2ChunkingForCausalLM": 6, + "Qwen2ForCausalLMPostBlockSteeringFixed": 6, + "Qwen2ForProcessRewardModel": 6, + "QwenForCausalLM": 6, + "RandyGPTForCausalLM": 6, + "ScratchLlamaForCausalLM": 6, + "SpatialLMQwenForCausalLM": 6, + "StarVectorForCausalLM": 6, + "T2MLRWrapper": 6, + "T5EncoderModel": 6, + "TPUGemma3ForCausalLM": 6, + "TrillionForCausalLM": 6, + "UMT5ForConditionalGeneration": 6, + "UrchinParallelForCausalLM": 6, + "UserLlamaForCausalLM": 6, + "ZambaForCausalLM": 6, + "AILOForCausalLM": 5, + "AnuLMForCausalLM": 5, + "BottleneckT5LMWithPerturb": 5, + "CWICForCausalLM": 5, + "CogVLMVideoForCausalLM": 5, + "DashQPhi3ForCausalLM": 5, + "DashQQwen2ForCausalLM": 5, + "DeciCoderForCausalLM": 5, + "DetikzifyForCausalLM": 5, + "EmbformerForCausalLM": 5, + "ErnieForCausalLM": 5, + "FWKVLanguageModel": 5, + "FineWebForCausalLM": 5, + "FlamingoForConditionalGeneration": 5, + "GearForCausalLM": 5, + "HGRN2ForCausalLM": 5, + "HawkForCausalLM": 5, + "HelloWorldModel": 5, + "INFLMForCausalLM": 5, + "IdeficsForVisionText2Text": 5, + "JetMoeForCausalLM": 5, + "KANGPT2LMHeadModel": 5, + "KirimForCausalLM": 5, + "Lfm2VlForConditionalGeneration": 5, + "LlamaGlideDecoderLayer": 5, + "LlavaGpt2ForCausalLM": 5, + "M2RForCausalLM": 5, + "MaincoderForCausalLM": 5, + "MambaInQwenForCausalLM": 5, + "MobileLlamaForCausalLM": 5, + "NGMEForCausalLM": 5, + "NanoTransformer": 5, + "OlmoHybridForCausalLM": 5, + "PhariaForCausalLM": 5, + "Phi3Model": 5, + "PhoneLMForCausalLM": 5, + "PinyinCodeForCausalLM": 5, + "QuasarLongForCausalLM": 5, + "Qwen2TSForCausalLM": 5, + "Qwen2_5OmniModel": 5, + "Qwen3AttnResForCausalLM": 5, + "RWKV6Qwen2ForCausalLM": 5, + "ReplitLM": 5, + "RosettaForCausalLM": 5, + "RwkvHybridForCausalLM": 5, + "Share4VLlamaForCausalLM": 5, + "SiQ_VLForCausalLM": 5, + "SmalLmForCausalLM": 5, + "SmallThinkerForCausalLM": 5, + "SolarForCausalLM": 5, + "StableLMAlphaForCausalLM": 5, + "TinyMixtralForCausalLM": 5, + "UniPhysGenQwen3ForCausalLM": 5, + "XQwen3ForCausalLM": 5, + "ActivationsGPTNeoForCausalLM": 4, + "AraGPT2LMHeadModel": 4, + "ArcticForCausalLM": 4, + "AsteriskForCausalLM": 4, + "AttnQwenForCausalLM": 4, + "BabyLMParaRNNForCausalLM": 4, + "BananaMind2MiniForCausalLM": 4, + "BertForSequenceClassification": 4, + "BitMamba2LM": 4, + "BlockFFNForCausalLM": 4, + "CableModel": 4, + "CagliostroForCausalLM": 4, + "CambrianLlamaForCausalLM": 4, + "ChessTRMForCausalLM": 4, + "Cohere2VisionForConditionalGeneration": 4, + "ComplexKDAForCausalLM": 4, + "CpmBeeForCausalLM": 4, + "CrystalCoderLMHeadModel": 4, + "CustomModel": 4, + "CustomPegasusForConditionalGeneration": 4, + "DashQQwen3_5MoeForCausalLM": 4, + "DeltaNetForCausalLM": 4, + "DiffuQwen35": 4, + "DuoLagunaForCausalLM": 4, + "EffT5ForConditionalGeneration": 4, + "EfficientDLM": 4, + "Emu3ForCausalLM": 4, + "EngGPTMoeForCausalLM": 4, + "EsmcARForCausalLM": 4, + "ExtendedLlamaForCausalLM": 4, + "ExtendedMptForCausalLM": 4, + "FP8LlamaForCausalLM": 4, + "Fast_dLLM_QwenForCausalLM": 4, + "Fuse3ForCausalLM": 4, + "G9v3ForCausalLM": 4, + "GPTCustomForCausalLM": 4, + "GPTModel": 4, + "GPTX2ForCausalLM": 4, + "Gemma2Model": 4, + "Gemma3TextModel": 4, + "GemmaModel": 4, + "GistLlamaForCausalLM": 4, + "HybridEchoForCausalLM": 4, + "HyenaDnaForCausalLM": 4, + "HyperNetEmbeddedQwen2ForCausalLM": 4, + "IdemFormerForCausalLM": 4, + "IsaacForConditionalGeneration": 4, + "KosineForCausalLM": 4, + "LCKVLlamaForCausalLM": 4, + "LLamaSynCABEL": 4, + "LatentMoELLaVAPhiForCausalLM": 4, + "Lfm2DSparkDraftModel": 4, + "LlamaWithIntervention": 4, + "LlavaCrystalForCausalLM": 4, + "LlavaMiniLlamaForCausalLM": 4, + "LlavaPhi3ForCausalLM": 4, + "LlavaQwen1_5ForCausalLM": 4, + "MPLUGDocOwlForConditionalGeneration": 4, + "MambaModelForCausalLM": 4, + "MegrezMoeForCausalLM": 4, + "MiniGPTForCausalLM": 4, + "MiniMaxForCausalLM": 4, + "MistralStarForCausalLM": 4, + "MoLMForCausalLM": 4, + "Muse2ForCausalLM": 4, + "MyLLaMa": 4, + "NGen3ForCasualLM": 4, + "NanoExpandForCausalLM": 4, + "NemotronFlashForCausalLM": 4, + "NexusForCausalLM": 4, + "NorT5ForConditionalGeneration": 4, + "OtterForConditionalGeneration": 4, + "OutlierMoEForCausalLM": 4, + "PawLlamaForCausalLM": 4, + "PebbleForCausalLM": 4, + "PersimmonForCausalLM": 4, + "Phi2MoeForCausalLM": 4, + "PipelinedBartForConditionalGeneration": 4, + "PlasmidLMForCausalLM": 4, + "PlusModelForCausalLM": 4, + "QaptaanForCausalLM": 4, + "QuarkForCausalLM": 4, + "Qwen2AudioTimeForConditionalGeneration": 4, + "Qwen3ASRForConditionalGeneration": 4, + "Qwen3ASVDForCausalLM": 4, + "Qwen3SharedMoeForCausalLM": 4, + "Qwen3SparseMoBEForCausalLM": 4, + "QwerkyLlamaMambaHybridForCausalLM": 4, + "RITAModelForCausalLM": 4, + "RaggedGemma4ForCausalLM": 4, + "RaggedGptOssForCausalLM": 4, + "RaggedQwen3_5MoeForCausalLM": 4, + "RecursiveCompressorLM": 4, + "RuGPT3XLForCausalLM": 4, + "SLMForCausalLM": 4, + "SeerAttnLlamaForCausalLM": 4, + "ShikraLlamaForCausalLM": 4, + "SlimMoEForCausalLM": 4, + "SparseASTForCausalLM": 4, + "SquaredReLUQwen3ForCausalLM": 4, + "StableDiffcoderForCausalLM": 4, + "Step3p7ForConditionalGeneration": 4, + "SurjoForCausalLM": 4, + "T5Gemma2ForConditionalGeneration": 4, + "T5LAForConditionalGeneration": 4, + "TPULlamaForCausalLM": 4, + "TRMTextISMForCausalLM": 4, + "TanukiForCausalLM": 4, + "TaoNetForCausalLM": 4, + "Telechat3ForCausalLM": 4, + "TransformerModel": 4, + "TrimKVQwen3ForCausalLM": 4, + "UrvashiForCausalLM": 4, + "VILAForCasualLM": 4, + "VaetkiForCausalLM": 4, + "ValleyLlamaForCausalLM": 4, + "VexionGPTForCausalLM": 4, + "VulavulaLlamaForCausalLM": 4, + "_SimpleLMForCausalLM": 4, + "i3HybridModel": 4, + "A2DGPTNeoXForCausalLM": 3, + "AETHERV27wayForCausalLM": 3, + "AlignGPTForCausalLM": 3, + "AndreaForCausalLM": 3, + "Apertus1p5ForConditionalGeneration": 3, + "AxionForCausalLM": 3, + "AxonForCausalLM": 3, + "BAAR25_3M": 3, + "BananaMind21Lite25MForCausalLM": 3, + "BananaMind21UnifiedForCausalLM": 3, + "Beit3LlavaLlamaForCausalLM": 3, + "BharatGPTForCausalLM": 3, + "BigBirdPegasusForCausalLM": 3, + "BlazeForCausalLM": 3, + "CFRDForCausalLM": 3, + "CoRMForCausalLM": 3, + "CodeLlamaForCausalLM": 3, + "CodeShellForCausalLM": 3, + "Cohere2Model": 3, + "ColMaskMoELLaVAQwen2ForCausalLM": 3, + "CosmicFish": 3, + "CustomDecoderOnlyT5": 3, + "CustomGPT2LMHeadModel": 3, + "DFlashLagunaForCausalLM": 3, + "DashQLlamaForCausalLM": 3, + "DeepseekV3ForCausalLMNextN": 3, + "DeltalmForConditionalGeneration": 3, + "DenseformerForCausalLM": 3, + "DharaARForCausalLM": 3, + "DiCoWForConditionalGeneration": 3, + "DistributedLlamaForCausalLM": 3, + "DockGenModel": 3, + "Doge2ForCausalLM": 3, + "EVELlamaForCausalLM": 3, + "EmovaQwen2ForCausalLM": 3, + "EnergyTransformer": 3, + "ErniePixelForCausalLM": 3, + "Evo1ForCausalLM": 3, + "FSDPQwen3ForCausalLM": 3, + "FineRMoeForCausalLM": 3, + "FlyForCausalLM": 3, + "ForCausalLM": 3, + "FreedomOmegaForCausalLM": 3, + "FuxiTranyuForCausalLM": 3, + "GADForAgenticModeling": 3, + "GFusionForDiffusionLM": 3, + "GLaMMForCausalLM": 3, + "GPT2": 3, + "GPT2ForCausalLM": 3, + "GPT2ForSequenceClassification": 3, + "GPT2MIMOLMHeadModel": 3, + "GPT2RoPEForCausalLM": 3, + "GPTLMHeadModel": 3, + "GPTModelForTextGeneration": 3, + "GPTNeoXJapaneseForCausalLM": 3, + "GPTRefactForCausalLM": 3, + "GRIN-MoE": 3, + "GRetrieverModel": 3, + "GazelleForConditionalGeneration": 3, + "GeluGPTForCausalLM": 3, + "Gemma2MoeForCausalLM": 3, + "Gemma3Model": 3, + "GiddForDiffusionLM": 3, + "GistT5ForConditionalGeneration": 3, + "Glm4vForConditionalGeneration": 3, + "GuppyLM": 3, + "HCXVisionForCausalLM": 3, + "HCXVisionV2ForCausalLM": 3, + "HebrewGPTForCausalLM": 3, + "HenlaConfedForCausalLM": 3, + "HgrnForCausalLM": 3, + "HrmTextMoEForCausalLM": 3, + "HybridForCausalLM": 3, + "InternS2PreviewForConditionalGeneration": 3, + "InternVLForConditionalGeneration": 3, + "Jais2ForCausalLM": 3, + "KW5V2ForCausalLM": 3, + "Kanana2TinyForCausalLM": 3, + "KayraForCausalLM": 3, + "KeyLM75M": 3, + "KeystoneFuseForCausalLM": 3, + "KlearMoeForCausalLM": 3, + "KoGumForCausalLM": 3, + "KoliberForCausalLM": 3, + "KonkanGPT": 3, + "LLTransformerForCausalLM": 3, + "LLaDAForCausalLM": 3, + "LLaMA": 3, + "LMHeadWithValueModel": 3, + "LOLALMHeadModel": 3, + "LSTLlamaForCausalLM": 3, + "LSTMForCausalLM": 3, + "LatexForCausalLM": 3, + "LayerWiseMiniCPMForCausalLM": 3, + "LightningForCausalLM": 3, + "Llama2BiasModel": 3, + "LlamaForConditionalGeneration": 3, + "LlamaForSequenceClassification": 3, + "Llamavision": 3, + "LlavaForCausalLM": 3, + "LongcatNextForCausalLM": 3, + "LongformerEncoderBARTDecoderForConditionalGeneration": 3, + "LowOnMindForCausalLM": 3, + "LumeesForCausalLM": 3, + "LummaForCausalLM": 3, + "MabaSparseForCausalLM": 3, + "Maira2ForConditionalGeneration": 3, + "MambaModel": 3, + "MatriochkaForCausalLM": 3, + "MeshModel": 3, + "MetisMoRLMHeadModel": 3, + "MidmLMHeadModel": 3, + "MiniBananaForCausalLM": 3, + "MiniBeatrixForCausalLM": 3, + "MiniMindOmni": 3, + "MoEGPTForCausalLM": 3, + "MoELLaVAPhiForCausalLM": 3, + "MoELLaVAQwen1_5ForCausalLM": 3, + "MoELLaVAStablelmForCausalLM": 3, + "MoLAForCausalLM": 3, + "ModuleFormerForCausalLM": 3, + "MolLLaMA": 3, + "MolmoForConditionalGeneration": 3, + "MonoFormerForCausalLM": 3, + "MoonfrostForCausalLM": 3, + "MuToRGemmaForCausalLM": 3, + "MultiscreenForCausalLM": 3, + "MyModelForCausalLM": 3, + "NGen4ForCausalLM": 3, + "NdmForCausalLM": 3, + "NemotronDenseAudexForConditionalGeneration": 3, + "NemotronHAudexForConditionalGeneration": 3, + "NeuroBLASTForCausalLM": 3, + "Nova1ForCausalLM": 3, + "OdinNextForCausalLM": 3, + "OffsetLlamaForCausalLM": 3, + "Olmo2ForSequenceClassification": 3, + "Open1BForCausalLM": 3, + "OpenMythosForCausalLM": 3, + "OpenThaiWilaiForCausalLM": 3, + "OrthrusLM": 3, + "OspreyLlamaForCausalLM": 3, + "PM_MiniFinLLM_Model": 3, + "PaloForCausalLM": 3, + "PeftModelForCausalLM": 3, + "PenguinVLQwen3ForCausalLM": 3, + "PipelinedGPT2LMHeadModel": 3, + "PipelinedT5ForConditionalGeneration": 3, + "ProphetNetForCausalLM": 3, + "PrunedQwen3_5MoeForCausalLM": 3, + "QUSSMForCausalLM": 3, + "QovaryxForCausalLM": 3, + "Qwen2AdapterForCausalLM": 3, + "Qwen2QuaRotForCausalLM": 3, + "Qwen3ForGuardModel": 3, + "Qwen3MoBEForCausalLM": 3, + "Qwen3TSForCausalLM": 3, + "Qwen3_5DLLMForConditionalGeneration": 3, + "Qwen3_Hybrid_OC_ForCausalLM": 3, + "QwenDFlareDraftModel": 3, + "RecombinationTransformerForCausalLM": 3, + "RecurrentQwenForCausalLM": 3, + "RoseX1ForCausalLM": 3, + "RostForCausalLM": 3, + "SDLMQwen2ForCausalLM": 3, + "SLM": 3, + "SMModelForCausalLM": 3, + "SeerAttnQwen2ForCausalLM": 3, + "SentinelBrainForCausalLM": 3, + "Sewy3ForCausalLM": 3, + "SewyV2ForCausalLM": 3, + "ShivikM4ForCausalLM": 3, + "SkyCRESTForCausalLM": 3, + "SkyForCausalLM": 3, + "SledForConditionalGeneration": 3, + "SoraForSLM": 3, + "SovythosModel": 3, + "SpatialLMLlamaForCausalLM": 3, + "Speech2TextTransformerForConditionalGeneration": 3, + "SphericalKANByteLM": 3, + "SrcProberForConditionalGeneration": 3, + "StieltjesGPT2ForCausalLM": 3, + "SusonoForCausalLM": 3, + "SwarmForCausalLM": 3, + "TRHashForCausalLM": 3, + "TTTLinearForCausalLM": 3, + "TTTMLPForCausalLM": 3, + "TarsierForConditionalGeneration": 3, + "TinyLM2": 3, + "TinyLlavaPhiForCausalLM": 3, + "TinyWayForCausalLM": 3, + "TranceptionLMHeadModel": 3, + "TreeFlashDraftModel": 3, + "TrtcV4ForCausalLM": 3, + "UpcycledLlamaForCausalLM": 3, + "VMistralForVisionText2Text": 3, + "VaswaniRoPEForConditionalGeneration": 3, + "VaultGemmaForCausalLM": 3, + "VideoBlipForConditionalGeneration": 3, + "WiolaForCausalLM": 3, + "YsnrfdForCausalLM": 3, + "Zarya": 3, + "ZsGPT2LMHeadModel": 3, + "modeling_camelidae.LlamaForCausalLM": 3, + "modeling_llama_butler.LlamaButlerForCausalLM": 3, + "vwLlamaForCausalLM": 3, + "xLSTMForCausalLM": 3, + "ABIRForCausalLM": 2, + "AILOLoopForCausalLM": 2, + "AQForCausalLM": 2, + "AROBabyLMForCausalLM": 2, + "ASGTransformerForCausalLM": 2, + "AWQCompatibleLlamaForCausalLM": 2, + "AWQCompatibleQwen2ForCausalLM": 2, + "AXK1ForCausalLM": 2, + "AXK2ForCausalLM": 2, + "AceRAG": 2, + "AdditionForCausalLM": 2, + "AdelicLlamaForCausalLM": 2, + "AlexaLlamaForCausalLM": 2, + "AmadablamForCausalLM": 2, + "Apriel2ForConditionalGeneration": 2, + "AquilaDenseForCausalLM": 2, + "AquilaMoeForCausalLM": 2, + "ArabicGPTModel": 2, + "ArlowForCausalLM": 2, + "AudioQwen2VLForConditionalGeneration": 2, + "AuroraForCausalLM": 2, + "AuroraGPT2ForCausalLM": 2, + "BDH": 2, + "BETForCausalLM": 2, + "BabyLMForCausalLM": 2, + "BacformerForCausalGM": 2, + "BackpackGPT2LMHeadModel": 2, + "BananaMind21CoderForCausalLM": 2, + "BananaMind2MediumForCausalLM": 2, + "BananaMind2MicroForCausalLM": 2, + "BarbetForCausalLM": 2, + "BareTorchForCausalLM": 2, + "BartForSequenceClassification": 2, + "BertGenerationDecoder": 2, + "BharatAI": 2, + "BigBirdPegasusForConditionalGeneration": 2, + "BijaForCausalLM": 2, + "BinaryLLMForCausalLM": 2, + "Bonsai2ForCausalLM": 2, + "BoomerForCausalLM": 2, + "BosunForDecision": 2, + "Braille256Model": 2, + "BrumbyForCausalLM": 2, + "BuddyGPTForCausalLM": 2, + "BunnyQwenForCausalLM": 2, + "BurtImmaForCausalLM": 2, + "C3QwenForCausalLM": 2, + "CBHybridLlamaForCausalLM": 2, + "CERPTForCausalLM": 2, + "CERPTForConditionalGeneration": 2, + "CIDDiffusionForMaskedLM": 2, + "CLIPVisionMarianForConditionalGeneration": 2, + "CMAForCausalLM": 2, + "CPTForConditionalGeneration": 2, + "CS336ForCausalLM": 2, + "CascadeForCausalLM": 2, + "CausalLM": 2, + "CheXagentForCausalLM": 2, + "CheXagentForConditionalGeneration": 2, + "ChemQ3MTPForCausalLM": 2, + "ClinamenForCausalLM": 2, + "CoDALanguageModel": 2, + "CoGPT2LMHeadModel": 2, + "CodifyForCausalLM": 2, + "Codva1ForCausalLM": 2, + "ConceptDominantGPTBertForPreTraining": 2, + "ContinuumForCausalLM": 2, + "CosmosForConditionalGeneration": 2, + "CostWiseGemmaForCausalLM": 2, + "CovSVDLlamaForCausalLM": 2, + "CubicHierLM": 2, + "CustomGPTForCausalLM": 2, + "CustomMixtralForCausalLM": 2, + "CustomModel3": 2, + "CustomQwen2ForCausalLM": 2, + "CustomQwen2Model": 2, + "CyclicFormerForCausalLM": 2, + "D3LMForMaskedLM": 2, + "DDLlamaForCausalLM": 2, + "DNikudModel": 2, + "DSHybridForCausalLM": 2, + "Data2VecTextForCausalLM": 2, + "DeCodon": 2, + "DebertaV2ForCausalLM": 2, + "DecoderOnlyT5Model": 2, + "DeepstackLlamaForCausalLM": 2, + "DenseLLM": 2, + "DibaForCausalLM": 2, + "DiffLlamaForCausalLM": 2, + "DistilBertForMaskedLM": 2, + "DomainTransformerForCausalLM": 2, + "DotsOCRForCausalLM": 2, + "DribbleForCausalLM": 2, + "DwarfForCausalLM": 2, + "E2TTTMLPForCausalLM": 2, + "E2TTTSwiGLUForCausalLM": 2, + "EchoesForCausalLM": 2, + "EcoacoForCausalLM": 2, + "EffLongT5ForConditionalGeneration": 2, + "EmberForCausalLM": 2, + "EmberProeliaForCausalLM": 2, + "EmenderForCausalLM": 2, + "EmuForCausalLM": 2, + "EnglishBaseForCausalLM": 2, + "EnhancedCustomLoRAModel": 2, + "EnsembleModelForCausalLM": 2, + "EveMoEForCausalLM": 2, + "ExpIvmeForDiffusionLMHub": 2, + "FSGPTMoEForCausalLM": 2, + "FabricForCausalLM": 2, + "FastPlus40mForCausalLM": 2, + "FegeoLlamaForCausalLM": 2, + "FlashGPTNeoXForCausalLM": 2, + "FleckModel": 2, + "Flex_Qwen2_5_VLMoeForConditionalGeneration": 2, + "Florence2ForConditionalGeneration": 2, + "FridayForCausalLM": 2, + "FuseGlmForCausalLM": 2, + "FusionInDecoderForConditionalGeneration": 2, + "G0NanoForCausalLM": 2, + "GAD2ForAgenticModeling": 2, + "GLUSForCausalLM": 2, + "GPT2CustomLMHeadModel": 2, + "GPT2MTP": 2, + "GPT2Vision": 2, + "GPT2WithHMForCausalLM": 2, + "GPT4DevForCausalLM": 2, + "GPTLanguageModel": 2, + "GPTMiniForCausalLM": 2, + "GTLMForCausalLM": 2, + "GaudiLlamaForCausalLM": 2, + "GeckoForConditionalGeneration": 2, + "Gemma3ForCausalLMTTT": 2, + "Gemma3MoEForCausalLM": 2, + "Gemma3MoeForCausalLM": 2, + "Gemma3nForCausalLM": 2, + "Gemma4DSparkModel": 2, + "Gemma4E2BItHybridForCausalLM": 2, + "Gemma4TextForCausalLM": 2, + "GeoChatLlamaForCausalLM": 2, + "GeoVForCausalLM": 2, + "GistGPTNeoForCausalLM": 2, + "GitLlamaForCausalLM": 2, + "GraniteMoeSWAForCausalLM": 2, + "Grok1ForCausalLM": 2, + "Grok1ModelForCausalLM": 2, + "H2OVLChatModel": 2, + "H3ForCausalLM": 2, + "HCAForCausalLM": 2, + "HFByteETM": 2, + "HFCausalModel": 2, + "HFTransformerModel": 2, + "HanForgeForCausalLM": 2, + "HelionForCausalLM": 2, + "HfMoondream": 2, + "HolographicQwenForCausalLM": 2, + "HrmForCausalLM": 2, + "HuazangQWenForCausalLM": 2, + "HumanVForCausalLM": 2, + "HybridFourierLM": 2, + "HybridGPTForCausalLM": 2, + "ILLaDAForCausalLM": 2, + "IdeficsForCausalLM": 2, + "IndraForCausalLM": 2, + "InnovatorVLForConditionalGeneration": 2, + "InternLM2ForRewardModel": 2, + "InternLMXComposerForCausalLM": 2, + "InternS1ForConditionalGeneration": 2, + "InternS2MobiusForConditionalGeneration": 2, + "IvmeXLForCausalLM": 2, + "JeevesForCausalLM": 2, + "JetNemotronForCausalLM": 2, + "JiRackTernaryModel": 2, + "JudgeXL": 2, + "KNKVFForCausalLM": 2, + "KORMoForCausalLMWithMTP": 2, + "Kanana2VecModel": 2, + "KimiK2ForCausalLM": 2, + "LEDForConditionalGeneration": 2, + "LLAMIAFluxForConditionalGeneration": 2, + "LLM": 2, + "LLMForCausalLM": 2, + "LLaMAVIDLlavaForConditionalGeneration": 2, + "LLaMA_model": 2, + "LLaVAOneVision1_5_ForConditionalGeneration": 2, + "LSWTForCausalLM": 2, + "LaCTForCausalLM": 2, + "LanceAI": 2, + "LaneformerForCausalLM": 2, + "LangFlow": 2, + "LatentMoELLaVAQwen2ForCausalLM": 2, + "LatentMoELLaVAQwen3ForCausalLM": 2, + "LightbrainHybridForCausalLM": 2, + "LilmForCausalLM": 2, + "LingLongForCausalLM": 2, + "LiquidForCausalLM": 2, + "LiveMemForCausalLM": 2, + "LlaaaLlamaForCausalLM": 2, + "Llama3ForCausalLM": 2, + "LlamaButlerForCausalLM": 2, + "LlamaForCausalLMEagle": 2, + "LlamaHydraForCausalLM": 2, + "LlamaMLAForCausalLM": 2, + "LlamaMoDForCausalLM": 2, + "LlamaMoEUpscalingForCausalLM": 2, + "LlamaMoeForCausalLM": 2, + "LlamaSeqEndMaskBMForCausalLM": 2, + "LlamaVarLayerForCausalLM": 2, + "LlavaGPT2ForCausalLM": 2, + "LlavaQWenForCausalLM": 2, + "LlavaStableLMEpochForCausalLM": 2, + "LlavaStablelmForCausalLM": 2, + "LocalLLMForCausalLM": 2, + "LocateAnythingForConditionalGeneration": 2, + "LongT5ForConditionalGeneration": 2, + "LongcatFlashOmniForCausalLM": 2, + "LongcatFlashSparseForCausalLM": 2, + "LongformerEncoderDecoderForConditionalGeneration": 2, + "MCGPTForCausalLM": 2, + "MCQHFModel": 2, + "MGMLlamaForCausalLM": 2, + "MGMOmniForCausalLM": 2, + "MMGPTLlamaForCausalLM": 2, + "MMGPTQwenForCausalLM": 2, + "MT5Model": 2, + "MUDDPythia": 2, + "MaccyForCausalLM": 2, + "MagicForConditionalGeneration": 2, + "MalayalamMoEForCausalLM": 2, + "Mamba3CausalLM": 2, + "MantaForConditionalGeneration": 2, + "MedusaModel": 2, + "MegLMForCausalLM": 2, + "MegatronForCausalLM": 2, + "MemBartForConditionalGeneration": 2, + "MetaDiffusionForCausalLM": 2, + "MetaLLMForCausalLM": 2, + "MiMoAudioForCausalLM": 2, + "MiMoMamba2ForCausalLM": 2, + "MicroLoopForDiffusionLM": 2, + "MinGRULMForCausalLM": 2, + "MinSparkForCausalLM": 2, + "MiniCPMO": 2, + "MiniDeepSeekV3ForCausalLM": 2, + "MiniDeepSeekV4ForCausalLM": 2, + "MiniGPT": 2, + "MiniGeminiGemmaForCausalLM": 2, + "MiniGeminiMixtralForCausalLM": 2, + "MiniLlama3ForCausalLM": 2, + "MiniLlamaForCausalLM": 2, + "MiniMaxM3VLForCausalLM": 2, + "MiniQwen3NextForCausalLM": 2, + "MinistralDualRopeForCausalLM": 2, + "MiphaPhiForCausalLM": 2, + "MistralDenseFormerForCausalLM": 2, + "MixsenseLlamaForCausalLM": 2, + "MoELLaVAQwen2ForCausalLM": 2, + "MobileReasoningLLM": 2, + "MobilintQwen3Eagle3ForCausalLM": 2, + "MochivaForCausalLM": 2, + "Model": 2, + "ModelForCausalLM": 2, + "MoeForCausalLM": 2, + "MoeLlamaForCausalLM": 2, + "MoeTransformerForCausalLM": 2, + "MoedlForCausalLM": 2, + "MolexarForCausalLM": 2, + "MolformerForCausalLM": 2, + "MonkeyLMHeadModel": 2, + "MonoidForCausalLM": 2, + "MplugOwlForConditionalGeneration": 2, + "MultimodalStarcoder2ForCausalLM": 2, + "MyBaichuanForCausalLM": 2, + "MyLlamaForTokenClassification": 2, + "MyMossForCausalLM": 2, + "NGen3Model": 2, + "NMMaskMoELLaVAQwen2ForCausalLM": 2, + "NanoChatTopKForCausalLM": 2, + "NanoGPTLMHeadModel": 2, + "NanoLM": 2, + "NanochatWasmFusedModel": 2, + "NekoMindMoeForCausalLM": 2, + "NeroXSAForCausalLM": 2, + "NilexForCausalLM": 2, + "Nushy5ForCausalLM": 2, + "OBILanguageModel": 2, + "OLMoModelForCausalLM": 2, + "OPT_PromptTuned_For_SentimentAnalysis": 2, + "OURSForCausalLM": 2, + "OmniLlamaForCausalLM": 2, + "OpenBAForConditionalGeneration": 2, + "OpenModel": 2, + "OryxQwen2ForCausalLM": 2, + "OryxQwenForCausalLM": 2, + "OtterLM": 2, + "OxMiniForCausalLM": 2, + "PKVGPT": 2, + "PLBartForCausalLM": 2, + "PLETinyLMForCausalLM": 2, + "PLMForCausalLM": 2, + "PaluLlamaForCausalLM": 2, + "ParallaxForCausalLM": 2, + "ParchmentForCausalLM": 2, + "Pebble10MLM": 2, + "PebbleGPTForCausalLM": 2, + "PebbleLMForCausalLM": 2, + "PersonaMiniForCausalLM": 2, + "PhixtralForCausalLM": 2, + "PicoLMV2ForCausalLM": 2, + "PinkElephantForCausalLM": 2, + "Pix2SeqForConditionalGeneration": 2, + "PlutoForCausalLM": 2, + "PolyLLaMAForCausalLM": 2, + "PoptorchPipelinedBartForConditionalGeneration": 2, + "PoptorchPipelinedT5ForConditionalGeneration": 2, + "PothanaForCausalLM": 2, + "PrajnaStudentMultiLayer": 2, + "QraXAiForCausalLM": 2, + "QuartzForCausalLM": 2, + "Qwen2AudioForConditionalGeneration": 2, + "Qwen2CapreseForCausalLM": 2, + "Qwen2FlatQuantForCausalLM": 2, + "Qwen2ForSequenceClassification": 2, + "Qwen2LayerwiseSAEForCausalLM": 2, + "Qwen2SteeringVectorForCausalLM": 2, + "Qwen2WithRegressionHead": 2, + "Qwen2_5OmniThinkerForCausalLM": 2, + "Qwen2_5OmniThinkerForConditionalGeneration": 2, + "Qwen2_5_VL_PGNForConditionalGeneration": 2, + "Qwen2blForCausalLM": 2, + "Qwen3MoeFusedForCausalLM": 2, + "Qwen3MoeModel": 2, + "Qwen3QuaRotForCausalLM": 2, + "Qwen3ReasoningForCausalLM": 2, + "Qwen3TerminatorForCausalLM": 2, + "Qwen3_5MoEForCausalLM": 2, + "Qwen3_5Model": 2, + "Qwen3_5TextForCausalLM": 2, + "Qwen3_5_MoeForCausalLM": 2, + "Qwen4ExpTextForCausalLM": 2, + "RACE": 2, + "RNSAQwen3ForCausalLM": 2, + "RPTForCausalLM": 2, + "RRForConditionalGeneration": 2, + "RRTForCausalLM": 2, + "RWKV6ForCausalLM": 2, + "RWModel": 2, + "RecGPTForCausalLM": 2, + "RewardModel": 2, + "RobertaModel": 2, + "RobertaPreLayerNormForCausalLM": 2, + "RotaryGPT2LMHeadModel": 2, + "SAIForCausalLM": 2, + "SCANForCausalLM": 2, + "SDLCSLM": 2, + "SESAMEForCausalLM": 2, + "SLMoEForCausalLM": 2, + "SMTModelForCausalLM": 2, + "SPILlavaMPTForCausalLM": 2, + "SakhiModel": 2, + "SeqaxLMHeadModel": 2, + "SerayukiForCausalLM": 2, + "ShinraForCausalLM": 2, + "ShrinkModelForCausalLM": 2, + "SimpleAttentionNetwork": 2, + "SimpleStories4MModel": 2, + "SimpleStoriesForCausalLM": 2, + "SliderGPT": 2, + "SmallLanguageModel": 2, + "SoloGPTForCausalLM": 2, + "SoloLLMForCausalLM": 2, + "SovereignMemoryTwinForCausalLM": 2, + "SparkForCausalLM": 2, + "SpecT1ForCausalLM": 2, + "SpectusForConditionalGeneration": 2, + "Starcoder2Model": 2, + "Step1ForCausalLM": 2, + "Step3VLForConditionalGeneration": 2, + "Step3vForConditionalGeneration": 2, + "StepAudio2ForCausalLM": 2, + "StepVLForConditionalGeneration": 2, + "StreamLfm2MoeForCausalLM": 2, + "SwitchGPT2ForCausalLM": 2, + "T5GraphForConditionalGeneration": 2, + "TALlavaGemmaForCausalLM": 2, + "TLiveOmniForConditionalGeneration": 2, + "TPUQwen3ForCausalLM": 2, + "TamazightForCausalLM": 2, + "TanAiForCausalLM": 2, + "TeleFLMForCausalLM": 2, + "TestGeniyForCausalLM": 2, + "TextToTextModel": 2, + "ThetaForCausalLM": 2, + "TinyBuddyForCausalLM": 2, + "TinyForCausalLM": 2, + "TinyGPT": 2, + "TinyLLM": 2, + "TinyMindForCausalLM": 2, + "TinyStoriesGPT": 2, + "TinyllmForCausalLM": 2, + "TrOCRForCausalLM": 2, + "TraXLMistralForCausalLM": 2, + "TransformerLMForCausalLM": 2, + "TridaForDLM": 2, + "UnboxForCausalLM": 2, + "UunoForCausalLM": 2, + "VCoderDSLlavaLlamaForCausalLM": 2, + "VCoderLlavaLlamaForCausalLM": 2, + "VLCLIPGPTNeoXForCausalLM": 2, + "VLite7Mini20mForCausalLM": 2, + "VaayuForCausalLM": 2, + "VibeVoiceRealTimeForConditionalGeneration": 2, + "Videollama2Qwen2ForCausalLM": 2, + "Videollama3Qwen2ForCausalLM": 2, + "VoRAForCausalLM": 2, + "VoxtralForConditionalGeneration": 2, + "WasmInterpreterTransformer": 2, + "WhaleyeForConditionalGeneration": 2, + "WhisperAccentForConditionalGeneration": 2, + "WindEdgeForCausalLM": 2, + "WinnowLagunaForCausalLM": 2, + "WisentQwen2ForCausalLM": 2, + "WyrmlingForCausalLM": 2, + "XLMRobertaForSequenceClassification": 2, + "XMistralForCausalLM": 2, + "XeDroplycheeForCausalLM": 2, + "YForCausalLM2": 2, + "YForCausalLM3": 2, + "YForCausalLM31": 2, + "YasinForCausalLM": 2, + "YatFullGPTForCausalLM": 2, + "YiVLForCausalLM": 2, + "ZYR3ForCausalLM": 2, + "ZZJRabbit22ForCausalLM": 2, + "ZZJRabbit2ForCausalLM": 2, + "ZeusForCausalLM": 2, + "axiomForCausalLM": 2, + "i3": 2, + "i3Model": 2, + "infllmv2_Qwen3ForCausalLM": 2, + "modeling_grove_moe.GroveMoeForCausalLM": 2, + "omFlaxT5ForConditionalGeneration": 2, + "smallLlamaForCausalLM": 2, + "tnl1-385m-10b-token_no-act": 2, + "A194LogitEnsembleForCausalLM": 1, + "A2DQwen2LMHeadModel": 1, + "A2DQwen3_5LMHeadModel": 1, + "A2DQwenLMHeadModel": 1, + "AAIEDDenseForCausalLM": 1, + "AAIEDDenseGFTForCausalLM": 1, + "AAIEMoEForCausalLM": 1, + "ABU2HEADMODEL": 1, + "ACSwiGLUForCausalLM": 1, + "AETHERMicroForCausalLM": 1, + "AETHERV211AttnForCausalLM": 1, + "AMITForConditionalGeneration": 1, + "ARAr381MForCausalLM": 1, + "ARMTForCausalLM": 1, + "ASVDOPTForCausalLM": 1, + "AXI with transformers": 1, + "AbcTransformer": 1, + "AbiaForCausalLM": 1, + "AdaVocabGemmaForCausalLM": 1, + "AdaVocabGemmaforCausalLM": 1, + "AdaVocabQwen2ForCausalLM": 1, + "AdaptiveRiverLM": 1, + "AdditionTransformer": 1, + "AeroForConditionalGeneration": 1, + "AetherMindForCausalLM": 1, + "AetherStoryModel": 1, + "AgnesForCausalLM": 1, + "AgoraForCausalLM": 1, + "Aicraftar-Tharo.G-ConditionalGeneration": 1, + "AlbertMoE": 1, + "AlfredUnimodelForConditionalGeneration": 1, + "AliceT5MoEForConditionalGeneration": 1, + "AlinlightForCausalLM": 1, + "AlphaErForCausalLM": 1, + "AncientAIV": 1, + "AntiHalForConditionalGeneration": 1, + "AnubisMoeForCausalLM": 1, + "ApolloForCausalLM": 1, + "ApproxDumbForCausalLM": 1, + "AprielHForCausalLM": 1, + "ArcanaLlamaForCausalLM": 1, + "ArcanalamaForCausalLM": 1, + "AresModel": 1, + "ArgonneModelParallel": 1, + "AriesForCausalLM": 1, + "ArkaV3ForCausalLM": 1, + "ArmanForCausalLM": 1, + "ArmoRMForSequenceClassification": 1, + "ArzLMForCausalLM": 1, + "AscleLMForCausalLM": 1, + "AtomK3DSparkModel": 1, + "AttnOnlyForCausalLM": 1, + "AudioOnlyThinker": 1, + "Aurora1.0": 1, + "Aurora80K": 1, + "AutoGUILMHeadModel": 1, + "AutoModelForSeq2SeqLM": 1, + "Autoencoder": 1, + "AveyDecoderMoEForCausalLM": 1, + "AveyForCausalLM": 1, + "AxonModel": 1, + "AyaVisionForConditionalGeneration": 1, + "BREENForCausalLM": 1, + "BacLMForCausalLM": 1, + "BackboneConceptLM": 1, + "Bagel": 1, + "BailingMoeLinearForCausalLM": 1, + "Baiwen3ForCausalLM": 1, + "BananaMind21TestForCausalLM": 1, + "BananaMind2MoEForCausalLM": 1, + "BanglaGSGForCausalLM": 1, + "BanglaGambaForCausalLM": 1, + "BartPrefixPropForConditionalGeneration": 1, + "BaselineModel": 1, + "BatGPTForCausalLM": 1, + "BeetleMoEHF": 1, + "BernaForCausalLM": 1, + "BertForCausalLM": 1, + "BertForMTSparseFFDEditModel": 1, + "BertForTokenClassification": 1, + "BetterGPTForCausalLM": 1, + "BharatGPTForConditionalGeneration": 1, + "BharataiForCausalLM": 1, + "BiRWKV7ForCausalLM": 1, + "BiatronForCausalLM": 1, + "BigBrainLanguageModel": 1, + "BioMedGPTForCausalLM": 1, + "BioPhysKimiDualBrainForCausalLM": 1, + "BitNetGPTForCausalLM": 1, + "BitSkipV1ForCausalLMWithEarlyExit": 1, + "BitSkipV2ForCausalLMWithEarlyExit": 1, + "BitSkipV3ForCausalLM": 1, + "BlastModelForCausalLM": 1, + "BlipForConditionalGeneration": 1, + "BltForCausalLM": 1, + "BoraMoEForCausalLM": 1, + "BranchyCausalModel": 1, + "BridgeVQAModel": 1, + "BucketMemoryModel": 1, + "BungeoForCausalLM": 1, + "BunnyMiniCPMForCausalLM": 1, + "BunnyQwen2ForCausalLM": 1, + "ByteGPTForCausalLM": 1, + "CALIForCausalLM": 1, + "CED": 1, + "CERPTMultimodalForConditionalGeneration": 1, + "CIDModel": 1, + "CLIPVisionMBartForConditionalGeneration": 1, + "COCONUTGPT2": 1, + "CODIGPT2": 1, + "COReForCausalLM": 1, + "CPMAntForCausalLM": 1, + "CRANEAIModel": 1, + "CXRMate2ForConditionalGeneration": 1, + "CXRMateEDModel": 1, + "CaReAQA": 1, + "CambrianPhi3ForCausalLM": 1, + "CampGPT": 1, + "CatsModel": 1, + "CausalLMoEForCausalLM": 1, + "CelerityLMHeadModel": 1, + "ChameleonXLLMXForConditionalGeneration": 1, + "CharNgramForCausalLM": 1, + "ChatGLM2NSAForCausalLM": 1, + "ChatGlmForCausalLM": 1, + "ChatTitleModel": 1, + "ChatUniViLlamaForCausalLM": 1, + "ChessGPTForCausalLM": 1, + "ChessLLM": 1, + "ChessModel": 1, + "CiModelForCausalLM": 1, + "CiloForCausalLM": 1, + "CinnabarLMForCausalLM": 1, + "CircuitGPTForCausalLM": 1, + "ClarityMR1ForCausalLM": 1, + "CleverModel": 1, + "ClinicalLlamaForCausalLM": 1, + "ClokCEMForCausalLM": 1, + "CloverLMForCausalLM": 1, + "CodeBharatForCausalLM": 1, + "CodeShell4bitForCausalLM": 1, + "CodeT5pBimodalModel": 1, + "CoffeeChatAI": 1, + "CogViewForCausalLM": 1, + "CognitiveAgentForCausalLM": 1, + "CoherenceMomentumModel": 1, + "ColarLlama": 1, + "CollisionForCausalLM": 1, + "CombinedLMForCausalLM": 1, + "CompliantLLMModel": 1, + "CompressedLlamaForCausalLM": 1, + "ConceptLMDFlashModel": 1, + "ConditionalGPT": 1, + "ConditionalGPT2LMHeadModel": 1, + "ConfigfarsModel": 1, + "Continue1ForCausalLM": 1, + "ContinuousQwen3ForCausalLM": 1, + "ConvGPTForCausalLM": 1, + "ConvaiCausalLM": 1, + "CraftlyRobotForCausalLM": 1, + "CroweLogicMiniForCausalLM": 1, + "CubicV11LongContext": 1, + "CubiczanMoEForCausalLM": 1, + "CuriousForCausalLM": 1, + "CustomBioGptForCausalLM": 1, + "CustomGPT": 1, + "CustomGPTModel": 1, + "CustomLoRAModel": 1, + "CustomModel4": 1, + "CustomModel5": 1, + "CustomModelForCausalLM": 1, + "CustomTagalogLLM": 1, + "CustomTransformerForCausalLM": 1, + "Custom_MPTForCausalLM": 1, + "D3PMSanskritModel": 1, + "DCFormer": 1, + "DCPythia": 1, + "DEXV1LMHeadModel": 1, + "DFMModel": 1, + "DGPT": 1, + "DIT": 1, + "DIVEdoc": 1, + "DLMLlamaForCausalLM": 1, + "DMTDQwen3ForCausalLM": 1, + "DPMMForCausalLM": 1, + "DakitariInstructModel": 1, + "DaltonForCausalLM": 1, + "DartForCausalLM": 1, + "DarwinDuoOrchestrator": 1, + "DashQGlm4MoeForCausalLM": 1, + "DashQNemotronHForCausalLM": 1, + "DashaHFForCausalLM": 1, + "DatForCausalLM": 1, + "DeTiME": 1, + "DeViLQwen2ForCausalLM": 1, + "DebertaV2ForSequenceClassification": 1, + "DebertaV2PairRM": 1, + "DecoderOnlyTransformer": 1, + "DecoderTransformerForCausalLM": 1, + "DeepLlamaForCausalLM": 1, + "DeepSeekNanoForCausalLM": 1, + "DeepSeekR1ForCausalLM_PolySurgery": 1, + "DeepSeekR1_UnifiedIdemFormer": 1, + "DeepSeekV4": 1, + "DeepseekFixedForCausalLM": 1, + "DeepseekOCRForCausalLM": 1, + "DeepseekOcrForConditionalGeneration": 1, + "DeepseekV2MoBEForCausalLM": 1, + "DeepseekV2SparseMoBEForCausalLM": 1, + "DeepseekVLV2ForCausalLM": 1, + "DendroForCausalLM": 1, + "DenseK3ForCausalLM": 1, + "DeplyzeMiniForCausalLM": 1, + "DexV1LMHeadModel": 1, + "DharaForMaskedDiffusion": 1, + "DiffuQwen3": 1, + "DiffusionLLM": 1, + "DiseaseGPTModel": 1, + "DistMoeForConditionalGeneration": 1, + "DistillixForCausalLM": 1, + "DizelLM": 1, + "DominoDraftModel": 1, + "DotLMForCausalLM": 1, + "DragonbornTransformerModel": 1, + "DribbleLlamaForCausalLM": 1, + "DropLycheeForCausalLM": 1, + "DuchifatForCausalLM": 1, + "DumbSetLanguageModel": 1, + "DummyLlamaForCausalLM": 1, + "DusMistralForCausalLM": 1, + "DynColMaskMoELLaVAQwen3ForCausalLM": 1, + "DynHierColMaskMoELLaVAQwen2ForCausalLM": 1, + "DynamicMindMoEForCausalLM": 1, + "DynamicNeuralNetwork": 1, + "EMGForCausalLM": 1, + "EMGMTSForCausalLM": 1, + "Eagle3LlamaForCausalLM": 1, + "EasyFormerLMHeadModel": 1, + "EditGPTMistralForCausalLM": 1, + "EditableQwen3ForCausalLM": 1, + "EfficientNetForImageClassification": 1, + "EgoLLM": 1, + "ElBartForConditionalGeneration": 1, + "ElasticGPT": 1, + "ElasticT5ForConditionalGeneration": 1, + "ElysiumForCausalLM": 1, + "Ember2ForCausalLM": 1, + "EmenderGDN2ForCausalLM": 1, + "EmployeeMicroSLMForCausalLM": 1, + "EncoderDecoderForConditionalGeneration": 1, + "EngramQwenForCausalLM": 1, + "EnigmaForCausalLM": 1, + "ErkLinearForCausalLM": 1, + "EryonForCausalLM": 1, + "EryonModel": 1, + "EshmunForCausalLM": 1, + "Esm2LlamaInstructForCausalLM": 1, + "EvafrillMoForCausalLM": 1, + "EvilModel": 1, + "EvoMistralForCausalLM": 1, + "Exaone4ForCausalLMConv": 1, + "Exaone4_5_ForConditionalGeneration": 1, + "ExaoneTDForCausalLM": 1, + "ExpandedJetMoEForCausalLM": 1, + "ExperiementalForCausalLM": 1, + "ExportableGPTScratchForCausalLM": 1, + "ExtraAI": 1, + "FEGeoQwen2ForCausalLM": 1, + "FERRETLlamaForCausalLM": 1, + "FHN_T4Max_199M": 1, + "FM9GForCausalLM": 1, + "FSDPGptOssForCausalLM": 1, + "FSDPLLaDAUPMModelLM": 1, + "FSDPLlamaForCausalLM": 1, + "FSDPT5ForConditionalGeneration": 1, + "FSGPTForCausalLM": 1, + "FaberGenesisForCausalLM": 1, + "FairseqT5ForConditionalGeneration": 1, + "FastLigerLlamaForCausalLM": 1, + "FastPlus125mForCausalLM": 1, + "FastPlusForCausalLM": 1, + "FastyForCausalLM": 1, + "FelaForCausalLM": 1, + "Fern3BModel": 1, + "FiPhi-NeuralArk-3.9-Ultra": 1, + "FiPhi-NeuralMark-V3": 1, + "FidelForCausalLM": 1, + "FieldsHubModel": 1, + "FinGPTForCausalLM": 1, + "FixedEnhancedHybridTransformer": 1, + "FlamingoForCausalLM": 1, + "FlaxGPTJForCausalLM": 1, + "FlexRankModel": 1, + "FlintForCausalLM": 1, + "FlyGPTForCausalLM": 1, + "FontaineLM": 1, + "ForgeLM": 1, + "ForwardBackwardRepairModel": 1, + "FrawdLLMForCausalLM": 1, + "FrenchLLMForCausalLM": 1, + "FrontD11M": 1, + "Fuse2ForCausalLM": 1, + "Fuse3V2ForCausalLM": 1, + "FutureGQ47MForCausalLM": 1, + "FutureGQ47qForCausalLM": 1, + "GAIForCausalLM": 1, + "GITForCausalLM": 1, + "GPT2CompetitiveMoE": 1, + "GPT2ForQuestionAnswering": 1, + "GPT2HeadWithValueModel": 1, + "GPT2LMHeadModelForMultiTokenPrediction": 1, + "GPT2LMHeadModelWithRoPE": 1, + "GPT2PrefixTuningWithLMHeadModel": 1, + "GPT2WithRoles": 1, + "GPT2hlcLMHeadModel": 1, + "GPT300M": 1, + "GPT3DevLM": 1, + "GPTBertForMaskedLM": 1, + "GPTBigCodeForSequenceClassification": 1, + "GPTBigCodeLMHeadModel": 1, + "GPTBigCodeModel": 1, + "GPTForHF": 1, + "GPTJLoraForCausalLM": 1, + "GPTJMoEForCausalLM": 1, + "GPTNeoModel": 1, + "GPTNeoXLongForCausalLM": 1, + "GPTOSSForCausalLM": 1, + "GPTOSSMoEModel": 1, + "GPTS14MForCausalLM": 1, + "GPTS25ForCausalLM": 1, + "GPTS3ForCausalLM": 1, + "GPTSDPReLULMHeadModel": 1, + "GPTSanJapaneseForConditionalGeneration": 1, + "GPTX3ForCausalLM": 1, + "GPTXForCausalLM": 1, + "GQAGPT2": 1, + "GalahadForCausalLM": 1, + "GatedDeltaProductForCausalLM": 1, + "GatorForCausalLM": 1, + "GeminiForCausalLM": 1, + "Gemma2BlockAttnResForCausalLM": 1, + "Gemma2ForCausalLM_PolySurgery": 1, + "Gemma3AbliteratedForConditionalGeneration": 1, + "Gemma3PXForCausalLM": 1, + "Gemma4NanoForCausalLM": 1, + "Gemma4TextModel": 1, + "Gemma4UnifiedForCausalLM": 1, + "GemmagainForCausalLM": 1, + "GenerativePromptLlama": 1, + "GeoMotionGPTForCausalLM": 1, + "GexQwenForCausalLM": 1, + "GiftOfGabForCausalLM": 1, + "GigaChatAudioForConditionalGeneration": 1, + "GliDeForCausalLM": 1, + "Glm4MoeLitePlusPlusForCausalLM": 1, + "GlmOcrForConditionalGeneration": 1, + "GlmasrForConditionalGeneration": 1, + "GlubLM": 1, + "GoatVVVForCausalLM": 1, + "GobbledygookForCausalLM": 1, + "GodQueenIVForCausalLM": 1, + "GomeForCausalLM": 1, + "GopuForCausalLM": 1, + "GptNeoForCausalLM": 1, + "GptOssAsideForCausalLM": 1, + "GptOssPuzzleForCausalLM": 1, + "GptOssVLForConditionalGeneration": 1, + "GraniteKVShareForCausalLM": 1, + "GraniteSWAForCausalLM": 1, + "GraphLlamaForCausalLM": 1, + "GraphT5TransformerForConditionalGeneration": 1, + "GraphTokenLM": 1, + "GreedyModel": 1, + "GritLM": 1, + "Grok2ForCausalLM": 1, + "GrokForCausalLM": 1, + "GroundedBLIPLMForConditionalGeneration": 1, + "GuppyLMForCausalLM": 1, + "HEDForConditionalGeneration": 1, + "HFBasicModel": 1, + "HFETM": 1, + "HFGPTModel": 1, + "HFHealthSLM": 1, + "HFLuminoLexForCausalLM": 1, + "HFOpenMoeForCausalLM": 1, + "HFPForCausalLM": 1, + "HIComQwen2ForCausalLM": 1, + "HLM5ForCausalLM": 1, + "HModel": 1, + "HNetForCausalLM": 1, + "HRMCosmicFish": 1, + "HRMForCausalLM": 1, + "HTDNModel": 1, + "HYV3VLForCausalLM": 1, + "HaikuForCausalLM": 1, + "HaipaiForCausalLM": 1, + "HaipaiLM": 1, + "HaloSForCausalLM": 1, + "HaltCoTForCausalLM": 1, + "HangulGemmaDeobfuscator": 1, + "HanseForCausalLM": 1, + "HanziForCausalLM": 1, + "HelionOSCForCausalLM": 1, + "HelloAgentForCausalLM": 1, + "HenyoModel": 1, + "HeteroLlamaForCausalLM": 1, + "HiLSForCausalLM": 1, + "HierColMaskMoELLaVAQwen2ForCausalLM": 1, + "HindiCausalLM": 1, + "HindiSLMForCausalLM": 1, + "HixtralForCausalLM": 1, + "HookedLlamaForCausalLM": 1, + "HubertGPTNeoXCrossForConditionalGeneration": 1, + "HuggingFaceCompatibleModel": 1, + "HumanGPTForCausalLM": 1, + "HummingbirdForCausalLM": 1, + "HunyuanImage3ForCausalMM": 1, + "HushNanoForCausalLM": 1, + "HuskyForConditionalGeneration": 1, + "HybriKoForCausalLM": 1, + "HybridGatedDeltaNetForCausalLM": 1, + "HybridLMForCausalLM": 1, + "HybridMoRMoEForCausalLM": 1, + "HybridTinyForCausalLM": 1, + "HybridTransformerV2ForCausalLM": 1, + "HydrogenForCausalLM": 1, + "Hyper0xModel": 1, + "HyperMambaLM": 1, + "HyperpartisanModel": 1, + "I3ForCausalLM": 1, + "IBCEGPT2LowRank": 1, + "ICONNForCausalLM": 1, + "IIIForCausalLM": 1, + "IKNN-Rl1-A1ForCausalLM": 1, + "IQuestPLTCoderForCausalLM": 1, + "ISO20022ForCausalLM": 1, + "ISOMR1Coder16BMoEForCausalLM": 1, + "ISOMR1Edge130MMoEForCausalLM": 1, + "ISOMR1Enterprise40BForCausalLM": 1, + "ISOMR1Reasoning15BForCausalLM": 1, + "IcarusForCausalLM": 1, + "Idefics2ForVisionText2Text": 1, + "IlamaForCausalLM": 1, + "IlluminatorLMHeadModel": 1, + "ImbalanceTexxWithBertForConditionalGeneration": 1, + "ImpPhi3ForCausalLM": 1, + "ImpQwen2ForCausalLM": 1, + "Imu1ForCausalLM": 1, + "IndexLM": 1, + "InductionForCausalLM": 1, + "InferenceMemoryWrapper": 1, + "InfiMMHDModel": 1, + "InfiMMVicunaModel": 1, + "InfiMMZephyrModel": 1, + "Int8LlamaForCausalLM": 1, + "IntellixForConditionalGeneration": 1, + "InversionFromHiddenStatesModel": 1, + "Ions1Model": 1, + "Isllmai50mForCausalLM": 1, + "ItaloForCausalLM": 1, + "IvmeCoderV1ForCausalLM": 1, + "IvmeConversateSModel": 1, + "IvmeConversateSV2InstructModel": 1, + "IvmeForCausalLM": 1, + "JeeneyModel": 1, + "JiRackTernary1B": 1, + "JiRackTernaryPro1B": 1, + "JibayAiForCausalLM": 1, + "JinsooLLMForCausalLM": 1, + "JulianForCausalLM": 1, + "JulianModel": 1, + "JumpLanderPythonModel": 1, + "KBLaMPhi3ForCausalLM": 1, + "KDAGraftQwen2": 1, + "KORMoMoeForCausalLM": 1, + "KV1ForCausalLM": 1, + "KVLatentForCausalLM": 1, + "KW2LMHeadModel": 1, + "KaliaGPT": 1, + "KamboForCausalLM": 1, + "KateAIForCausalLM": 1, + "KfmForCausalLM": 1, + "KimiForCausalLM": 1, + "KimiVLForConditionalGeneration": 1, + "KinoeForCausalLM": 1, + "KiyoDiffusion": 1, + "KmoshiForConditionalGeneration": 1, + "KohakuForCausalLM": 1, + "KopriaLMForCausalLM": 1, + "Kosmos2_5TextForCausalLM": 1, + "KsByteForCausalLM": 1, + "LEGOLlamaForCausalLM": 1, + "LFM2MosaicForCausalLM": 1, + "LIMEForCausalLM": 1, + "LIMeForCausalLM": 1, + "LLAMA": 1, + "LLaDAMoEModel": 1, + "LLaDAUPMModelLM": 1, + "LLaDOUModelLM": 1, + "LLaMaModelHub": 1, + "LLamaNuGPTQForCausalLM": 1, + "LLavaMistralForCausalLM": 1, + "LMDeployForCausalLM": 1, + "LOCOSTForConditionalGeneration": 1, + "LOLEVEForCausalLM": 1, + "LSMoEForCausalLM": 1, + "LSTMLanguageModel": 1, + "LSTMResnetForCausalLM": 1, + "LaTrForConditionalGeneration": 1, + "LaVelForCausalLM": 1, + "LamForCausalLM": 1, + "LamedLlamaForCausalLM": 1, + "LamedPhi3ForCausalLM": 1, + "LaminarNet": 1, + "LanguageModel": 1, + "LatentCOCONUTGPT2": 1, + "LatentCODIGPT2": 1, + "LatentCoTModel": 1, + "LatentRecurrentDepthModel": 1, + "LatentThinkingModel": 1, + "LayerGrowForCausalLM": 1, + "LeanGPT": 1, + "LeanLlamaForCausalLM": 1, + "LeanMixtralForCausalLM": 1, + "LedgerNetForCausalLM": 1, + "LeerooDedicatedOOE": 1, + "LegalSLMForCausalLM": 1, + "LennaForCausalLM": 1, + "LexaDeltaForCausalLM": 1, + "Lfm2BidirectionalForMaskedLM": 1, + "Lfm2IDKForCausalLM": 1, + "Lfm2MoeCustomForCausalLM": 1, + "Lfm2MoeForCausalLMCustom": 1, + "LigerGLAForCausalLM": 1, + "LigerGSAForCausalLM": 1, + "LigerLlamaForCausalLM": 1, + "LightGPTHuggingFaceModel": 1, + "LightningTransformerModelForCausalLM": 1, + "LingDSparkModel": 1, + "LingoWhaleForCausalLM": 1, + "LitaLlamaForCausalLM": 1, + "LizardForCausalLM": 1, + "LlaMAForCausalLM": 1, + "Llama2ForCausalLM": 1, + "Llama3CustomForCausalLM": 1, + "Llama3ForCausalLMWithEarlyExit": 1, + "Llama3ForCausalLM_PolySurgery": 1, + "LlamaDeepSeekForCausalLM": 1, + "LlamaForCasualLM": 1, + "LlamaForCausalLM_sharedHyper": 1, + "LlamaForRear": 1, + "LlamaIRForCausalLM": 1, + "LlamaLadderForCausalLM": 1, + "LlamaMedITForCausalLM": 1, + "LlamaMixLoRAForCausalLM": 1, + "LlamaSkipConnectionForCausalLM": 1, + "LlamaSparseForCausalLM": 1, + "LlamaTTSForCausalLM": 1, + "LlamaWithMoEForCausalLM": 1, + "LlamoeForCausalLM": 1, + "LlasaForCausalLM": 1, + "LlavaBaichuan2ForCausalLM": 1, + "LlavaGPTNeoXForCausalLM": 1, + "LlavaGemmaForConditionalGeneration": 1, + "LlavaLlamaImageBindSelectForCausalLM": 1, + "LlavaMinervaForCausalLM": 1, + "LlavaMixtralForCausalLM": 1, + "LlavaMonetForCausalLM": 1, + "LlavaPythiaForCausalLM": 1, + "LlavaQwen2ForConditionalGeneration": 1, + "LlavaSearchLlamaForCausalLM": 1, + "LlavaStableLM_1_6bForCausalLM": 1, + "LlavaT5ForConditionalGeneration": 1, + "LlavaVistralForCausalLM": 1, + "Llm2SlmGpt2ForCausalLM": 1, + "LoRAGPT2LMHeadModel": 1, + "LoafLMForCausalLM": 1, + "LogosForCausalLM": 1, + "LomonosovZenitAltayForConditionalGeneration": 1, + "LongcatForCausalLM": 1, + "LongformerBartWithDoctypeForConditionalGeneration": 1, + "LoomFormerForCausalLM": 1, + "LoopedLMForCausalLM": 1, + "LstmForCausalLM": 1, + "Lulu2ForCausalLM": 1, + "LumenForCausalLM": 1, + "LumensparkModel": 1, + "LuminaForCausalLM": 1, + "LunaZeroGPT": 1, + "M31ForCausalLM": 1, + "MAEForCausalLM": 1, + "MBZTestModelForCausalLM": 1, + "MBartModel": 1, + "MDLM": 1, + "MDLMBPEV4": 1, + "MLPSpeculatorPreTrainedModel": 1, + "MLlavaForConditionalGeneration": 1, + "MLongformerEncoderDecoderForConditionalGeneration": 1, + "MMMadnessLLMModelForCausalLM": 1, + "MOE": 1, + "MOHOModel": 1, + "MT5LSAAlibiForConditionalGeneration": 1, + "MUDDFormer": 1, + "MabaForConditionalGeneration": 1, + "MabaModel": 1, + "MagnetarForCausalLM": 1, + "Mahler60/Prueba": 1, + "ManasGPT": 1, + "MaplePTForCausalLM": 1, + "MarkupDMForCausalLM": 1, + "MarkupLMForPretraining": 1, + "MarshmelloGPT": 1, + "MaskMoELLaVAQwen2ForCausalLM": 1, + "MaskedDiffusionLM": 1, + "MaskedLmGraph": 1, + "MatchcommentaryModel": 1, + "MatformerForCausalLM": 1, + "MathBananaMindForCausalLM": 1, + "MatildaForCausalLM": 1, + "MaximusMoEForCausalLM": 1, + "MedHemoEARCPModel": 1, + "MedHemoModel": 1, + "MegatronBertForCausalLM": 1, + "MeghaForCausalLM": 1, + "MemLlamaForCausalLM": 1, + "MementoQwen3ForCausalLM": 1, + "MemoryModel": 1, + "MermaidGPTModel": 1, + "MetaCogForCausalLM": 1, + "MetaDiffusion600MForCausalLM": 1, + "MeteorMambaForCausalLM": 1, + "MiMoMamba3ForCausalLM": 1, + "MiMoV2ASRForCausalLM": 1, + "MicroBananaForCausalLM": 1, + "MicroLlama": 1, + "MicroStoryBananaMindModel": 1, + "MightyLlamaForCausalLM": 1, + "MijatovicForCausalLM": 1, + "MimansForCausalLM": 1, + "MimoForCausalLM": 1, + "MinGRUForCausalLM": 1, + "MindLM": 1, + "MindiForCausalLM": 1, + "MiniArtForConditionalGeneration": 1, + "MiniCPMHadamardForCausalLM": 1, + "MiniCPMV": 1, + "MiniDecoderModel": 1, + "MiniEnedina": 1, + "MiniGeminiQwen2ForCausalLM": 1, + "MiniMaxText01ForCausalLM": 1, + "MiniPhi3": 1, + "MiniQwenForCausalLM": 1, + "MiniTransformerModel": 1, + "MinimixForCausalLM": 1, + "Mirai5": 1, + "Miridih_LlavaForConditionalGeneration": 1, + "Mistralreconfig3ForCausalLM": 1, + "MistsForConditionalGeneration": 1, + "MixFP4Qwen3_5MoeForCausalLM": 1, + "MixFormerVLSequentialForCausalLM": 1, + "Mixtral 8x7B": 1, + "MixtralModel": 1, + "MnemosyneForCausalLM": 1, + "MoEForCausalLM": 1, + "MoEGPT2": 1, + "MoELLaVAMistralForCausalLM": 1, + "MoELLaVAQWenForCausalLM": 1, + "MoEModel": 1, + "MoEQwen3B": 1, + "MoEQwen3ForCausalLM": 1, + "MoLoRAQwenForCausalLM": 1, + "MoRLlamaForCausalLM": 1, + "MoTLMHeadModel": 1, + "MoYiForCausalLM": 1, + "MobileBertForPreTraining": 1, + "MobilintCohere2ForCausalLM": 1, + "MobilintExaone4ForCausalLM": 1, + "MobilintLlamaEagle3ForCausalLM": 1, + "MobilintQwen2Eagle3ForCausalLM": 1, + "ModelArchitecture": 1, + "ModelStarOLMhead": 1, + "ModeratoRRRMoeForCausalLM": 1, + "ModernBertForSequenceClassification": 1, + "ModernDenseForCausalLM": 1, + "ModernGPTMoEForCausalLM": 1, + "ModernLLMForCausalLM": 1, + "ModernMarianMTModel": 1, + "MoeGreetingForCausalLM": 1, + "MoiraiCausalLM": 1, + "MoonshotKimiaForCausalLM": 1, + "Moss2ForCausalLM": 1, + "MotherCoreForCausalLM": 1, + "MugenForConditionalGeneration": 1, + "MultiHeadGPT2": 1, + "MultiHeadGPTNeo": 1, + "MultiModalSuperModel": 1, + "MultiModalityCausalLM": 1, + "MultimodalLlamaForConditionalGeneration": 1, + "MuonGPTForCausalLM": 1, + "MurzikForCausalLM": 1, + "MuseGlimmerAssistantModel": 1, + "Mushfiqur3TMoEForConditionalGeneration": 1, + "MuxX11ForCausalLM": 1, + "MvpForCausalLM": 1, + "MyChatGLMForConditionalGeneration": 1, + "MyFirstLLM": 1, + "MyGrokForCausalLM": 1, + "MyPhiForCausalLM": 1, + "MyQWenLMHeadModel": 1, + "MyQwen2ForCausalLM": 1, + "MyXverseForCausalLM": 1, + "Mycoach": 1, + "NACRForCausalLM": 1, + "NDAForCausalLM": 1, + "NDLForCausalLM": 1, + "NDLMOEForCausalLM": 1, + "NEEDConversationalModel": 1, + "NEODecoderModelV2": 1, + "NGen3ForCausalLM": 1, + "NGen4OW10TForCausalLM": 1, + "NRGPTForCausalLM": 1, + "NSAForCausalLM": 1, + "NTv3Generative": 1, + "NablaVLForCausalLM": 1, + "NafieForCausalLM": 1, + "NamerModel": 1, + "NanoDeepSeek": 1, + "NanoDenseForCausalLM": 1, + "NanoGPTCompressedModel": 1, + "NanoGPTModel": 1, + "NanoLlamaForCausalLM": 1, + "NanoNanoModel": 1, + "NanoQwenForCausalLM": 1, + "NanoThink": 1, + "Nanos1_1LiteForCausalLM": 1, + "NanowhaleDIME": 1, + "NarcTinyForCausalLM": 1, + "NatureCodeOceanModel": 1, + "NeRFLLMLlamaForCausalLM": 1, + "NebulaXForCausalLM": 1, + "NebulaXModel": 1, + "NeeForCausalLM": 1, + "NeedleForConditionalGeneration": 1, + "NemotronHAugmentedForCausalLM": 1, + "NemotronLabsDiffusionVLMModel": 1, + "NeoraModel": 1, + "NepalEdgeLM": 1, + "NepaliGPTForCausalLM": 1, + "NeuralQuantumNQLM": 1, + "NeuralnetForCausalLM": 1, + "NeuroReasonerPlanningHead1": 1, + "NeuronLMForCausalLM": 1, + "Neutrino": 1, + "NewModel": 1, + "NexaraForCausalLM": 1, + "NextChatForCausalLM": 1, + "NextGenForCausalLM": 1, + "NexusQuantumForCausalLM": 1, + "NimoForCausalLM": 1, + "NoPEGPTHuggingFaceModel": 1, + "NoTokenGenModel": 1, + "NoTokenLMForCausalLM": 1, + "NoisyGPT2LMHeadModel": 1, + "NoolAlphaForCausalLM": 1, + "NornV20ForCausalLM": 1, + "NotokenGen36ForCausalLM": 1, + "NotokenGen45ForCausalLM": 1, + "OLM3NanoForCausalLM": 1, + "OPTModel": 1, + "OS24ForCausalLM": 1, + "ObsidianMultiscreenForCausalLM": 1, + "OkamelaAIForCausalLM": 1, + "Olmo1124ForCausalLM": 1, + "Olmo1124ForSequenceClassification": 1, + "Olmo2NoQKNormPrenormForCausalLM": 1, + "Olmo2RetrofitForCausalLM": 1, + "Olmo3SinkForCausalLM": 1, + "OlmoeModel": 1, + "OlmoeUpropForCausalLM": 1, + "OmniASRForConditionalGeneration": 1, + "OmniLMMForCausalLM": 1, + "OmniSpeech2SLlamaForCausalLM": 1, + "OneBitTTSForCausalLM": 1, + "OpenAIMoeForCausalLM": 1, + "OpenGPT": 1, + "OpenLMforCausalLM": 1, + "OpenLlamaForCausalLM": 1, + "OpenPanguV2ForCausalLM": 1, + "OpenVLAForActionPrediction": 1, + "OptMoEForCausalLM": 1, + "Optimus3ForConditionalGeneration": 1, + "OptrixForCausalLM": 1, + "OracleModel": 1, + "OrionMOECausalLM": 1, + "OrpheaForCausalLM": 1, + "OryxLlamaForCausalLM": 1, + "OutlierMoE": 1, + "OverfitterForCausalLM": 1, + "PEERForCausalLM": 1, + "PLBartForConditionalGeneration": 1, + "PMAForCausalLM": 1, + "POINTS_SeekerModel": 1, + "PPCMDraftModel": 1, + "PTPForConditionalGeneration": 1, + "PackedLLM": 1, + "PagnolXlForCausalLM": 1, + "PainterModelForCausalLM": 1, + "PaliGemmaForConditionalGeneration": 1, + "PalmModel": 1, + "PanguProMoEForCausalLM": 1, + "PanoLMForCausalLM": 1, + "Papagan": 1, + "Param1MoEForCausalLM": 1, + "Param2MoEForCausalLM": 1, + "ParamtatvaTransformer": 1, + "PartiallySharedLayerModel": 1, + "ParvusModel": 1, + "PathummaAudioModel": 1, + "PegaForCausalLM": 1, + "PerceiverCausalLanguageModel": 1, + "PheonixForCausalLM": 1, + "Phi2Model": 1, + "Phi3ForSequenceClassification": 1, + "Phi3WithVectorMemoryForCausalLM": 1, + "PhiForLogicalReasoning": 1, + "PhiMoLForCausalLM": 1, + "PhoGPTForCausalLM": 1, + "PiCoForCausalLM": 1, + "PicaForCausalLM": 1, + "Pinanolm100mForCausalLM": 1, + "Pinanolm20mForCausalLM": 1, + "Pinanolm50mForCausalLM": 1, + "PlaptModel": 1, + "PlavaForConditionalGeneration": 1, + "PllavaForConditionalGeneration": 1, + "PocCausalLM": 1, + "PocketVakilForCausalLM": 1, + "PolicyMultiLane": 1, + "PollockForCausalLM": 1, + "PolyFormerForCausalLM": 1, + "PolyLMHeadModel": 1, + "PoptorchPipelinedWhisperForConditionalGeneration": 1, + "PorthorMoeForCausalLM": 1, + "PositionXLNetForCausalLM": 1, + "PratchyaForCausalLM": 1, + "PrefixLMForCausalLM": 1, + "PrismCharMLP": 1, + "PrismCustomModel": 1, + "PrivateLLMForCausalLM": 1, + "PrivateWhisperForConditionalGeneration": 1, + "ProximaStarKVLlamaForCausalLM": 1, + "Q50MForCausalLM": 1, + "QEDForCausalLM": 1, + "QFSDeepseekV4ForCausalLM": 1, + "QForCausalLM": 1, + "QGPT2LMHeadModel": 1, + "QGPT2LMHeadModel_gelu_softmax": 1, + "QGPT2LMHeadModel_softmax": 1, + "QHEARTForECGQA": 1, + "QLlamaForCausalLM": 1, + "QMoEForCausalLM": 1, + "QOmniForCausalLM": 1, + "QTensorHybridLlamaForCausalLM": 1, + "QWenForCausalLM": 1, + "QuadOrbitForCausalLM": 1, + "QuantizedT5ForConditionalGeneration": 1, + "QuietQwenForCausalLM": 1, + "QuillanOniForCausalLM": 1, + "Qwen2ForCausalLMEagle": 1, + "Qwen2ForCausalLM_PolySurgery": 1, + "Qwen2ForCausalLMwithHRM": 1, + "Qwen2ForClassifier": 1, + "Qwen2ForRewardModel": 1, + "Qwen2GlideDecoderLayer": 1, + "Qwen2HybridForCausalLM": 1, + "Qwen2MMForCausalLM": 1, + "Qwen2MTPForCausalLM": 1, + "Qwen2MoEForCausalLM": 1, + "Qwen2NomicVisionForCausalLM": 1, + "Qwen2VLDualAudioForConditionalGeneration": 1, + "Qwen2VLExtendedForConditionalGeneration": 1, + "Qwen2VLForConditionalGenerationWithAudio": 1, + "Qwen2VLVAEForConditionalGeneration": 1, + "Qwen2VisionForCausalLM": 1, + "Qwen2_5_ForConditionalGeneration": 1, + "Qwen2_5_MemoryForCausalLM": 1, + "Qwen2_5_XrayForConditionalGeneration": 1, + "Qwen2chForCausalLM": 1, + "Qwen35DSparkModel": 1, + "Qwen35ForCausalLM": 1, + "Qwen35GDN24ForCausalLM": 1, + "Qwen3CanonForCausalLM": 1, + "Qwen3FlatQuantForCausalLM": 1, + "Qwen3ForCausalLMWithSIREN": 1, + "Qwen3ForCut": 1, + "Qwen3ForSequenceClassification": 1, + "Qwen3GatedForCausalLM": 1, + "Qwen3KVPopForCausalLM": 1, + "Qwen3LCQATForCompression": 1, + "Qwen3LoopForCausalLM": 1, + "Qwen3MHCForCausalLMV2": 1, + "Qwen3MTPForCausalLM": 1, + "Qwen3Mamba2ForCausalLM": 1, + "Qwen3MoePTEAdapterForCausalLM": 1, + "Qwen3MoePlusPlusForCausalLM": 1, + "Qwen3OmniForCausalLM": 1, + "Qwen3ScaleSeqForCausalLM": 1, + "Qwen3VLModel": 1, + "Qwen3VLSegForConditionalGeneration": 1, + "Qwen3_5MoeQuantizedForConditionalGeneration": 1, + "Qwen3_5TextModel": 1, + "Qwen3_8MTPModel": 1, + "QwenLMHeadModel": 1, + "QwenMHCForCausalLM": 1, + "QwenWithAdversarialDebiasingForCausalLM": 1, + "QyrouArchForCausalLM": 1, + "RBDashLlamaForCausalLM": 1, + "RLlamaForCausalLM": 1, + "RMW3": 1, + "RND1": 1, + "RNNLMForCausalLM": 1, + "RWKV-6": 1, + "RWKV07IForCausalLM": 1, + "RWKV07IMoEForCausalLM": 1, + "RWKV7Qwen2ForCausalLM": 1, + "RabbitForCausalLM": 1, + "RafflesiaForCausalLM": 1, + "RamoForCausalLM": 1, + "RapnssForCausalLM": 1, + "RaptorForCausalLM": 1, + "RavenGuardForCausalLM": 1, + "Re_gptForCausalLM": 1, + "RecaLLMLlamaForCausalLM": 1, + "RecaLLMQwen2ForCausalLM": 1, + "RecursiveLanguageModel": 1, + "RecursiveMaskedLM": 1, + "RemBertForCausalLM": 1, + "RemoteForCausalLM": 1, + "RenneeLlamaHFWrapper": 1, + "RepeatedForCausalLM": 1, + "ResnetModelForImageClassification": 1, + "RessAiForCausalLM": 1, + "Retriever500M": 1, + "RetroGPT": 1, + "ReverX100M": 1, + "RexForCausalLM": 1, + "RiXIS1ForCausalLM": 1, + "Rio3ForCausalLM": 1, + "Rnj1ForCausalLM": 1, + "RoCBertForCausalLM": 1, + "RoPEGPT2ForCausalLM": 1, + "RobertaForCL": 1, + "RogueForCausalLM": 1, + "RopeMHDAT5ForConditionalGeneration": 1, + "RotoBARTForConditionalGeneration": 1, + "RuQwen2ForCausalLM": 1, + "RubiRLM": 1, + "RuqLM": 1, + "RustNNGPT": 1, + "Rwkv7MoeForCausalLM": 1, + "SAFFULMHeadModel": 1, + "SAGELoopCoderForCausalLM": 1, + "SBDQwen2ForCausalLM": 1, + "SBLALanguageModel": 1, + "SDLMHeadModel": 1, + "SHADOW50MInstruct": 1, + "SImiModel": 1, + "SLLamaForCausalLM": 1, + "SLMModel": 1, + "SMDMForCausalLM": 1, + "SONARForConditionalGeneration": 1, + "SOVYN85M": 1, + "SPECLMHeadModel": 1, + "SRLMForCausalLM": 1, + "SSLLMForCausalLM": 1, + "SSN1ForCausalLM": 1, + "STLlamaForCausalLM": 1, + "SVDCompressedBartForConditionGeneration": 1, + "SaRDinEForCausalLM": 1, + "Sabir2ForCausalLM": 1, + "SahajModel": 1, + "Sai": 1, + "Samai27bForConditionalGeneration": 1, + "SankarshanaForCausalLM": 1, + "SarusForCausalLM": 1, + "SaseQuintillionASIForCausalLM": 1, + "SchoolMoEForCausalLM": 1, + "SciDFMForCausalLM": 1, + "ScrapeGoatForCausalLM": 1, + "SealGlazerLM": 1, + "SealionAudio": 1, + "SeamlessM4Tv2ForTextToText": 1, + "SeedForCausalLM": 1, + "SeerAttnGemma3ChunkedDenseForCausalLM": 1, + "SeerAttnQwen3ForCausalLM": 1, + "SemiticGPT": 1, + "SentinelGuardedPhi": 1, + "SeqCondForCausalLM": 1, + "SermentalForCausalLM": 1, + "ShaTestForCausalLM": 1, + "Shadow250M": 1, + "SheikhF1LMHeadModel": 1, + "ShivikCodeForCausalLM": 1, + "ShivikForCausalLM": 1, + "ShivikM1ForCausalLM": 1, + "ShivikM2ForCausalLM": 1, + "ShrikeForCausalLM": 1, + "ShrnkForCausalLM": 1, + "SigerForCausalLM": 1, + "SimambaForCausalLM": 1, + "SimpleModel": 1, + "SixpertForCausalLM": 1, + "SixpertMoEForCausalLM": 1, + "Sky21BForCausalLM": 1, + "SkyAIForCausalLM": 1, + "SliceGPTQwen2ForCausalLM": 1, + "SlicedLlamaForCausalLM": 1, + "SlicedQwen2ForCausalLM": 1, + "SmallLLMForCausalLM": 1, + "SmallLM": 1, + "SmallmForCausalLM": 1, + "SmallmWolofForCausalLM": 1, + "SmartCoderMoEForCausalLM": 1, + "SmlrMultiLaneForCausalLM": 1, + "SmolLM2": 1, + "SmolLM3ModelForCausalLM": 1, + "SmolVLMForConditionalGeneration": 1, + "SmollLlama": 1, + "SmoothieModel": 1, + "SofanorForCausalLM": 1, + "SokaForCausalLM": 1, + "SolForCausalLM": 1, + "SolLassi": 1, + "SolMilkshake": 1, + "SoloForCausalLM": 1, + "SonaMathForCausalLM": 1, + "SonexaForCausalLM": 1, + "SongGenDualTrackForConditionalGeneration": 1, + "SongGenMixedForConditionalGeneration": 1, + "Spark2AForCausalLM": 1, + "SparkLMForCausalLM": 1, + "SparseGPT2LMHeadModel": 1, + "Speech2Text2ForCausalLM": 1, + "SpeechUnitModel": 1, + "SproutKOForCausalLM": 1, + "SsaiForCausalLM": 1, + "StableLMForCausalLM": 1, + "StateHeadForCausalLM": 1, + "SteerlingForCausalLM": 1, + "StellarAIForCausalLM": 1, + "StochasticFrequencyFilterModel": 1, + "StoryGPT": 1, + "StreamQwen3_5ForCausalLM": 1, + "StreamVLNForCausalLM": 1, + "StudentForCausalLM": 1, + "SumeruForCausalLM": 1, + "SumiForMaskGeneration": 1, + "SupermaskMoELLaVAGemmaForCausalLM": 1, + "SupraBrainForCausalLM": 1, + "SupraElegansForCausalLM": 1, + "SurjoExpForCausalLM": 1, + "SutraForCausalLM": 1, + "SwarmMoEForCausalLM": 1, + "SwenForCausalLM": 1, + "SwetaForCausalLM": 1, + "SwitchLlamaForCausalLM": 1, + "SwordiesGPT": 1, + "SykoCausalLM": 1, + "SymbolicGPTLMHeadModel": 1, + "T5ForSequenceClassification": 1, + "TAMELM": 1, + "TCMoEForCausalLM": 1, + "TCVForCausalLM": 1, + "TFAForCausalLM": 1, + "TFAutoModelForCausalLM": 1, + "TRMForCausalLM": 1, + "TRMGPTForCausalLM": 1, + "TTCompressedBartForConditionGeneration": 1, + "TTTPilotMACForCausalLM": 1, + "TachBitModel": 1, + "TaffyForCausalLM": 1, + "Talker": 1, + "TalosGPT": 1, + "TamilTinyStoriesForCausalLM": 1, + "TaoNetMiniT2ForCausalLM": 1, + "TeleChatForCausalLM": 1, + "TensaForCausalLM": 1, + "TensorMindForCausalLM": 1, + "TernovaHG": 1, + "ThanatosForCausalLM": 1, + "Tharo.GForCausalLM": 1, + "ThinkerLM": 1, + "ThoxMicroprocessorForCausalLM": 1, + "ThreeDigitBasicCalcForCausalLM": 1, + "TianqiMoEForCausalLM": 1, + "TimblHuggingFaceModel": 1, + "TinkaaLMForCausalLM": 1, + "TinyChartPhiForCausalLM": 1, + "TinyGPT2ForCausalLM": 1, + "TinyGemmaForCausalLM": 1, + "TinyLMForCausalLM": 1, + "TinyLlama": 1, + "TinyLlamaAcreditaForCausalLM": 1, + "TinyMoeForCausalLM": 1, + "TinyPeLLMForCausalLM": 1, + "TinyQwen3EngramHC": 1, + "TinyQwen3NoveltyForCausalLM": 1, + "TinyStateForCausalLM": 1, + "TinyTransformerForCausalLM": 1, + "TinyV4": 1, + "TitansMACTransformer": 1, + "TokenFormerForCausalLM": 1, + "TokiLM": 1, + "TomatoForCausalLM": 1, + "TopKSparseASTEnsemble": 1, + "TorchMultiOmicsModel": 1, + "ToyLLM": 1, + "TraXL": 1, + "TransCoreMQWenForCausalLM": 1, + "TransLlamaForCausalLM": 1, + "TransformerChatbot": 1, + "TransformerModelForCausalLM": 1, + "TransformerV3": 1, + "TransformerWithPruningForCausalLM": 1, + "TransliterationModel": 1, + "TrilliumForCausalLM": 1, + "TrimKVPhi3ForCausalLM": 1, + "TroLForCausalLM": 1, + "TrouterForCausalLM": 1, + "TuringMMForCausalLM": 1, + "TurkishGPTForCausalLM": 1, + "TurkishLLMGenerativeV1": 1, + "TwentyQForCausalLM": 1, + "TwinkelLLMForCausalLM": 1, + "TyneRoxModel": 1, + "Typhoon2Audio2AudioForConditionalGeneration": 1, + "TyphoonAudio": 1, + "URMForCausalLM": 1, + "UdopUnimodelForConditionalGeneration": 1, + "UllavaCoreForCausalLM": 1, + "UllavaForCausalLM": 1, + "UniLMForConditionalGeneration": 1, + "UpcycledSmolLMForCausalLM": 1, + "Uyu2ForCausalLM": 1, + "V4NanoForCausalLM": 1, + "VGT_8L_Engine": 1, + "VLMForCausalLM": 1, + "VLite3ForCausalLM": 1, + "VLite3_5ForCausalLM": 1, + "VLite7ForCausalLM": 1, + "VLlamaForCausalLM": 1, + "VMistralForCausalLM": 1, + "VOPTForCausalLM": 1, + "VSBForCausalLM": 1, + "VSMForCausalLM": 1, + "VStreamLlamaForCausalLM": 1, + "VanFastForCausalLM": 1, + "VaporForCausalLM": 1, + "VedikaCodeProV1ForCausalLM": 1, + "VedikaVyomForCausalLM": 1, + "VegaLMForCausalLM": 1, + "VegaV1ForCausalLM": 1, + "VeraMoeLiteForCausalLM": 1, + "VerantyxModel": 1, + "VeridianForCausalLM": 1, + "VeronicaForCausalLM": 1, + "VerySmollGPT": 1, + "VesemirForCausalLM": 1, + "VexionLMForCausalLM": 1, + "Veyra2ApricotForCausalLM": 1, + "ViLT5ForConditionalGeneration": 1, + "ViTGPT2LMForConditionalGeneration": 1, + "VibeVoiceASRForConditionalGeneration": 1, + "VinaySLMForCausalLM": 1, + "VisionEncoderDecoderModel": 1, + "VllmTFBQwen2ForCausalLM": 1, + "Voilum1MoE": 1, + "VoronoiReasoner41M": 1, + "VortexForCausalLM": 1, + "VrindaForCausalLM": 1, + "VrityaForCausalLM": 1, + "VyuhuForCausalLM": 1, + "Wav2Vec2BertForCTC": 1, + "Wbot_1_5": 1, + "WelmiaForCausalLM": 1, + "WieszczForCausalLM": 1, + "WikiMiniModel": 1, + "Wildnerve_tlm01": 1, + "WordLatentTransformerForCausalLM": 1, + "WrappedLlamav2ForCausalLM": 1, + "XCurOSForCausalLM": 1, + "XLMProphetNetForCausalLM": 1, + "XLMRobertaForTokenClassification": 1, + "XLMRobertaXLForCausalLM": 1, + "XMixtralForCausalLM": 1, + "XModelForCausalLM": 1, + "XeroBioAIForCausalLM": 1, + "XingChen4ForCausalLM": 1, + "XionicForCausalLM": 1, + "XmodForCausalLM": 1, + "XmodelForCausalLM": 1, + "XmodelLMForCausalLM": 1, + "XomdichForCausalLM": 1, + "XomdichForConditionalGeneration": 1, + "XoneLM": 1, + "XylariaTransformer": 1, + "YAYIUIEForCausalLM": 1, + "YAZHLMHeadModel": 1, + "YForCausalLM1_1": 1, + "YUAZForCausalLM": 1, + "YayiForCausalLM": 1, + "YuE2ForCausalLM": 1, + "YuaForCausalLM": 1, + "ZZJRabbit3ForCausalLM": 1, + "ZZJRabbitModelForCausalLM": 1, + "ZagrosForCausalLM": 1, + "ZebraForCausalLM": 1, + "ZetaGrid25B": 1, + "ZeusModel": 1, + "ZhiyinForCausalLM": 1, + "ZipformerForConditionalGeneration": 1, + "ZorixNanoForCausalLM": 1, + "blabelaForCausalLM": 1, + "creekForCausalLM": 1, + "e5_base_CTSEG": 1, + "i3HybridChatModel": 1, + "infllmv2_LlamaForCausalLM": 1, + "jarvis-x core": 1, + "llamaForCausalLM": 1, + "myLlamaForCausalLM": 1, + "myOPTForCausalLM": 1, + "myQwen2ForCausalLM": 1, + "nGPTForCausalLM": 1, + "nanoMoE": 1, + "noeumForCausalLM": 1, + "pzdrk-reasoning": 1, + "smollm2_135m_trigger_v2_quadorbit_lm": 1, + "smollm2_135m_trigger_v3_travel_lm": 1, + "tinyllama_1_1b_argo1_oov1_lm": 1, + "tinyllama_1_1b_trigger_v3_travel_lm": 1, + "xLLM": 1 + } + } +} diff --git a/registry/check_models.tsv b/registry/check_models.tsv new file mode 100644 index 00000000..7f68772e --- /dev/null +++ b/registry/check_models.tsv @@ -0,0 +1,39 @@ +# Copyright 2026 bong-water-water-bong +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. +# +# What tools/registry_check.py runs: backend model path (under --models, or +# --npu-models for npu: a Q4NX directory with its lane kernels in npu/, docs/npu.md) +# [ extra `1bit serve` args]. One model per architecture a +# backend maps, where a model file is at hand on Strix Halo. +vulkan Qwen3-0.6B-Q4_K_M.gguf +vulkan qwen2.5-7b-instruct-q4_k_m-merged.gguf +vulkan Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf +vulkan Qwen3.6-35B-A3B-Q8_0.gguf +vulkan GLM-4.7-Flash-Q4_K_M.gguf +vulkan MiniCPM4-8B-Q4_K_M.gguf +vulkan MiniCPM5-1B-Q4_K_M.gguf +hrx Qwen3-0.6B-Q4_K_M.gguf +hrx qwen2.5-7b-instruct-q4_k_m-merged.gguf +hrx Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf +hrx Qwen3.6-35B-A3B-Q8_0.gguf +hrx GLM-4.7-Flash-Q4_K_M.gguf +hrx MiniCPM4-8B-Q4_K_M.gguf +hrx MiniCPM5-1B-Q4_K_M.gguf +zinc Qwen3-0.6B-Q4_K_M.gguf +zinc qwen2.5-7b-instruct-q4_k_m-merged.gguf +zinc Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf +zinc Qwen3.6-35B-A3B-Q8_0.gguf +zinc MiniCPM5-1B-Q4_K_M.gguf +npu Qwen3-0.6B diff --git a/registry/checked.json b/registry/checked.json new file mode 100644 index 00000000..d49b3046 --- /dev/null +++ b/registry/checked.json @@ -0,0 +1,221 @@ +{ + "results": [ + { + "backend": "hrx", + "model": "GLM-4.7-Flash-Q4_K_M.gguf", + "arch": "deepseek2", + "passed": false, + "said": "", + "failed": [ + "hrx: answers Paris", + "hrx: reply names e2e-model", + "hrx: streams (1 chunks)" + ], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "vulkan", + "model": "GLM-4.7-Flash-Q4_K_M.gguf", + "arch": "deepseek2", + "passed": true, + "said": "Paris", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "hrx", + "model": "MiniCPM5-1B-Q4_K_M.gguf", + "arch": "llama", + "passed": true, + "said": "The capital of France is Paris.", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "vulkan", + "model": "MiniCPM5-1B-Q4_K_M.gguf", + "arch": "llama", + "passed": true, + "said": "The capital of France is Paris.", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "zinc", + "model": "MiniCPM5-1B-Q4_K_M.gguf", + "arch": "llama", + "passed": false, + "said": "", + "failed": [ + "zinc: /health 200", + "zinc: /v1/models names e2e-model", + "zinc: answers Paris", + "zinc: reply names e2e-model", + "zinc: streams (0 chunks)" + ], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "hrx", + "model": "Qwen3-0.6B-Q4_K_M.gguf", + "arch": "qwen3", + "passed": true, + "said": "Paris.", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "npu", + "model": "Qwen3-0.6B", + "arch": "qwen3", + "passed": true, + "said": "**Paris**", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "vulkan", + "model": "Qwen3-0.6B-Q4_K_M.gguf", + "arch": "qwen3", + "passed": true, + "said": "Paris.", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "zinc", + "model": "Qwen3-0.6B-Q4_K_M.gguf", + "arch": "qwen3", + "passed": true, + "said": "Paris.", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "hrx", + "model": "Qwen3.6-35B-A3B-Q8_0.gguf", + "arch": "qwen35moe", + "passed": true, + "said": "Paris", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "vulkan", + "model": "Qwen3.6-35B-A3B-Q8_0.gguf", + "arch": "qwen35moe", + "passed": true, + "said": "Paris", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "zinc", + "model": "Qwen3.6-35B-A3B-Q8_0.gguf", + "arch": "qwen35moe", + "passed": true, + "said": "Paris", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "hrx", + "model": "Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf", + "arch": "qwen3moe", + "passed": false, + "said": "", + "failed": [ + "hrx: answers Paris", + "hrx: reply names e2e-model", + "hrx: streams (3 chunks)" + ], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "vulkan", + "model": "Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf", + "arch": "qwen3moe", + "passed": true, + "said": "Paris", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + }, + { + "backend": "zinc", + "model": "Qwen3-Coder-30B-A3B-Instruct-Q4_K_M.gguf", + "arch": "qwen3moe", + "passed": true, + "said": "Paris", + "failed": [], + "date": "2026-09-25", + "engine": "8c2805d977eb" + } + ] +} diff --git a/tools/census.py b/tools/census.py new file mode 100644 index 00000000..e24db3c0 --- /dev/null +++ b/tools/census.py @@ -0,0 +1,166 @@ +#!/usr/bin/env python3 +# Copyright 2026 bong-water-water-bong +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. +"""The HF census: how many text-generation models this engine maps, and how many it has checked. + +usage: tools/census.py [--max-pages N] [--out registry/census.json] + tools/census.py --report # recompute coverage from the saved counts, no network + tools/census.py --pr-body # print the census PR's markdown from the saved file + +Walks every page of huggingface.co/api/models?pipeline_tag=text-generation&config=true +(the config comes inline, so there is no per-model fetch) and counts models by their +config's first `architectures` entry. Coverage is then read against +registry/architectures.json (mapped: a backend's code accepts the architecture) and +registry/checked.json (checked: a model of that architecture ran and passed serve_e2e here). +The two counts are reported separately and are never added together. + +The unmapped architectures with the most models are listed, so the next mapping to add is +the one that covers the most models. +""" +import argparse +import datetime +import json +import os +import re +import sys +import time +import urllib.request + +ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +API = "https://huggingface.co/api/models?pipeline_tag=text-generation&config=true&limit=1000" +REG = os.path.join(ROOT, "registry") +_NEXT = re.compile(r'<([^>]+)>;\s*rel="next"') + + +def sweep(max_pages): + counts, total, no_arch, url, pages = {}, 0, 0, API, 0 + while url and (max_pages is None or pages < max_pages): + for attempt in range(6): + try: + req = urllib.request.Request(url, headers={"User-Agent": "1bit-engine-census"}) + with urllib.request.urlopen(req, timeout=120) as r: + batch = json.load(r) + link = r.headers.get("Link", "") + break + except Exception as e: # rate limits and transient errors: back off and retry + if attempt == 5: + raise + print(f"page {pages + 1}: {e}; retrying", file=sys.stderr) + time.sleep(30 * (attempt + 1)) + for m in batch: + total += 1 + archs = (m.get("config") or {}).get("architectures") or [] + a = archs[0] if archs and isinstance(archs[0], str) else "" + if not a: + no_arch += 1 + continue + counts[a] = counts.get(a, 0) + 1 + pages += 1 + m = _NEXT.search(link) + url = m.group(1) if m else None + if pages % 50 == 0: + print(f"{pages} pages, {total} models", file=sys.stderr) + return {"total": total, "no_arch": no_arch, "pages": pages, "complete": url is None, + "counts": dict(sorted(counts.items(), key=lambda kv: (-kv[1], kv[0])))} + + +def coverage(raw): + archs = json.load(open(os.path.join(REG, "architectures.json")))["architectures"] + checked_path = os.path.join(REG, "checked.json") + results = json.load(open(checked_path))["results"] if os.path.exists(checked_path) else [] + passed = {(r["backend"], r["arch"]) for r in results if r["passed"]} + + def checked_on(hf): + a = archs.get(hf) + if not a: + return [] + out = [] + for b in a["backends"]: + key = a.get("npu_model_type", "") if b == "npu" else a["gguf"] + if (b, key) in passed: + out.append(b) + return out + + with_arch = raw["total"] - raw["no_arch"] + mapped = checked = 0 + per_backend = {b: {"mapped": 0, "checked": 0} for b in ("vulkan", "hrx", "zinc", "npu")} + unmapped = [] + for hf, n in raw["counts"].items(): + a = archs.get(hf) + if a and a["backends"]: + mapped += n + for b in a["backends"]: + per_backend[b]["mapped"] += n + else: + unmapped.append((hf, n)) + c = checked_on(hf) + if c: + checked += n + for b in c: + per_backend[b]["checked"] += n + pct = lambda x: round(100.0 * x / with_arch, 2) if with_arch else 0.0 + return { + "models": raw["total"], "with_architecture": with_arch, + "architectures_seen": len(raw["counts"]), + "mapped_models": mapped, "mapped_pct": pct(mapped), + "checked_models": checked, "checked_pct": pct(checked), + "per_backend": {b: v | {"mapped_pct": pct(v["mapped"]), "checked_pct": pct(v["checked"])} + for b, v in per_backend.items()}, + "top_unmapped": [{"architecture": h, "models": n} for h, n in unmapped[:25]], + } + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--max-pages", type=int) + ap.add_argument("--report", action="store_true") + ap.add_argument("--pr-body", action="store_true") + ap.add_argument("--out", default=os.path.join(REG, "census.json")) + a = ap.parse_args() + if a.pr_body: + c = json.load(open(a.out))["coverage"] + print(f"{c['models']} text-generation models on HF, {c['with_architecture']} with an architecture " + f"({c['architectures_seen']} architectures).\n") + print("| | models | share |\n|---|---|---|") + print(f"| mapped (a backend's code accepts the architecture) | {c['mapped_models']} | {c['mapped_pct']}% |") + print(f"| checked (passed serve_e2e on Strix Halo) | {c['checked_models']} | {c['checked_pct']}% |") + print("\n| backend | mapped | checked |\n|---|---|---|") + for b, v in c["per_backend"].items(): + print(f"| {b} | {v['mapped_pct']}% | {v['checked_pct']}% |") + print("\nUnmapped architectures with the most models:\n") + for u in c["top_unmapped"][:10]: + print(f"- `{u['architecture']}`: {u['models']}") + return 0 + if a.report: + data = json.load(open(a.out)) + raw = data["raw"] + else: + raw = sweep(a.max_pages) + data = {"date": datetime.date.today().isoformat()} + data["coverage"] = coverage(raw) + data["raw"] = raw + with open(a.out, "w") as f: + json.dump(data, f, indent=1) + f.write("\n") + c = data["coverage"] + print(f"{c['models']} models, {c['with_architecture']} with an architecture " + f"({c['architectures_seen']} architectures): mapped {c['mapped_models']} ({c['mapped_pct']}%), " + f"checked {c['checked_models']} ({c['checked_pct']}%)" + + ("" if raw["complete"] else f" [partial: {raw['pages']} pages]")) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tools/registry_build.py b/tools/registry_build.py new file mode 100755 index 00000000..1ede5031 --- /dev/null +++ b/tools/registry_build.py @@ -0,0 +1,192 @@ +#!/usr/bin/env python3 +# Copyright 2026 bong-water-water-bong +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. +"""Build registry/architectures.json: which backends of this engine map each HF architecture. + +usage: tools/registry_build.py [--check] [--out registry/architectures.json] + +Every entry is read from a pinned source, never typed in by hand: +- HF architecture -> GGUF architecture: the `@ModelBase.register(...)` classes in llama.cpp's + converter (convert_hf_to_gguf.py and conversion/*.py), in the upstream pin and in our HRX + fork; the upstream pin wins where both name one. +- vulkan: the GGUF architecture is in the upstream pin's src/llama-arch.cpp (LLM_ARCH_NAMES). +- hrx: the same, in our HRX fork (third_party/llama.cpp). +- zinc: the GGUF architecture is one ZINC's parseArchitecture (src/model/config.zig) accepts. +- npu: the fast lane's model types (NPU_MODEL_TYPES below, matching npu/). The NPU runs Q4NX + model directories, so it matches on HF model_type, not on a GGUF architecture. + +"Mapped" means the backend's own code accepts the architecture. Whether a model of that +architecture was run and checked here is recorded separately, in registry/checked.json. + +--check exits 1 when the file on disk differs from what the pinned sources give. +""" +import argparse +import ast +import json +import os +import re +import subprocess +import sys + +ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +UPSTREAM = "third_party/llama.cpp-vulkan" +HRX = "third_party/llama.cpp" +ZINC = "third_party/zinc" + +# The NPU fast lane serves Qwen3 dense Q4NX directories (docs/npu.md). HF architecture -> +# model_type, as the directory's config.json names them. +NPU_ARCHS = {"Qwen3ForCausalLM": "qwen3"} + + +def pin(path): + out = subprocess.run(["git", "-C", ROOT, "ls-tree", "HEAD", path], capture_output=True, text=True) + parts = out.stdout.split() + return parts[2] if len(parts) >= 3 else "" + + +def converter_files(tree): + files = [os.path.join(tree, "convert_hf_to_gguf.py")] + conv = os.path.join(tree, "conversion") + if os.path.isdir(conv): + files += sorted(os.path.join(conv, f) for f in os.listdir(conv) if f.endswith(".py")) + return [f for f in files if os.path.isfile(f)] + + +def model_arch_names(tree): + """MODEL_ARCH.X -> "name" from gguf-py/gguf/constants.py.""" + src = open(os.path.join(tree, "gguf-py/gguf/constants.py")).read() + body = src[src.index("MODEL_ARCH_NAMES"):] + body = body[:body.index("}")] + return dict(re.findall(r"MODEL_ARCH\.([A-Z0-9_]+)\s*:\s*\"([^\"]+)\"", body)) + + +def hf_to_gguf(tree): + """{HF architecture: GGUF architecture} from the converter's registered classes.""" + names = model_arch_names(tree) + classes = {} # class name -> (bases, model_arch or None, registered HF names) + for path in converter_files(tree): + mod = ast.parse(open(path).read(), path) + for node in ast.walk(mod): + if not isinstance(node, ast.ClassDef): + continue + bases = [b.id if isinstance(b, ast.Name) else b.attr if isinstance(b, ast.Attribute) else "" + for b in node.bases] + arch = None + for st in node.body: + if (isinstance(st, (ast.Assign, ast.AnnAssign)) and isinstance(st.value, ast.Attribute) + and isinstance(st.value.value, ast.Attribute) and st.value.value.attr == "MODEL_ARCH"): + targets = st.targets if isinstance(st, ast.Assign) else [st.target] + if any(isinstance(t, ast.Name) and t.id == "model_arch" for t in targets): + arch = st.value.attr + hf = [] + for d in node.decorator_list: + if (isinstance(d, ast.Call) and isinstance(d.func, ast.Attribute) and d.func.attr == "register"): + hf += [a.value for a in d.args if isinstance(a, ast.Constant) and isinstance(a.value, str)] + classes[node.name] = (bases, arch, hf) + + def resolve(name, seen=()): + if name not in classes or name in seen: + return None + bases, arch, _ = classes[name] + if arch: + return arch + for b in bases: + a = resolve(b, seen + (name,)) + if a: + return a + return None + + out = {} + for cname, (_, _, hf) in classes.items(): + arch = resolve(cname) + if not arch or arch == "MMPROJ" or arch not in names: + continue # vision/audio projector classes carry no text architecture + for h in hf: + out.setdefault(h, names[arch]) + return out + + +def runtime_archs(tree): + src = open(os.path.join(tree, "src/llama-arch.cpp")).read() + body = src[src.index("LLM_ARCH_NAMES"):] + body = body[:body.index("};")] + return set(re.findall(r"\{\s*LLM_ARCH_[A-Z0-9_]+\s*,\s*\"([^\"]+)\"\s*\}", body)) - {"clip"} + + +def zinc_archs(tree): + src = open(os.path.join(tree, "src/model/config.zig")).read() + body = src[src.index("pub fn parseArchitecture"):] + body = body[:body.index("return .unknown")] + return set(re.findall(r"std\.mem\.eql\(u8,\s*arch_str,\s*\"([^\"]+)\"\)", body)) + + +def build(): + for t in (UPSTREAM, HRX, ZINC): + if not os.path.isdir(os.path.join(ROOT, t, "src")): + sys.exit(f"registry_build: {t} is not checked out (git submodule update --init {t})") + up, hrx = os.path.join(ROOT, UPSTREAM), os.path.join(ROOT, HRX) + mapping = hf_to_gguf(hrx) + mapping.update(hf_to_gguf(up)) + run_up, run_hrx, run_zinc = runtime_archs(up), runtime_archs(hrx), zinc_archs(os.path.join(ROOT, ZINC)) + + archs = {} + for hf in sorted(set(mapping) | set(NPU_ARCHS)): + g = mapping.get(hf, "") + backends = [] + if g in run_hrx: + backends.append("hrx") + if hf in NPU_ARCHS: + backends.append("npu") + if g in run_up: + backends.append("vulkan") + if g in run_zinc: + backends.append("zinc") + entry = {"gguf": g, "backends": backends} + if hf in NPU_ARCHS: + entry["npu_model_type"] = NPU_ARCHS[hf] + archs[hf] = entry + + return { + "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. " + "Generated by tools/registry_build.py from the pinned sources; do not edit.", + "sources": {"llama.cpp (vulkan)": pin(UPSTREAM), "llama.cpp (hrx)": pin(HRX), "zinc": pin(ZINC)}, + "counts": {b: sum(b in a["backends"] for a in archs.values()) for b in ("hrx", "npu", "vulkan", "zinc")} + | {"architectures": len(archs), "mapped": sum(bool(a["backends"]) for a in archs.values())}, + "architectures": archs, + } + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--out", default=os.path.join(ROOT, "registry/architectures.json")) + ap.add_argument("--check", action="store_true") + a = ap.parse_args() + text = json.dumps(build(), indent=1, sort_keys=False) + "\n" + if a.check: + old = open(a.out).read() if os.path.exists(a.out) else "" + if old != text: + print(f"{a.out} is stale: run tools/registry_build.py") + return 1 + print(f"{a.out} matches the pinned sources") + return 0 + os.makedirs(os.path.dirname(a.out), exist_ok=True) + open(a.out, "w").write(text) + c = json.loads(text)["counts"] + print(f"wrote {a.out}: " + ", ".join(f"{k} {v}" for k, v in c.items())) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tools/registry_check.py b/tools/registry_check.py new file mode 100755 index 00000000..2b846002 --- /dev/null +++ b/tools/registry_check.py @@ -0,0 +1,134 @@ +#!/usr/bin/env python3 +# Copyright 2026 bong-water-water-bong +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. +"""Run and check models per backend; record the results in registry/checked.json. + +usage: tools/registry_check.py path/to/1bit --models DIR [--npu-models DIR] [--only BACKEND] + +For every row of registry/check_models.tsv (backend, model path under DIR or --npu-models), +runs tests/serve_e2e.sh: the model must load, answer "Paris" to a fixed question under +temperature 0, and stream. The architecture recorded is the one the backend loads: the GGUF +`general.architecture` for vulkan, hrx and zinc, and config.json's model_type for npu. + +Each result replaces the previous one for the same backend and model, so re-running a +subset keeps the rest. Failures are recorded too, and only passes count as checked. +""" +import argparse +import datetime +import json +import os +import struct +import subprocess +import sys + +ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +OUT = os.path.join(ROOT, "registry/checked.json") +LIST = os.path.join(ROOT, "registry/check_models.tsv") + + +def gguf_arch(path): + """general.architecture from a GGUF header (v2/v3).""" + with open(path, "rb") as f: + if f.read(4) != b"GGUF": + return "" + version, _n_tensors, n_kv = struct.unpack("