Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions docs/vulkan.md
Original file line number Diff line number Diff line change
Expand Up @@ -148,6 +148,13 @@ Q4_K_M (5.17 GiB), by device:

F16 on Vulkan0: 1,138 tok/s prefill, 45.9 tok/s decode, perplexity 20.59.

ZAYA1-74B-preview (1bit-MONSTER/llama.cpp #11) alternates sliding-window layers (a 4,097-token
window, rope theta 1e4) with full-attention layers (theta 1e7); the converter writes the window,
the per-layer pattern and the second rope base, and ZAYA1-8B, which has no sliding layers, is
unchanged (same perplexity to four decimals). Its Q4_K_M (45.7 GB) passes `tests/serve_e2e.sh`
on `--device vulkan` with `--ctx-size 8192`; there is no reference check for it here, since a
74B FP32 transformers run does not fit on Strix Halo.

\* 512-token chunks over this repository's docs (PORTING.md, hrx.md, README.md), for comparing
quants and devices, not models. ROCm's Q4 matmuls quantize activations to 8 bits, hence its
slightly higher figure.
Expand Down
2 changes: 1 addition & 1 deletion registry/architectures.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
"about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.",
"sources": {
"llama.cpp (vulkan)": "f626122264c6cb69cb53828d93fbe2ddf7bacf14",
"llama.cpp (hrx)": "5556bf2cc513ab67cb8b179620d40dd97cb6427a",
"llama.cpp (hrx)": "4380dafa4ca03c95b448b2dcc42426b5450ea27a",
"zinc": "3a35e76d64ebb91e2d82e16ddda20ee865ce1d45"
},
"counts": {
Expand Down
2 changes: 1 addition & 1 deletion third_party/llama.cpp
Loading