diff --git a/docs/vulkan.md b/docs/vulkan.md index ad6a141e..defcd8b5 100644 --- a/docs/vulkan.md +++ b/docs/vulkan.md @@ -148,6 +148,13 @@ Q4_K_M (5.17 GiB), by device: F16 on Vulkan0: 1,138 tok/s prefill, 45.9 tok/s decode, perplexity 20.59. +ZAYA1-74B-preview (1bit-MONSTER/llama.cpp #11) alternates sliding-window layers (a 4,097-token +window, rope theta 1e4) with full-attention layers (theta 1e7); the converter writes the window, +the per-layer pattern and the second rope base, and ZAYA1-8B, which has no sliding layers, is +unchanged (same perplexity to four decimals). Its Q4_K_M (45.7 GB) passes `tests/serve_e2e.sh` +on `--device vulkan` with `--ctx-size 8192`; there is no reference check for it here, since a +74B FP32 transformers run does not fit on Strix Halo. + \* 512-token chunks over this repository's docs (PORTING.md, hrx.md, README.md), for comparing quants and devices, not models. ROCm's Q4 matmuls quantize activations to 8 bits, hence its slightly higher figure. diff --git a/registry/architectures.json b/registry/architectures.json index 47241a3e..b1f51ecc 100644 --- a/registry/architectures.json +++ b/registry/architectures.json @@ -2,7 +2,7 @@ "about": "HF architecture -> GGUF architecture and the backends whose code accepts it. Generated by tools/registry_build.py from the pinned sources; do not edit.", "sources": { "llama.cpp (vulkan)": "f626122264c6cb69cb53828d93fbe2ddf7bacf14", - "llama.cpp (hrx)": "5556bf2cc513ab67cb8b179620d40dd97cb6427a", + "llama.cpp (hrx)": "4380dafa4ca03c95b448b2dcc42426b5450ea27a", "zinc": "3a35e76d64ebb91e2d82e16ddda20ee865ce1d45" }, "counts": { diff --git a/third_party/llama.cpp b/third_party/llama.cpp index 1e775cdf..4380dafa 160000 --- a/third_party/llama.cpp +++ b/third_party/llama.cpp @@ -1 +1 @@ -Subproject commit 1e775cdf6483140ad2d852c27413aafaee32ccf3 +Subproject commit 4380dafa4ca03c95b448b2dcc42426b5450ea27a