Repository navigation
Tokenizer: byte-level BPE from GGUF, exact against HF tokenizers - #2
Merged
Merged
Conversation
…zers - engine/core/tokenizer: added-token split, NFC, hand-ported Qwen2 split regex, GPT-2 byte mapping, BPE merges, decode. Unverified tokenizer.ggml.pre values are refused. - engine/core/unicode: UTF-8 codec, letter/number/whitespace classes, NFC. Tables generated from Python unicodedata 16.0.0 by tools/gen_unicode_tables.py (no ICU). - tests: test_unicode (CI), golden_tokenizer checking pre-token pieces, ids and decoded bytes. The Qwen3 vocab-only GGUF (5.9 MB) and 481 cases are committed so it runs in CI. 481 committed cases and 20,081 random cases: 0 mismatches on split, encode and decode. Nine of ten single-rule mutations fail the test; the tenth is equivalent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 22, 2026
bong-water-water-bong
pushed a commit
that referenced
this pull request
Sep 26, 2026
…ama.cpp #2, #5, #6) The pin adds the zaya architecture and its converter (#2, #5) on top of #104's #95 fixes, and carries #106's scalar-scatter SET_ROWS onto the tracked branch (#6): #106 pinned c358334, which is on no branch of the fork. registry/architectures.json, regenerated: ZayaForCausalLM maps to hrx and vulkan (266 architectures, all mapped). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
added a commit
that referenced
this pull request
Sep 26, 2026
…nly architectures (#107) * ZAYA1 (Zyphra) on Vulkan: serve routes fork-only architectures to our llama.cpp `1bit serve --device vulkan` reads the GGUF's general.architecture (app/gguf_meta.h) and runs an architecture only our llama.cpp implements (zaya) on the HRX build's Vulkan0, even when the upstream Vulkan build is present. registry_build counts such architectures for vulkan the same way, and check_models.tsv checks ZAYA1-8B Q4_K_M on vulkan. docs/vulkan.md: converting and serving ZAYA1, with the checks against transformers and the measured speeds on Strix Halo. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ZAYA1 on HRX and ROCm: docs, and serve_e2e on vulkan and hrx in the registry Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Pin llama.cpp 5556bf2: ZAYA1 on Vulkan, HRX and ROCm (1bit-MONSTER/llama.cpp #2, #5, #6) The pin adds the zaya architecture and its converter (#2, #5) on top of #104's #95 fixes, and carries #106's scalar-scatter SET_ROWS onto the tracked branch (#6): #106 pinned c358334, which is on no branch of the fork. registry/architectures.json, regenerated: ZayaForCausalLM maps to hrx and vulkan (266 architectures, all mapped). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * registry: vulkan and hrx rows re-checked at b9812e0 (16/16 pass, ZAYA1-8B included) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: build gguf_meta_test Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Text in, token ids out, with no Python and no ICU. It reads the vocabulary from GGUF metadata.
What
engine/core/tokenizer.*: the same steps as HFtokenizers, in order:tokenizer.ggml.prevalues it has no golden for.engine/core/unicode.*: UTF-8 codec, character classes and NFC. The tables come from Pythonunicodedata16.0.0 viatools/gen_unicode_tables.py.docs/tokenizer.mdexplains the design, the verification and the known limits.Verification
golden_tokenizerchecks three things per case, all exact against HFtokenizers0.22.2:Mutation sweep: ten single-rule deletions, nine caught (for example, dropping NFC gives 187 mismatches). The tenth, the
\p{N}branch, is an equivalent mutant. An earlier version of the cases missed contraction rules, because" 'd"is split by the punctuation branch before the contraction branch runs. The cases now put contractions after letters.Test plan
cteston strixhalo: 5/5 (incl.golden_cpu)ctestlocally: 4/4 (no model needed)-Wall -Wextra -Wpedantic🤖 Generated with Claude Code