Skip to content

Tokenizer: byte-level BPE from GGUF, exact against HF tokenizers - #2

Merged
bong-water-water-bong merged 1 commit into
mainfrom
tokenizer
Sep 22, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
tokenizer

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Text in, token ids out, with no Python and no ICU. It reads the vocabulary from GGUF metadata.

What

  • engine/core/tokenizer.*: the same steps as HF tokenizers, in order:
    • added-token split (leftmost-longest)
    • NFC normalization
    • the Qwen2 split regex, ported by hand one branch at a time
    • GPT-2 byte-level mapping, then BPE merges
    • decode
  • It refuses tokenizer.ggml.pre values it has no golden for.
  • engine/core/unicode.*: UTF-8 codec, character classes and NFC. The tables come from Python unicodedata 16.0.0 via tools/gen_unicode_tables.py.
  • docs/tokenizer.md explains the design, the verification and the known limits.

Verification

golden_tokenizer checks three things per case, all exact against HF tokenizers 0.22.2:

  • pre-token pieces, checked separately because BPE can hide a wrong split
  • ids
  • decoded bytes
case set cases mismatches
committed, runs in CI (vocab-only GGUF 5.9 MB + cases) 481 0
random, seed 7 20,081 0 (0.9 s)

Mutation sweep: ten single-rule deletions, nine caught (for example, dropping NFC gives 187 mismatches). The tenth, the \p{N} branch, is an equivalent mutant. An earlier version of the cases missed contraction rules, because " 'd" is split by the punctuation branch before the contraction branch runs. The cases now put contractions after letters.

Test plan

  • ctest on strixhalo: 5/5 (incl. golden_cpu)
  • ctest locally: 4/4 (no model needed)
  • zero warnings at -Wall -Wextra -Wpedantic

🤖 Generated with Claude Code

…zers

- engine/core/tokenizer: added-token split, NFC, hand-ported Qwen2 split
  regex, GPT-2 byte mapping, BPE merges, decode. Unverified
  tokenizer.ggml.pre values are refused.
- engine/core/unicode: UTF-8 codec, letter/number/whitespace classes, NFC.
  Tables generated from Python unicodedata 16.0.0 by
  tools/gen_unicode_tables.py (no ICU).
- tests: test_unicode (CI), golden_tokenizer checking pre-token pieces,
  ids and decoded bytes. The Qwen3 vocab-only GGUF (5.9 MB) and 481 cases
  are committed so it runs in CI.

481 committed cases and 20,081 random cases: 0 mismatches on split,
encode and decode. Nine of ten single-rule mutations fail the test; the
tenth is equivalent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit 8331046 into main Sep 22, 2026
1 check passed
@bong-water-water-bong
bong-water-water-bong deleted the tokenizer branch September 22, 2026 19:14
bong-water-water-bong pushed a commit that referenced this pull request Sep 26, 2026
…ama.cpp #2, #5)

The pin adds the zaya architecture and its converter on top of 96f6b89 (#104's #95 fixes).
registry/architectures.json, regenerated: ZayaForCausalLM maps to hrx and vulkan
(266 architectures, all mapped).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong pushed a commit that referenced this pull request Sep 26, 2026
…ama.cpp #2, #5, #6)

The pin adds the zaya architecture and its converter (#2, #5) on top of #104's #95 fixes, and
carries #106's scalar-scatter SET_ROWS onto the tracked branch (#6): #106 pinned c358334,
which is on no branch of the fork. registry/architectures.json, regenerated:
ZayaForCausalLM maps to hrx and vulkan (266 architectures, all mapped).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 26, 2026
…nly architectures (#107)

* ZAYA1 (Zyphra) on Vulkan: serve routes fork-only architectures to our llama.cpp

`1bit serve --device vulkan` reads the GGUF's general.architecture (app/gguf_meta.h) and
runs an architecture only our llama.cpp implements (zaya) on the HRX build's Vulkan0, even
when the upstream Vulkan build is present. registry_build counts such architectures for
vulkan the same way, and check_models.tsv checks ZAYA1-8B Q4_K_M on vulkan.

docs/vulkan.md: converting and serving ZAYA1, with the checks against transformers and
the measured speeds on Strix Halo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ZAYA1 on HRX and ROCm: docs, and serve_e2e on vulkan and hrx in the registry

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Pin llama.cpp 5556bf2: ZAYA1 on Vulkan, HRX and ROCm (1bit-MONSTER/llama.cpp #2, #5, #6)

The pin adds the zaya architecture and its converter (#2, #5) on top of #104's #95 fixes, and
carries #106's scalar-scatter SET_ROWS onto the tracked branch (#6): #106 pinned c358334,
which is on no branch of the fork. registry/architectures.json, regenerated:
ZayaForCausalLM maps to hrx and vulkan (266 architectures, all mapped).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* registry: vulkan and hrx rows re-checked at b9812e0 (16/16 pass, ZAYA1-8B included)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: build gguf_meta_test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant