Skip to content

Tokenizers: Hugging Face tokenizers v0.23.2 behind our C ABI, exact on 18 models - #19

Merged
bong-water-water-bong merged 2 commits into
mainfrom
hf-tokenizers
Sep 23, 2026
Merged

bong-water-water-bong merged 2 commits into
mainfrom
hf-tokenizers

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

Any model's tokenizer.json, through Hugging Face tokenizers itself. This is the Rust library behind the tokenizers npm and PyPI packages, at the same release (v0.23.2).

  • third_party/tokenizers = huggingface/tokenizers v0.23.2 (88a4498a). There's no official C binding upstream, and mlc-ai/tokenizers-cpp is stuck on 0.21.2.
  • hf_tokenizers/ is our own C ABI over it (about 100 lines of Rust), plus a small C++ owner, onebit::HfTokenizer. It builds as a static library with only the onig feature.
  • The Rust release is pinned in rust-toolchain.toml (1.98.1) and the crate graph in Cargo.lock (--locked).
  • -DONEBIT_HF_TOKENIZERS=ON builds it. With -DONEBIT_TOKENIZER_ROOT=<model dirs>, ctest runs one golden test per model.
  • tools/make_tokenizer_golden.py makes the goldens with Python tokenizers==0.23.2. The corpus covers eight scripts, code, numbers, whitespace runs, emoji with ZWJ and chat markup.
  • bump-tokenizers.yml moves the pin to each new release and refreshes the lock.

Verified

18/18 models, 17/17 cases each: identical ids and identical decoded text. The models are Qwen3 0.6B/1.7B/4B/8B, Qwen3-VL-4B, Qwen3.5-4B, Qwen3.6-35B-A3B, Llama 3.1-8B, 3.2-1B and 3.2-3B, Gemma3 1B/4B, Gemma4 E2B/E4B, Phi4-mini, LFM2 1.2B/2.6B and Nanbeige4.1-3B, with vocabularies from 64k to 262k. The engine's own Qwen golden also matches, 14/14. That ran through the real CMake build: ctest -R hf_tokenizer passes 18/18.

🤖 Generated with Claude Code

@bong-water-water-bong
bong-water-water-bong force-pushed the hf-tokenizers branch 2 times, most recently from f991cbd to d7b2d8d Compare September 23, 2026 20:45
bong-water-water-bong and others added 2 commits September 23, 2026 17:55
…e-exact on 18 models

third_party/tokenizers = huggingface/tokenizers v0.23.2 (88a4498a), the
release the tokenizers npm/PyPI packages ship. hf_tokenizers/ is our C ABI
(from_json/encode/decode/vocab_size/last_error) built as a static lib
(onig only), with Rust pinned (rust-toolchain.toml 1.98.1) and
Cargo.lock (--locked). -DONEBIT_HF_TOKENIZERS=ON builds it; hf_tokenizer_test
runs golden cases made by the same Python release for 18 models under
ONEBIT_TOKENIZER_ROOT. bump-tokenizers.yml keeps it current.

Verified: 18/18 models x 17/17 cases identical ids and decoded text
(vocab 64k-262k: Qwen3/3.5/3.6/VL, Llama 3.1/3.2, Gemma3/4, Phi4-mini, LFM2,
Nanbeige); the engine's Qwen golden 14/14.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…exempt)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bong-water-water-bong
bong-water-water-bong merged commit d844552 into main Sep 23, 2026
1 check passed
@bong-water-water-bong
bong-water-water-bong deleted the hf-tokenizers branch September 23, 2026 21:08
bong-water-water-bong pushed a commit that referenced this pull request Sep 26, 2026
Bumps third_party/llama.cpp-vulkan to 57c22a3 (fork PR #19, 1bit/moe-stream).
ExpertCache takes caller-owned slot memory (CacheOptions::region), as the fork's copy does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong added a commit that referenced this pull request Sep 26, 2026
* serve: --moe-slots streams MoE experts from the model file (Vulkan)

Bumps third_party/llama.cpp-vulkan to 57c22a3 (fork PR #19, 1bit/moe-stream).
ExpertCache takes caller-owned slot memory (CacheOptions::region), as the fork's copy does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: --moe-slots in serve, streaming in the inference path measured

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant