Repository navigation
Tokenizers: Hugging Face tokenizers v0.23.2 behind our C ABI, exact on 18 models - #19
Merged
Merged
Conversation
bong-water-water-bong
force-pushed
the
hf-tokenizers
branch
2 times, most recently
from
September 23, 2026 20:45
f991cbd to
d7b2d8d
Compare
…e-exact on 18 models third_party/tokenizers = huggingface/tokenizers v0.23.2 (88a4498a), the release the tokenizers npm/PyPI packages ship. hf_tokenizers/ is our C ABI (from_json/encode/decode/vocab_size/last_error) built as a static lib (onig only), with Rust pinned (rust-toolchain.toml 1.98.1) and Cargo.lock (--locked). -DONEBIT_HF_TOKENIZERS=ON builds it; hf_tokenizer_test runs golden cases made by the same Python release for 18 models under ONEBIT_TOKENIZER_ROOT. bump-tokenizers.yml keeps it current. Verified: 18/18 models x 17/17 cases identical ids and decoded text (vocab 64k-262k: Qwen3/3.5/3.6/VL, Llama 3.1/3.2, Gemma3/4, Phi4-mini, LFM2, Nanbeige); the engine's Qwen golden 14/14. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…exempt) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
force-pushed
the
hf-tokenizers
branch
from
September 23, 2026 20:55
d7b2d8d to
1ef449d
Compare
This was referenced Sep 26, 2026
bong-water-water-bong
pushed a commit
that referenced
this pull request
Sep 26, 2026
Bumps third_party/llama.cpp-vulkan to 57c22a3 (fork PR #19, 1bit/moe-stream). ExpertCache takes caller-owned slot memory (CacheOptions::region), as the fork's copy does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bong-water-water-bong
added a commit
that referenced
this pull request
Sep 26, 2026
* serve: --moe-slots streams MoE experts from the model file (Vulkan) Bumps third_party/llama.cpp-vulkan to 57c22a3 (fork PR #19, 1bit/moe-stream). ExpertCache takes caller-owned slot memory (CacheOptions::region), as the fork's copy does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: --moe-slots in serve, streaming in the inference path measured Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: bong-water-water-bong <bong-water-water-bong@1bit.gg> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Any model's
tokenizer.json, through Hugging Facetokenizersitself. This is the Rust library behind thetokenizersnpm and PyPI packages, at the same release (v0.23.2).third_party/tokenizers=huggingface/tokenizersv0.23.2 (88a4498a). There's no official C binding upstream, and mlc-ai/tokenizers-cpp is stuck on 0.21.2.hf_tokenizers/is our own C ABI over it (about 100 lines of Rust), plus a small C++ owner,onebit::HfTokenizer. It builds as a static library with only theonigfeature.rust-toolchain.toml(1.98.1) and the crate graph inCargo.lock(--locked).-DONEBIT_HF_TOKENIZERS=ONbuilds it. With-DONEBIT_TOKENIZER_ROOT=<model dirs>, ctest runs one golden test per model.tools/make_tokenizer_golden.pymakes the goldens with Pythontokenizers==0.23.2. The corpus covers eight scripts, code, numbers, whitespace runs, emoji with ZWJ and chat markup.bump-tokenizers.ymlmoves the pin to each new release and refreshes the lock.Verified
18/18 models, 17/17 cases each: identical ids and identical decoded text. The models are Qwen3 0.6B/1.7B/4B/8B, Qwen3-VL-4B, Qwen3.5-4B, Qwen3.6-35B-A3B, Llama 3.1-8B, 3.2-1B and 3.2-3B, Gemma3 1B/4B, Gemma4 E2B/E4B, Phi4-mini, LFM2 1.2B/2.6B and Nanbeige4.1-3B, with vocabularies from 64k to 262k. The engine's own Qwen golden also matches, 14/14. That ran through the real CMake build:
ctest -R hf_tokenizerpasses 18/18.🤖 Generated with Claude Code