Repository navigation
Conversation
openbmb/MiniCPM5-2B ships a byte-identical tokenizer.json to openbmb/MiniCPM5-1B (md5 ee55db96827d21929c7c5db2092596a8), so it hashes to the chkhsh already registered for "minicpm5" and needs no new entry. Document this on both the generated entry in conversion/base.py and the model list in convert_hf_to_gguf_update.py, which is the source the block is generated from. Kept to a single comment line so the get_existing_models regex still matches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Hi @cyxu0401, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
It doesn't work like that, the comment is autogenerated and we will notice if anyone tries to add a duplicate. |
Follow-up to #23384 (MiniCPM5-1B), now that
openbmb/MiniCPM5-2B and
openbmb/MiniCPM5-2B-GGUF
are public.
Comment-only change — no behaviour change.
MiniCPM5-2B ships a byte-identical
tokenizer.jsonto 1B(md5
ee55db96827d21929c7c5db2092596a8), so running theconvert_hf_to_gguf_update.pyrecipe against it reproduces the hash that isalready in the table:
Verified against four copies of the file — HF 1B, HF 2B, and two local
checkpoints — all identical.
So 2B already converts correctly today and needs no code change. This commit
only records why, because the obvious move for the next person adding 2B is
to append a second entry, and that would be a silent regression:
get_vocab_base_pre()are sequentialifstatements, notelif, so a second entry carrying the same hash placed later wins forboth models — see the existing
mpt/olmoandbert-bge/jina-v2-enpairs, where the later entry already shadows the earlier one;
minicpm5would break every GGUF alreadypublished with
tokenizer.ggml.pre = "minicpm5", including openbmb's own1B and 2B releases.
The note is deliberately kept on one line:
get_existing_models()inconvert_hf_to_gguf_update.pymatches theif chkhsh == .../res = ...pair with a regex that tolerates only a single line between them, so splitting
the comment across two lines breaks regeneration. Confirmed by trying it.
The chat template for 2B is a separate matter and is in its own PR.