Skip to content

common : store the tokens after each n-gram in a sorted vector - #6

Closed
jadidbourbaki wants to merge 1 commit into
ngram-cache-outer-mapfrom
ngram-cache-inner-vector
Closed

jadidbourbaki wants to merge 1 commit into
ngram-cache-outer-mapfrom
ngram-cache-inner-vector

Conversation

@jadidbourbaki

@jadidbourbaki jadidbourbaki commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #5.

I replaced the inner std::unordered_map of the n-gram caches with a vector of (token, count) pairs sorted by token. The inner map sends each token that follows an n-gram to its count. Lookups binary search the vector.

common_ngram_cache_part keeps the find, emplace, iteration, and size calls of common/ngram-cache.cpp, so common/ngram-cache.cpp is unchanged. C++23 has std::flat_map for a sorted vector with a map interface, but llama.cpp builds as C++17.

The drafting logic and the cache file format are unchanged.

Results

The benchmark is running.

@jadidbourbaki
jadidbourbaki deleted the ngram-cache-inner-vector branch September 26, 2026 14:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant