Skip to content

common : store the ngram caches in unordered_dense maps - #3

Closed
jadidbourbaki wants to merge 3 commits into
ngram-cache-no-copyfrom
ngram-cache-flat-map
Closed

jadidbourbaki wants to merge 3 commits into
ngram-cache-no-copyfrom
ngram-cache-flat-map

Conversation

@jadidbourbaki

@jadidbourbaki jadidbourbaki commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #2.

I replaced the nested std::unordered_maps of the n-gram caches with two flatter structures.

  • The outer map from each n-gram to its following tokens is now an ankerl::unordered_dense::map. I vendored unordered_dense v5.0.1 under vendor/ankerl and pinned it in scripts/sync_vendor.py.
  • The inner map from each following token to its count is now a vector of (token, count) pairs sorted by token. Lookups binary search the vector.

I first made the inner map an unordered_dense map as well. Memory went up, because an unordered_dense map with one entry takes about 430 bytes. A pair in the sorted vector takes 8 bytes.

common/ngram-cache.cpp, the drafting logic, and the cache file format are unchanged.

Results

I ran llama-lookup-stats on WikiText-103 test with static caches built from prefixes of WikiText-103 train. Each value is the median of 3 runs on an Apple M4 Pro. A corpus size of 0 means no static cache. #2 is the baseline.

Corpus Size (MB) #2 Latency (µs / tok) PR Latency (µs / tok) #2 Load Time (ms) PR Load Time (ms) #2 Memory (MB) PR Memory (MB)
0 1.89 0.74
25 4.12 3.82 474 241 249 110
50 4.42 4.20 845 488 451 184
100 4.64 4.74 1297 989 743 357
200 5.62 5.00 2517 1921 1284 650
541 6.47 5.96 5283 4483 2652 1341
  • Drafting is 2.56x faster without a static cache and 0.98x to 1.12x as fast with one.
  • Loading the static cache is 1.18x to 1.97x faster.
  • The static cache uses 1.98x to 2.45x less memory.
  • Acceptance changes by at most 0.08 percentage points, because the sorted vectors break ties between equally frequent tokens in a different order.

Drafting latency per drafted token

Static cache load time

Static cache memory

The benchmark code and full tables are in ngram-cache-bench.

@jadidbourbaki
jadidbourbaki added this pull request to stack #4 September 26, 2026 14:15
@jadidbourbaki

Copy link
Copy Markdown
Owner Author

Replaced by #5.

@jadidbourbaki
jadidbourbaki removed this pull request from stack #4 September 26, 2026 16:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant