Skip to content

common : store the ngram caches in an unordered_dense map - #5

Open
jadidbourbaki wants to merge 3 commits into
ngram-cache-no-copyfrom
ngram-cache-outer-map
Open

jadidbourbaki wants to merge 3 commits into
ngram-cache-no-copyfrom
ngram-cache-outer-map

Conversation

@jadidbourbaki

@jadidbourbaki jadidbourbaki commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #2.

I replaced the outer std::unordered_map of the n-gram caches with an ankerl::unordered_dense::segmented_map. The inner maps are unchanged.

A segmented map grows in blocks of 4096 bytes. A plain ankerl::unordered_dense::map keeps its entries in one vector that doubles as it fills. The 541 MB static cache has 8.9 million 2-grams, so the last doubling held both vectors at once and raised the peak memory by about 0.65 GB.

I vendored unordered_dense v5.0.1 by @martinus under vendor/ankerl and pinned it in scripts/sync_vendor.py.

Results

I borrowed the benchmark setup from ggml-org/llama.cpp#5479, which builds the static cache from WikiText-103 and runs with a context of 4096 tokens. Since this PR makes no algorithmic changes to lookup decoding, the dataset mainly matters for the acceptance rate, which stays identical. The metrics that change are the latency per drafted token, the load time of the static cache, and the memory used by the static cache.

I ran llama-lookup-stats on WikiText-103 test with static caches built from prefixes of WikiText-103 train. Each value is the median of 3 runs on the CPU of an Apple M4 Pro with 14 cores and 48 GB of memory, running macOS 26.5.1. A corpus size of 0 means no static cache. #2 is the baseline.

Corpus Size (MB) #2 Latency (µs / tok) PR Latency (µs / tok) #2 Load Time (ms) PR Load Time (ms) #2 Peak Memory (MB) PR Peak Memory (MB)
0 1.89 1.72 896 883
25 4.12 3.94 474 288 1145 1108
50 4.42 4.11 845 528 1347 1292
100 4.64 4.55 1297 922 1638 1563
200 5.62 4.96 2517 1648 2180 2074
541 6.47 5.81 5283 3508 3548 3364

Drafting latency per drafted tokenStatic cache load timePeak memory

The benchmark code and full tables are in ngram-cache-bench.

@jadidbourbaki
jadidbourbaki added this pull request to stack #8 September 26, 2026 16:43
@jadidbourbaki
jadidbourbaki removed this pull request from stack #8 September 26, 2026 16:47
@jadidbourbaki
jadidbourbaki added this pull request to stack #9 September 26, 2026 16:47
@jadidbourbaki
jadidbourbaki removed this pull request from stack #9 September 26, 2026 17:07
@jadidbourbaki
jadidbourbaki added this pull request to stack #11 September 26, 2026 17:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant