Skip to content

common : back the static ngram cache with a constmap - #7

Open
jadidbourbaki wants to merge 2 commits into
ngram-cache-inner-vectorfrom
ngram-cache-constmap
Open

jadidbourbaki wants to merge 2 commits into
ngram-cache-inner-vectorfrom
ngram-cache-constmap

Conversation

@jadidbourbaki

@jadidbourbaki jadidbourbaki commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #10.

I replaced the static n-gram cache with a read-only cache built on fastconstmap by @lemire. A VerifiedConstMap maps each 2-gram to a span of one contiguous array of (token, count) pairs sorted by token. The context and dynamic caches are unchanged.

  • I vendored fastconstmap at 990afd0 under vendor/constmap and pinned it in scripts/sync_vendor.py.
  • llama-lookup-create writes the static cache in a new file format. llama-lookup and llama-lookup-stats read it.
  • common_ngram_cache_draft takes the static cache as a const common_ngram_cache_static *, and all sequences in common/speculative.cpp share one static cache.

llama-lookup-merge still reads and writes the old format, so it keeps working for dynamic caches.

Results

I borrowed the benchmark setup from ggml-org/llama.cpp#5479, which builds the static cache from WikiText-103 and runs with a context of 4096 tokens. Since this PR makes no algorithmic changes to lookup decoding, the dataset mainly matters for the acceptance rate, which stays almost identical. The metrics that change are the latency per drafted token, the load time of the static cache, and the memory used by the static cache.

I ran llama-lookup-stats on WikiText-103 test with static caches built from prefixes of WikiText-103 train. Each value is the median of 3 runs on the CPU of an Apple M4 Pro with 14 cores and 48 GB of memory, running macOS 26.5.1. A corpus size of 0 means no static cache. #10 is the baseline.

Corpus Size (MB) #10 Latency (µs / tok) PR Latency (µs / tok) #10 Load Time (ms) PR Load Time (ms) #10 Peak Memory (MB) PR Peak Memory (MB)
0 0.82 0.89 846 845
25 3.31 3.06 234 37 937 890
50 3.46 3.25 482 53 1015 924
100 3.68 3.32 874 82 1097 982
200 3.95 3.47 1650 133 1302 1073
541 4.78 3.98 3756 233 1707 1308
  • Loading the static cache is 6.32x to 16.12x faster.
  • Peak memory is up to 1.30x lower, and the static cache itself uses 1.83x to 2.13x less memory, about the size of its file.
  • Drafting is 1.06x to 1.20x faster with a static cache and about the same speed without one.
  • Acceptance is identical, because both versions store the followers of each 2-gram sorted by token.

Drafting latency per drafted tokenStatic cache load timePeak memory

The benchmark code and full tables are in ngram-cache-bench.

@jadidbourbaki
jadidbourbaki added this pull request to stack #8 September 26, 2026 16:43
@jadidbourbaki
jadidbourbaki removed this pull request from stack #8 September 26, 2026 16:47
@jadidbourbaki
jadidbourbaki added this pull request to stack #9 September 26, 2026 16:47
@jadidbourbaki
jadidbourbaki removed this pull request from stack #9 September 26, 2026 17:07
@jadidbourbaki
jadidbourbaki changed the base branch from ngram-cache-outer-map to ngram-cache-inner-vector September 26, 2026 17:07
@jadidbourbaki
jadidbourbaki added this pull request to stack #11 September 26, 2026 17:07
@jadidbourbaki
jadidbourbaki force-pushed the ngram-cache-constmap branch 2 times, most recently from 948eb38 to bdd4ea7 Compare September 26, 2026 17:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant