Skip to content

common : read ngram cache parts by reference instead of copying - #2

Open
jadidbourbaki wants to merge 1 commit into
masterfrom
ngram-cache-no-copy
Open

jadidbourbaki wants to merge 1 commit into
masterfrom
ngram-cache-no-copy

Conversation

@jadidbourbaki

@jadidbourbaki jadidbourbaki commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

The drafting loop in common/ngram-cache.cpp copied an inner common_ngram_cache_part map in three places on every drafting step. I changed all three to read the part through a const reference.

  • try_draft copied the part of the looked-up 2-gram from the static cache.
  • try_draft copied the part of every n-gram size it tried in the context and dynamic caches.
  • common_ngram_cache_draft copied the static part of the current 2-gram. It now binds a const reference to the part, or to an empty part when the 2-gram is missing.

Results

I borrowed the benchmark setup from ggml-org/llama.cpp#5479, which builds the static cache from WikiText-103 and runs with a context of 4096 tokens. Since this PR makes no algorithmic changes to lookup decoding, the dataset mainly matters for the acceptance rate, which stays almost identical. The metrics that change are the latency per drafted token, the load time of the static cache, and the memory used by the static cache. I ran llama-lookup-stats on WikiText-103 test with static caches built from prefixes of WikiText-103 train. Each value is the median of 3 runs on the CPU of an Apple M4 Pro with 14 cores and 48 GB of memory, running macOS 26.5.1. A corpus size of 0 means no static cache.

Corpus Size (MB) Baseline Latency (µs / tok) PR Latency (µs / tok)
0 8.54 1.89
25 45.61 4.12
50 59.73 4.42
100 83.46 4.64
200 113.46 5.62
541 165.48 6.47
  • Drafting is 4.5x to 25.6x faster.
  • Loading the static cache takes the same time, about 5.3 s for the 541 MB cache.
  • Acceptance changes by at most 0.15 percentage points. Copying a libc++ std::unordered_map changes its iteration order, so the two versions break ties between equally frequent tokens differently.

Drafting latency per drafted tokenStatic cache load timePeak memory

The benchmark code and full tables are in ngram-cache-bench.

@jadidbourbaki
jadidbourbaki added this pull request to stack #4 September 26, 2026 14:15
@jadidbourbaki
jadidbourbaki removed this pull request from stack #4 September 26, 2026 16:46
@jadidbourbaki
jadidbourbaki added this pull request to stack #9 September 26, 2026 16:47
@jadidbourbaki
jadidbourbaki removed this pull request from stack #9 September 26, 2026 17:07
@jadidbourbaki
jadidbourbaki added this pull request to stack #11 September 26, 2026 17:07
@JohannesGaessler

Copy link
Copy Markdown

Blocked for making me waste my time.

@jadidbourbaki

Copy link
Copy Markdown
Owner Author

Hi @JohannesGaessler my apologies, this is something I am testing in my own fork before sending a PR upstream. The tag of your username was accidental. Apologies for wasting your time here.

@DesolateIntention

Copy link
Copy Markdown

Blocked for making me waste my time.

Hey, please don't be pathetic on GitHub.

alainnothere added a commit to alainnothere/llama.cpp that referenced this pull request Sep 28, 2026
…ast customer's homework

try_draft copied whole unordered_map parts by value on every drafted token; they are read by reference now, 8.54 to 0.64 us per drafted token in lookup-stats with a static cache (jadidbourbaki#2, plus lemire's threshold pre-check from ggml-org#12). begin() was a no-op, so a reused slot drafted from the previous request's n-grams at 13% acceptance; it now clears the context cache (ggml-org#27866), T2 15.2 to 39.3 t/s. finished context caches go to ngram_cache_done so the -lcd write-back still sees every request.
@Bits360

Bits360 commented Sep 28, 2026

Copy link
Copy Markdown

Blocked for making me waste my time.

Real professional.

@WalterBarrett

Copy link
Copy Markdown

@jadidbourbaki I recommend you limit conversation to collaborators.

Repository owner locked and limited conversation to collaborators Sep 28, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants