common : read ngram cache parts by reference instead of copying - #2
Open
jadidbourbaki wants to merge 1 commit into
Open
jadidbourbaki wants to merge 1 commit into
jadidbourbaki wants to merge 1 commit into
Conversation
jadidbourbaki
added this pull request to stack #4
September 26, 2026 14:15
jadidbourbaki
removed this pull request from stack #4
September 26, 2026 16:46
jadidbourbaki
added this pull request to stack #9
September 26, 2026 16:47
jadidbourbaki
removed this pull request from stack #9
September 26, 2026 17:07
jadidbourbaki
added this pull request to stack #11
September 26, 2026 17:07
JohannesGaessler
approved these changes
Sep 26, 2026
|
Blocked for making me waste my time. |
Owner
Author
|
Hi @JohannesGaessler my apologies, this is something I am testing in my own fork before sending a PR upstream. The tag of your username was accidental. Apologies for wasting your time here. |
Hey, please don't be pathetic on GitHub. |
alainnothere
added a commit
to alainnothere/llama.cpp
that referenced
this pull request
Sep 28, 2026
…ast customer's homework try_draft copied whole unordered_map parts by value on every drafted token; they are read by reference now, 8.54 to 0.64 us per drafted token in lookup-stats with a static cache (jadidbourbaki#2, plus lemire's threshold pre-check from ggml-org#12). begin() was a no-op, so a reused slot drafted from the previous request's n-grams at 13% acceptance; it now clears the context cache (ggml-org#27866), T2 15.2 to 39.3 t/s. finished context caches go to ngram_cache_done so the -lcd write-back still sees every request.
Real professional. |
|
@jadidbourbaki I recommend you limit conversation to collaborators. |
Repository owner
locked and limited conversation to collaborators
Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The drafting loop in
common/ngram-cache.cppcopied an innercommon_ngram_cache_partmap in three places on every drafting step. I changed all three to read the part through a const reference.try_draftcopied the part of the looked-up 2-gram from the static cache.try_draftcopied the part of every n-gram size it tried in the context and dynamic caches.common_ngram_cache_draftcopied the static part of the current 2-gram. It now binds a const reference to the part, or to an empty part when the 2-gram is missing.Results
I borrowed the benchmark setup from ggml-org/llama.cpp#5479, which builds the static cache from WikiText-103 and runs with a context of 4096 tokens. Since this PR makes no algorithmic changes to lookup decoding, the dataset mainly matters for the acceptance rate, which stays almost identical. The metrics that change are the latency per drafted token, the load time of the static cache, and the memory used by the static cache. I ran
llama-lookup-statson WikiText-103 test with static caches built from prefixes of WikiText-103 train. Each value is the median of 3 runs on the CPU of an Apple M4 Pro with 14 cores and 48 GB of memory, running macOS 26.5.1. A corpus size of 0 means no static cache.std::unordered_mapchanges its iteration order, so the two versions break ties between equally frequent tokens differently.The benchmark code and full tables are in ngram-cache-bench.