Skip to content

Option to quant KV cache to q8_0 lazily (i.e. when fp16 cache fills up) - #28267

Draft
eapache wants to merge 3 commits into
ggml-org:masterfrom
eapache:ehuus/lazy-kv-cache-quant
Draft

eapache wants to merge 3 commits into
ggml-org:masterfrom
eapache:ehuus/lazy-kv-cache-quant

Conversation

@eapache

@eapache eapache commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Note: There are still aspects of this I am working to understand, test, and fix before it will be ready for maintainer review.

Overview

Original idea at #27926.

Adds a LLAMA_KV_CACHE_LAZY_QUANT env var - when set, and q8_0 kv cache is used, the cache is initially overlayed with an fp16 view of about ~half the tokens. Once that fills up, the cells are quanted down to q8_0, the view is discarded, and processing continues with traditional quanted kv.

Additional information

  • Pro: Short sessions with a "q8" kv cache now get higher precision basically for free.
  • Pro: Long sessions get higher precision for the first ~half of the session, which means less trajectory drift.
  • Con: Adds a moderate pause in decoding at the point where the cache has to be requanted.
  • Con: On systems where the memory bandwidth of reading the whole cache is a limiting factor, more slowdown is seen (since more memory is actually in use) prior to the quant breakpoint.

Current limits/constraints (mostly for simplicity) include:

  • q8_0 only
  • conversion stages through host memory, which adds significant (temporary) host memory usage and a round-trip
  • gemma4 arch is excluded because it does some things with shared tensor pointers I don't totally understand fixed

Obvious follow-ups would be support for more quant types, and support for backend-native conversion without going through host memory.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - idea and basic design was mine; Sol did the initial draft; Astra fixed a lot of the edge cases

@feal87

feal87 commented Sep 3, 2026

Copy link
Copy Markdown

Would be nice this to be a ladder if possible F16->Q8->Q4 even though I'm not sure what impact double quantization might have on PPL. Worth a test though. It would make things smoother for constrained deployments

@alankila

alankila commented Sep 3, 2026

Copy link
Copy Markdown

Would it be possible to rather design a method that runs the cache in f16 for some specific length and then quantizes to q8_0 in batches of 128 tokens. For example, imagine at first you go from 0 to 255 slots in f16, but when inserting the 256th f16 token, you also quantize the 128 first values to q8_0, and then proceed with the next 128..255 token block in f16.

I'm just throwing this as an idea. I've no idea if it's easy, nor do I know whether 128..255 is the enough for keeping accuracy in the KV cache. There's probably an optimal length where degradation from KV quantization is virtually indistinguishable from having it all in f16, and it would be interesting to know whether it's relatively short run, like just 100 tokens, or if it must be comparable to a full assistant turn of perhaps 10000 tokens. Either way, there should be considerable memory savings there, and perhaps it would make coarser KV quantization levels more feasible.

@eapache

eapache commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

@feal87 @alankila both of those (and many more ideas) are possible, but greatly increase the complexity of the solution. I want to get the simple version working and proposed before adding more things.

@eapache
eapache force-pushed the ehuus/lazy-kv-cache-quant branch from a650430 to 69804e5 Compare September 5, 2026 13:05
@github-actions github-actions Bot added the testing Everything test related label Sep 5, 2026
@eapache

eapache commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor Author

This is getting closer, but I need to split out a pre-req PR first to reject restoring a cache with incompatible attention rotation, e.g. when LLAMA_ATTN_ROT_DISABLE differs between save and restore. This becomes more obvious when doing save/restore tests across lazy quants (since f16 is not rotated by default and q8 is).

Note from Astra: Rotation-state fix: record K/V rotation sizes, validate them during restore, update format versions, and test mismatched rotation using ordinary Q8 caches.

@eapache
eapache force-pushed the ehuus/lazy-kv-cache-quant branch from e5b4e53 to 86b8fe1 Compare October 5, 2026 19:29

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants