Repository navigation
Conversation
|
Would be nice this to be a ladder if possible F16->Q8->Q4 even though I'm not sure what impact double quantization might have on PPL. Worth a test though. It would make things smoother for constrained deployments |
|
Would it be possible to rather design a method that runs the cache in f16 for some specific length and then quantizes to q8_0 in batches of 128 tokens. For example, imagine at first you go from 0 to 255 slots in f16, but when inserting the 256th f16 token, you also quantize the 128 first values to q8_0, and then proceed with the next 128..255 token block in f16. I'm just throwing this as an idea. I've no idea if it's easy, nor do I know whether 128..255 is the enough for keeping accuracy in the KV cache. There's probably an optimal length where degradation from KV quantization is virtually indistinguishable from having it all in f16, and it would be interesting to know whether it's relatively short run, like just 100 tokens, or if it must be comparable to a full assistant turn of perhaps 10000 tokens. Either way, there should be considerable memory savings there, and perhaps it would make coarser KV quantization levels more feasible. |
a650430 to
69804e5
Compare
|
This is getting closer, but I need to split out a pre-req PR first to reject restoring a cache with incompatible attention rotation, e.g. when LLAMA_ATTN_ROT_DISABLE differs between save and restore. This becomes more obvious when doing save/restore tests across lazy quants (since f16 is not rotated by default and q8 is). Note from Astra: |
7703565 to
e5b4e53
Compare
Assisted-by: Codex
…lls up) Assisted-by: Codex
Assisted-by: Codex
e5b4e53 to
86b8fe1
Compare
Note: There are still aspects of this I am working to understand, test, and fix before it will be ready for maintainer review.
Overview
Original idea at #27926.
Adds a
LLAMA_KV_CACHE_LAZY_QUANTenv var - when set, and q8_0 kv cache is used, the cache is initially overlayed with an fp16 view of about ~half the tokens. Once that fills up, the cells are quanted down to q8_0, the view is discarded, and processing continues with traditional quanted kv.Additional information
Current limits/constraints (mostly for simplicity) include:
gemma4 arch is excluded because it does some things with shared tensor pointers I don't totally understandfixedObvious follow-ups would be support for more quant types, and support for backend-native conversion without going through host memory.
Requirements