馃殌 The feature, motivation and pitch
Only found a discussion asking about it but from evaluation it seems that Q4 is now better than FP8 and closer/almost equal to fp16 cache. I personally don't use this engine and am just looking from the outside, but I believe this may benefit some of its users who may be trying to squeeze in a bit more context without reducing the overall accuracy by much.
Additional context
Here's the evaluation between the different cache types: turboderp/exllamav2/doc/qcache_eval.md
馃殌 The feature, motivation and pitch
Only found a discussion asking about it but from evaluation it seems that Q4 is now better than FP8 and closer/almost equal to fp16 cache. I personally don't use this engine and am just looking from the outside, but I believe this may benefit some of its users who may be trying to squeeze in a bit more context without reducing the overall accuracy by much.
Additional context
Here's the evaluation between the different cache types: turboderp/exllamav2/doc/qcache_eval.md