Repository navigation
q4_k_m #47
Description
Activity
There is a chance for Q4_K_M but the quantization logic will have to be revised (i.e. which keys to keep in low/high precision). qkv is fused into one in the reference model so deciding based on to_q/to_k/to_v is not an option and the rest ot the logic doesn't apply 1:1.
Q8_K is not a valid output quantization format for llama.cpp (at least not at the moment).
The code you show is only for dequantization, and we're already using it. The K quants need C++ code to quantize and my solution to generate the current ones is really hacky, I'll upload it as a diff from the original llama.cpp repo once I'm done with everything else.
I'll upload it as a diff from the original llama.cpp repo once I'm done with everything else.
Oh cool, are you working on a way to move the slow operations to c++ or something?
i'm tried to convert flux to K types , but , it give convert.py: error: Unknown quant/file type Q4_K
can u help me to pypass this step , i'm already strugle inside codes !
i tried to edit the code but it return same size when i choose K type@RandomGitUser321 Yes that's one thing I'm looking at, as well as T5 in gguf (this is easier since llama.cpp supports it lol)
@al-swaiti Please re read the comment above:
The K quants need C++ code
Reacted by Abdallah and RandomGitUser321@al-swaiti Again, re-read the original comment and be patient.
The K quants need C++ code to quantize and my solution to generate the current ones is really hacky,
I'll upload it as a diff from the original llama.cpp repo once I'm done with everything else.Specifically this part:
once I'm done with everything else.
thank you
@RandomGitUser321 Yes that's one thing I'm looking at, as well as T5 in gguf (this is easier since llama.cpp supports it lol)
Awesome, I kind of figured that was also going to be a next step as well. It would be pretty sweet to be able to use t5 GGUFs and shouldn't be that much more work to implement, now that you've laid all this other groundwork. Keep up the great work!
@RandomGitUser321 First version of GGUF text encoder is up, check #49 for more info
Reacted by RandomGitUser321@city96 I'll test q4_k_m and q8_0. Q4_k_m is a big go-to in the LLM world and I imagine a lot of people will likely gravitate toward it. In theory, it should be pretty close in performance to the fp8 versions.
The T5 Q8_0 is working fine for me.
Similar size, but better results than T5 FP8.
Didn't tried the T5 q4_k_m, but it may be usefull if you only have 16GB RAM
Updated instruction with patch, see updated readme in the tools folder. Has logic for _M quants too, although the effects are diminished due to q/k/v not being separated.
THANK YOU TO MUCH I ALREADY QUANTIZED 3 MODEL ,,,,
REACH HERE
./llama.cpp/build/bin/llama-quantize Stable-diffusion/mege-BF16.gguf q4ks.gguf
IQ1_M🤣
Please do not use IQ1_S, IQ1_M, IQ2_S, IQ2_XXS, IQ2_XS or Q2_K_S quantization without an importance matrix ANY IDEA !
It says right there, those quantization types are unusable without an imatrix. Don't make IQ quants either they won't work in ComfyUI.
Reacted by Abdallah@city96 your opinion is very important , this is my special merge (coding)
https://civitai.com/models/682369?modelVersionId=764918
thanks@al-swaiti Seems to work fine. I've looked at Q4_K_M and the keys are all in the right format, and looks like the _M logic worked with some being in Q4_K and some other keys in Q5_K as intended. Two things to nitpick:
- "Q_8" as you named it isn't a quantization. It should be called "Q8_0".
- I think the license needs to be the same as the dev model if you used any part of the dev model instead of just "Apache2"
Other than that it looks good to me, output is perfectly coherent. Thank you for following the guide and not disabling any of the key checks/messing with the format lol.
@city96
okay ,,, about license i think you already know the flux dev is aura flow model trained on better quality dataset , is there any part of the model showed using of dev!? ,- here i was trying another way of quantization ,, but i struggled
could i apply it for normal comfyui models ?
"torchao_serialized"
https://gist.github.com/sayakpaul/e1f28e86d0756d587c0b898c73822c47
- here i was trying another way of quantization ,, but i struggled
okay ,,, about license i think you already know the flux dev is aura flow model trained on better quality dataset , is there any part of the model showed using of dev!? ,
I'm not sure what you're saying here. Are you claiming you've merged AuraFlow with Flux1-schnell?
could i apply it for normal comfyui models ?
You could implement it as a custom node but windows support would most likely need extra work based on just the comments there. As this is a different quantization method, it's also out of scope for this repository/issues and you should do your own research on this topic.

hi ,,, thanks for this amazing job , adding the features of llama.cpp to sota models is big step , i spent all night searching about this topic ,
Firstly is there chance to make q4_k_m ,
secondly : T5 encoder already supported do u planing to convert it
i,m waitin for view lines of codes to convert to Q8_K and others
i think its already supported in gguf/quantise code