Skip to content

q4_k_m #47

Description

@al-swaiti

hi ,,, thanks for this amazing job , adding the features of llama.cpp to sota models is big step , i spent all night searching about this topic ,
Firstly is there chance to make q4_k_m ,
secondly : T5 encoder already supported do u planing to convert it
i,m waitin for view lines of codes to convert to Q8_K and others
i think its already supported in gguf/quantise code

image

Activity

  1. city96 commented on Aug 19, 2024

    @city96
    Owner

    There is a chance for Q4_K_M but the quantization logic will have to be revised (i.e. which keys to keep in low/high precision). qkv is fused into one in the reference model so deciding based on to_q/to_k/to_v is not an option and the rest ot the logic doesn't apply 1:1.

    Q8_K is not a valid output quantization format for llama.cpp (at least not at the moment).

    The code you show is only for dequantization, and we're already using it. The K quants need C++ code to quantize and my solution to generate the current ones is really hacky, I'll upload it as a diff from the original llama.cpp repo once I'm done with everything else.

  2. RandomGitUser321 commented on Aug 19, 2024

    @RandomGitUser321
    Contributor

    I'll upload it as a diff from the original llama.cpp repo once I'm done with everything else.

    Oh cool, are you working on a way to move the slow operations to c++ or something?

  3. al-swaiti commented on Aug 19, 2024

    @al-swaiti
    Author

    i'm tried to convert flux to K types , but , it give convert.py: error: Unknown quant/file type Q4_K
    can u help me to pypass this step , i'm already strugle inside codes !
    i tried to edit the code but it return same size when i choose K type

  4. city96 commented on Aug 19, 2024

    @city96
    Owner

    @RandomGitUser321 Yes that's one thing I'm looking at, as well as T5 in gguf (this is easier since llama.cpp supports it lol)

    @al-swaiti Please re read the comment above:

    The K quants need C++ code

  5. al-swaiti commented on Aug 19, 2024

    @al-swaiti
    Author

    image
    do i need to wait for llama-cpp update !

  6. city96 commented on Aug 19, 2024

    @city96
    Owner

    @al-swaiti Again, re-read the original comment and be patient.

    The K quants need C++ code to quantize and my solution to generate the current ones is really hacky,
    I'll upload it as a diff from the original llama.cpp repo once I'm done with everything else.

    Specifically this part:

    once I'm done with everything else.

  7. al-swaiti commented on Aug 19, 2024

    @al-swaiti
    Author

    thank you

  8. RandomGitUser321 commented on Aug 19, 2024

    @RandomGitUser321
    Contributor

    @RandomGitUser321 Yes that's one thing I'm looking at, as well as T5 in gguf (this is easier since llama.cpp supports it lol)

    Awesome, I kind of figured that was also going to be a next step as well. It would be pretty sweet to be able to use t5 GGUFs and shouldn't be that much more work to implement, now that you've laid all this other groundwork. Keep up the great work!

  9. city96 commented on Aug 20, 2024

    @city96
    Owner

    @RandomGitUser321 First version of GGUF text encoder is up, check #49 for more info

  10. RandomGitUser321 commented on Aug 20, 2024

    @RandomGitUser321
    Contributor

    @city96 I'll test q4_k_m and q8_0. Q4_k_m is a big go-to in the LLM world and I imagine a lot of people will likely gravitate toward it. In theory, it should be pretty close in performance to the fp8 versions.

  11. JorgeR81 commented on Aug 20, 2024

    @JorgeR81

    The T5 Q8_0 is working fine for me.

    Similar size, but better results than T5 FP8.

    Didn't tried the T5 q4_k_m, but it may be usefull if you only have 16GB RAM

  12. city96 commented on Aug 21, 2024

    @city96
    Owner

    Updated instruction with patch, see updated readme in the tools folder. Has logic for _M quants too, although the effects are diminished due to q/k/v not being separated.

  13. al-swaiti commented on Aug 21, 2024

    @al-swaiti
    Author

    THANK YOU TO MUCH I ALREADY QUANTIZED 3 MODEL ,,,,
    REACH HERE
    ./llama.cpp/build/bin/llama-quantize Stable-diffusion/mege-BF16.gguf q4ks.gguf
    IQ1_M

    🤣

    Please do not use IQ1_S, IQ1_M, IQ2_S, IQ2_XXS, IQ2_XS or Q2_K_S quantization without an importance matrix ANY IDEA !

  14. city96 commented on Aug 21, 2024

    @city96
    Owner

    It says right there, those quantization types are unusable without an imatrix. Don't make IQ quants either they won't work in ComfyUI.

  15. al-swaiti commented on Aug 25, 2024

    @al-swaiti
    Author

    @city96 your opinion is very important , this is my special merge (coding)
    https://civitai.com/models/682369?modelVersionId=764918
    thanks

  16. city96 commented on Aug 26, 2024

    @city96
    Owner

    @al-swaiti Seems to work fine. I've looked at Q4_K_M and the keys are all in the right format, and looks like the _M logic worked with some being in Q4_K and some other keys in Q5_K as intended. Two things to nitpick:

    • "Q_8" as you named it isn't a quantization. It should be called "Q8_0".
    • I think the license needs to be the same as the dev model if you used any part of the dev model instead of just "Apache2"

    Other than that it looks good to me, output is perfectly coherent. Thank you for following the guide and not disabling any of the key checks/messing with the format lol.

  17. al-swaiti commented on Aug 27, 2024

    @al-swaiti
    Author

    @city96
    okay ,,, about license i think you already know the flux dev is aura flow model trained on better quality dataset , is there any part of the model showed using of dev!? ,

  18. city96 commented on Aug 27, 2024

    @city96
    Owner

    okay ,,, about license i think you already know the flux dev is aura flow model trained on better quality dataset , is there any part of the model showed using of dev!? ,

    I'm not sure what you're saying here. Are you claiming you've merged AuraFlow with Flux1-schnell?

    could i apply it for normal comfyui models ?

    You could implement it as a custom node but windows support would most likely need extra work based on just the comments there. As this is a different quantization method, it's also out of scope for this repository/issues and you should do your own research on this topic.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions