Skip to content

Is GGUF over? #463

Description

@m8rr

Native formats are definitely faster, though my hard drive would beg to differ.

Image "If you use gguf we recommend keeping dynamic vram enabled and using native ComfyUI model formats instead."

Activity

  1. m8rr commented on Jun 23, 2026

    @m8rr
    Author
    reader = gguf.GGUFReader(path)
    

    Currently with GGUF, executing just a single line, takes 20 seconds for Gemma4 E2B. On top of that, it executes once more in the tokenizer loader.

    m8rr@736fc9c
    Optimized GGUFReader tokenizer array parsing, delivering a noticeable speedup in initial loading time.
    (Further optimization is possible by skip list processing.
    Lists are handled by the tokenizer generator, and the generated tokenizer is saved as a file cache.
    If a tokenizer file already exists, it is retrieved and reused.)

    m8rr@7187555
    Performance is further enhanced by reusing the existing GGUFReader, preventing unnecessary reloads.

    m8rr@1c113ea
    Improves inference speed by torch-compiling dequantize_functions. In some cases, it delivers up to a 30% speedup.

  2. m8rr commented on Jun 24, 2026

    @m8rr
    Author

    https://github.com/m8rr/ComfyUI-GGUF/tree/Dynamic-VRAM
    Qwen3-VL Vision Support, gemma4-E2B Vision Support.
    Qwen3-VL-4B is correctly recognized as Krea2TEModel_, and it will probably work without mmproj.
    https://huggingface.co/realrebelai/KREA-2_GGUFs/tree/main/TURBO

    D:\AI\ComfyUI_windows_portable>.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --fast fp16_accumulation --use-sage-attention --disable-api-nodes -cache-ram 0 --fast-disk
    [INFO] setup plugin alembic.autogenerate......
    [INFO] Found comfy_kitchen backend.....
    [INFO] Checkpoint files will always be loaded safely.
    [INFO] Total VRAM 12282 MB, total RAM 32085 MB
    [INFO] pytorch version: 2.12.1+cu130
    [INFO] Enabled fp16 accumulation.
    [INFO] Set vram state to: NORMAL_VRAM
    [INFO] Device: cuda:0 NVIDIA GeForce RTX 4070 SUPER : cudaMallocAsync
    [INFO] Using async weight offloading with 2 streams
    [INFO] Enabled pinned memory 12834.0
    [INFO] Using sage attention
    aimdo: .....
    [INFO] DynamicVRAM support detected and enabled
    [INFO] Python version: 3.13.12 (tags/v3.13.12:1cbe481, Feb  3 2026, 18:22:25) [MSC v.1944 64 bit (AMD64)]
    [INFO] ComfyUI version: 0.26.0
    [INFO] comfy-aimdo version: 0.4.10
    [INFO] comfy-kitchen version: 0.2.10
    [INFO] comfyui-frontend-package version: 1.45.19
    [INFO] comfyui-workflow-templates version: 0.10.3
    [INFO] comfyui-embedded-docs version: 0.5.5
    
    [INFO] got prompt
    [INFO] Using pytorch attention in VAE
    [INFO] Using pytorch attention in VAE
    [INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
    [INFO] got prompt
    [INFO] gguf qtypes: F32 (145), Q6_K (37), Q4_K (216)
    [WARNING] Dequantizing token_embd.weight to prevent runtime OOM.
    [INFO] Attenpting to find mmproj file for text encoder...
    [INFO] Using mmproj 'Qwen3-VL-4B-Instruct.mmproj-Q8_0.gguf' for text encoder 'Qwen3-VL-4B-Instruct-Q4_K_M.gguf'.
    [INFO] gguf qtypes: F32 (212), Q8_0 (104)
    [INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
    [INFO] Requested to load Krea2TEModel_
    [INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 2808MB Staged. 0 patches attached. Force pre-loaded 243 weights: 1158 KB.
    [INFO] gguf qtypes: Q4_K (262), F32 (166), F16 (2)
    [INFO] model weight dtype torch.float16, manual cast: None
    [INFO] model_type FLUX
    [INFO] Requested to load Krea2
    [INFO] Model Krea2 prepared for dynamic VRAM loading. 6877MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
    100%|████████████████████████████████████████████████████████████████████████████████████| 8/8 [00:11<00:00,  1.49s/it]
    [INFO] Requested to load WanVAE
    [INFO] 0 models unloaded.
    [INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
    [INFO] Prompt executed in 25.12 seconds
    [INFO] Model Krea2 prepared for dynamic VRAM loading. 6877MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
    100%|████████████████████████████████████████████████████████████████████████████████████| 8/8 [00:12<00:00,  1.52s/it]
    [INFO] 0 models unloaded.
    [INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
    [INFO] Prompt executed in 12.60 seconds
    
  3. m8rr commented on Jun 24, 2026

    @m8rr
    Author

    ComfyUI now supports INT8. Is GGUF really dead?

    https://huggingface.co/silveroxides/models

    4070s, windows
    I tested Ideogram4 INT8 with Turbo LoRA, generating 768x1280 images in 6 steps.
    GGUF Q4K, 5.7G, 12s
    FP8, 9G, 6s
    INT8 9.3G 5s

    krea2
    Q4KM, 7G, 10s
    INT8 14G, 8s

    Is the gap narrowing because INT8 exceeds the VRAM, or is the BF16 of Ideogram 4 GGUF being handled incorrectly?

  4. SRStwo commented on Jun 27, 2026

    @SRStwo

    The "comfyanonymous" dev always talks about performance when making changes like this but not about quality. I have yet to see anyone produce data that shows any version of FP8 beating or equaling Q8_0 in perplexity. Performance is meaningless if the quality is inferior.

  5. m8rr commented on Jun 28, 2026

    @m8rr
    Author

    @SRStwo I always use Turbo models and treat it like a slot machine, but I totally agree with you. There definitely should be an option for higher quality.

  6. Ph0rk0z commented on Jun 28, 2026

    @Ph0rk0z

    I don't think it's over. GGUF was always slower but fidelity and compatibility is better. FP8 loses on quality time after time. FP4 is better but you better have $$blackwell$$. Int8 is good if you only want that size. GGUF can be sped up a few % with cublas too.. but I gotta make it a checkbox. Especially for stuff like text encoders, there really is no other sane option.

  7. kingp0dd commented on Jun 29, 2026

    @kingp0dd

    I tried int8 ltx2.3 but all the models won't fit in my 32gb ram 12gb vram so it reloads from disk, making GGUF faster. Pretty sure there are other similar setups like mine, which still benefit from gguf.

    Unless I'm doing something wrong with int8.

  8. m8rr commented on Jul 5, 2026

    @m8rr
    Author

    https://huggingface.co/Kijai/LTX2.3_comfy/blob/main/diffusion_models/ltx-2.3-22b-distilled-1.1_transformer_only_int8_convrot.safetensors

    I tested LTX INT8 (20G). Both the loading and generation speeds were faster than GGUF (17G). Shared GPU memory usage increased by pretty much the difference in their sizes.

    768x1280x129, 6/3 steps, fast fp16, sage
    INT8
    
    [INFO] got prompt
    [INFO] Model LTXAV prepared for dynamic VRAM loading. 20486MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB.
    100%|████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:10<00:00,  1.77s/it]
    [INFO] 0 models unloaded.
    [INFO] Model LTXAV prepared for dynamic VRAM loading. 20486MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB.
    100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:17<00:00,  5.73s/it]
    [INFO] Requested to load AudioVAE
    [INFO] loaded completely;  693.46 MB loaded, full load: True
    [INFO] 0 models unloaded.
    [INFO] Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
    [INFO] Prompt executed in 46.15 seconds
    
    
    
    GGUF
    
    [INFO] got prompt
    [INFO] Model LTXAV prepared for dynamic VRAM loading. 16915MB Staged. 0 patches attached. Force pre-loaded 608 weights: 6567 KB.
    100%|████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:11<00:00,  1.94s/it]
    [INFO] 0 models unloaded.
    [INFO] Model LTXAV prepared for dynamic VRAM loading. 16915MB Staged. 0 patches attached. Force pre-loaded 608 weights: 6567 KB.
    100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:22<00:00,  7.51s/it]
    [INFO] Requested to load AudioVAE
    [INFO] loaded completely;  693.46 MB loaded, full load: True
    [INFO] 0 models unloaded.
    [INFO] Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached.
    [INFO] Prompt executed in 52.20 seconds
    
    Image Image

    Looking forward to INT4.

  9. SRStwo commented on Jul 5, 2026

    @SRStwo

    INT8 has higher perplexity than Q8_0 GGUF, as I recall. I don't recall how much of a difference there was, though. I think INT8 was better than FP8. What I'd like to see is Q8_1 GGUF to see if it's worthwhile to get the lowest possible perplexity (when compared with 16-bit) without the size (particularly VRAM use) of 16-bit. However, although that format was described early on, it looks like no one bothered to do anything with it for consumer image/video models.

    There is a Wan checkpoint maker who adamantly evangelizes FP8 over "outdated" GGUF and yet his/her most recent release is much larger, apparently because even more 16-bit data has been put into it to try to compensate for the perplexity of FP8, which I find a bit droll.

  10. Ph0rk0z commented on Jul 7, 2026

    @Ph0rk0z

    I thought you could quantize to Q8_1 if you want to. You need it's numerical ID.

  11. SRStwo commented on Jul 7, 2026

    @SRStwo

    I thought you could quantize to Q8_1 if you want to. You need it's numerical ID.

    I've looked everywhere and I've never seen any image/video model offered in anything above Q8_0. Even if it's possible no one has done it.

    Someone on Reddit said there is also a problem with the nomenclature because apparently something unrelated has already taken the Q8_1 name. I think they suggested using Q8_1_L or Q8_0_L instead; I don't recall exactly which was suggested. Personally, I think using the Q8_1 name is okay as long as it's in the context of that being offered alongside the various others (Q8_0, Q6_K, etc.)

  12. Ph0rk0z commented on Jul 7, 2026

    @Ph0rk0z

    The name doesn't matter so much, there are only so many tensor quant types in the header where it lists them. I remember trying a model like this before and not seeing much in the way of quality improvements. Pretty easy to replicate if you feed it the correct numerical value. Just because nobody released it, doesn't mean you can't quantize yourself.

  13. SRStwo commented on Jul 8, 2026

    @SRStwo

    The name doesn't matter so much, there are only so many tensor quant types in the header where it lists them. I remember trying a model like this before and not seeing much in the way of quality improvements. Pretty easy to replicate if you feed it the correct numerical value. Just because nobody released it, doesn't mean you can't quantize yourself.

    I'd be interested to see how its perplexity compares (to FP32, BF16/FP16, Q8_0, INT8, FP8, and one of the Q6 variants). The idea is to get the closest possible quality to 16-bit without the VRAM usage. I know Q8_0 is supposed to be close but even closer might have a use case I think.

  14. Ph0rk0z commented on Jul 8, 2026

    @Ph0rk0z

    I don't think I ever measured perplexity on an image model, only on LLM.

  15. SRStwo commented on Jul 8, 2026

    @SRStwo

    I don't think I ever measured perplexity on an image model, only on LLM.

    Where I can see it being particularly important is with video models like Wan 2.2. That can be run with 12 GB of VRAM via Q8_0 but the quality doesn't seem to be quite at the level of 16-bit. If Q8_1 were to narrow the gap further it might be useful for low-VRAM systems.

  16. 6 remaining items

  17. kingp0dd commented on Jul 13, 2026

    @kingp0dd

    Thanks mate I'll give this a try.

    For the record and for anyone following, here's my comfy versions

    OS
    linux
    Python Version
    3.13.11 | packaged by Anaconda, Inc. | (main, Dec 10 2025, 21:28:48) [GCC 14.3.0]
    Embedded Python
    false
    Pytorch Version
    2.12.0+cu130
    Arguments
    main.py --listen 0.0.0.0 --reserve-vram=0.05
    RAM Total
    31.28 GB
    RAM Free
    14.47 GB
    Templates Version
    0.11.6
    Devices
    Name
    cuda:0 NVIDIA GeForce RTX 4070 : cudaMallocAsync
    Type
    cuda
    VRAM Total
    11.6 GB
    VRAM Free
    8.22 GB
    Torch VRAM Total
    96 MB
    Torch VRAM Free
    59.87 MB

  18. m8rr commented on Jul 13, 2026

    @m8rr
    Author

    TE isn't being reset, but it's true that it's reading from the drive. Since LTX files are large, that makes sense.
    These are the results from krea2. Memory usage is almost the same and there seems to be enough RAM, but due to a 4G difference in TE, BF16 is being read from the drive.
    UNET INT8 13G
    TE INT8 4.7G
    TE BF16 8.6G

    TE INT8 4.7G
    Image
    TE BF16 8.6G
    Image

  19. kingp0dd commented on Jul 13, 2026

    @kingp0dd

    Tried to install the same versions as yours (pytorch, comfy) and used your workflow but TE still doesn't want to get pinned. Only difference now i think is just the OS. Weirdly though in the issue thread i linked, they're using Windows. Will think of other differences.

    Thanks for sharing!

    Image

    [INFO] Total VRAM 11873 MB, total RAM 32027 MB
    [INFO] pytorch version: 2.13.0+cu130
    [INFO] Enabled fp16 accumulation.
    [INFO] Set vram state to: NORMAL_VRAM
    [INFO] Device: cuda:0 NVIDIA GeForce RTX 4070 : native
    [INFO] Using async weight offloading with 2 streams
    [INFO] Enabled pinned memory 28823.0
    [INFO] Using sage attention
    aimdo: /project/src-posix/cuda-funchooks.c:52:DEBUG:aimdo_setup_hooks: hooks succe
    ssfully installed
    aimdo: /project/src/control.c:247:INFO:comfy-aimdo inited for GPU: NVIDIA GeForce
    RTX 4070 (VRAM: 11873 MB)
    [INFO] DynamicVRAM support detected and enabled
    [INFO] Python version: 3.13.11 | packaged by Anaconda, Inc. | (main, Dec 10 2025,
    21:28:48) [GCC 14.3.0]
    [INFO] ComfyUI version: 0.27.0
    [INFO] comfy-aimdo version: 0.4.10
    [INFO] comfy-kitchen version: 0.2.18
    [INFO] comfyui-frontend-package version: 1.45.20
    [INFO] comfyui-workflow-templates version: 0.11.6
    [INFO] comfyui-embedded-docs version: 0.5.7
    [INFO] comfy-kitchen version: 0.2.18
    [INFO] comfy-aimdo version: 0.4.10

  20. kingp0dd commented on Jul 13, 2026

    @kingp0dd

    TE isn't being reset, but it's true that it's reading from the drive. Since LTX files are large, that makes sense.

    But your ltx (UNET INT8(20G)+TE INT8(13G).) is not being re-read from your disk right?

  21. m8rr commented on Jul 13, 2026

    @m8rr
    Author

    But your ltx (UNET INT8(20G)+TE INT8(13G).) is not being re-read from your disk right?

    @kingp0dd It looks like TE INT8 is being read almost entirely from the drive, whereas TE GGUF barely reads from the drive at all, if ever. Compare the initial loading state with the state where both samplers are bypassed (TE+VAE only).

    Initial loading for INT8+INT8
    Image

    TE INT8
    Image

    TE GGUF
    Image

  22. kingp0dd commented on Jul 13, 2026

    @kingp0dd

    It looks like they're both being read from disk but since gguf is smaller, hence the disk activity is shorter? Does it seem like you're also experiencing issue wherein TE is not pinned and reloaded from disk

  23. m8rr commented on Jul 13, 2026

    @m8rr
    Author

    It looks like they're both being read from disk but since gguf is smaller, hence the disk activity is shorter?

    That's possible.
    To sum it up: I changed the prompt for all of them, ran them 3 times each, then connected 'Preview as Text' to the 'CLIP Text Encode' and ran just the CLIP.

    krea2 INT8 (13G), TE INT8 (4.7G), there was no disk activity, so I skipped it (or maybe it was just too fast to notice).

    krea2 INT8 (13G), TE BF16 (8.6G)
    Image

    LTX INT8 (20G), TE GGUF (7.2G)
    Image

    LTX INT8 (20G), TE INT8 (13G)
    Image

  24. m8rr commented on Jul 13, 2026

    @m8rr
    Author

    krea2 INT8 (13G), TE INT8 (4.7G), No drive activity
    Image

    krea2 INT8 (13G), TE BF16 (8.6G), TE drive loading
    Image

    LTX INT8 (20G), TE GGUF (7.2G), No TE drive loading, but loading text projection
    Image

    LTX INT8 (20G), TE INT8 (13G), TE drive loading is too obvious.

  25. kingp0dd commented on Jul 20, 2026

    @kingp0dd

    @m8rr it seems you also experience the same disk reloading issue. BTW i also tried int4 ltx and int4 TE which should both fit in RAM, but they also reload from disk.

    Right now, GGUF Q4KM is faster for me (and better quality than int4) but have to use --high-ram. Without --high-ram, comfy only uses ~85% of RAM.

  26. kingp0dd commented on Jul 23, 2026

    @kingp0dd

    @m8rr using new pull, my models seem to stay pinned now.

    Comfy-Org/ComfyUI#14705 (comment)

  27. m8rr commented on Jul 29, 2026

    @m8rr
    Author

    @kingp0dd The PR was officially merged and I tested it.
    The previous issue with TE disk loading in the krea2 Unet INT8 + TE BF16 combination is gone. It's using all of the RAM instead of leaving it empty.

    LTX GGUF+GGUF, INT8+GGUF, and INT8+INT8 combinations didn't show any changes. With 32GB of RAM, it's probably still too large to make a difference.

    Anyway, it's improved.

  28. kingp0dd commented on Jul 29, 2026

    @kingp0dd
  29. Ph0rk0z commented on Jul 29, 2026

    @Ph0rk0z

    I am still turning dynamic vram off. All it gives me is overhead. Whether I use GGUF or int8 or anything. This whole time, from all this effort, the only thing it provided was slowdown or compile breakage. Agree that it's finally not massively detrimental but it's no help either.

  30. kingp0dd commented on Jul 29, 2026

    @kingp0dd
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions