Repository navigation
Is GGUF over? #463
Description
Activity
reader = gguf.GGUFReader(path)Currently with GGUF, executing just a single line, takes 20 seconds for Gemma4 E2B. On top of that, it executes once more in the tokenizer loader.
m8rr@736fc9c
Optimized GGUFReader tokenizer array parsing, delivering a noticeable speedup in initial loading time.
(Further optimization is possible by skip list processing.
Lists are handled by the tokenizer generator, and the generated tokenizer is saved as a file cache.
If a tokenizer file already exists, it is retrieved and reused.)m8rr@7187555
Performance is further enhanced by reusing the existing GGUFReader, preventing unnecessary reloads.m8rr@1c113ea
Improves inference speed by torch-compiling dequantize_functions. In some cases, it delivers up to a 30% speedup.Reacted by tom-mReacted by tom-mReacted by kingp0dd and tom-mhttps://github.com/m8rr/ComfyUI-GGUF/tree/Dynamic-VRAM
Qwen3-VL Vision Support, gemma4-E2B Vision Support.
Qwen3-VL-4B is correctly recognized as Krea2TEModel_, and it will probably work without mmproj.
https://huggingface.co/realrebelai/KREA-2_GGUFs/tree/main/TURBOD:\AI\ComfyUI_windows_portable>.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --fast fp16_accumulation --use-sage-attention --disable-api-nodes -cache-ram 0 --fast-disk [INFO] setup plugin alembic.autogenerate...... [INFO] Found comfy_kitchen backend..... [INFO] Checkpoint files will always be loaded safely. [INFO] Total VRAM 12282 MB, total RAM 32085 MB [INFO] pytorch version: 2.12.1+cu130 [INFO] Enabled fp16 accumulation. [INFO] Set vram state to: NORMAL_VRAM [INFO] Device: cuda:0 NVIDIA GeForce RTX 4070 SUPER : cudaMallocAsync [INFO] Using async weight offloading with 2 streams [INFO] Enabled pinned memory 12834.0 [INFO] Using sage attention aimdo: ..... [INFO] DynamicVRAM support detected and enabled [INFO] Python version: 3.13.12 (tags/v3.13.12:1cbe481, Feb 3 2026, 18:22:25) [MSC v.1944 64 bit (AMD64)] [INFO] ComfyUI version: 0.26.0 [INFO] comfy-aimdo version: 0.4.10 [INFO] comfy-kitchen version: 0.2.10 [INFO] comfyui-frontend-package version: 1.45.19 [INFO] comfyui-workflow-templates version: 0.10.3 [INFO] comfyui-embedded-docs version: 0.5.5[INFO] got prompt [INFO] Using pytorch attention in VAE [INFO] Using pytorch attention in VAE [INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16 [INFO] got prompt [INFO] gguf qtypes: F32 (145), Q6_K (37), Q4_K (216) [WARNING] Dequantizing token_embd.weight to prevent runtime OOM. [INFO] Attenpting to find mmproj file for text encoder... [INFO] Using mmproj 'Qwen3-VL-4B-Instruct.mmproj-Q8_0.gguf' for text encoder 'Qwen3-VL-4B-Instruct-Q4_K_M.gguf'. [INFO] gguf qtypes: F32 (212), Q8_0 (104) [INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16 [INFO] Requested to load Krea2TEModel_ [INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 2808MB Staged. 0 patches attached. Force pre-loaded 243 weights: 1158 KB. [INFO] gguf qtypes: Q4_K (262), F32 (166), F16 (2) [INFO] model weight dtype torch.float16, manual cast: None [INFO] model_type FLUX [INFO] Requested to load Krea2 [INFO] Model Krea2 prepared for dynamic VRAM loading. 6877MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB. 100%|████████████████████████████████████████████████████████████████████████████████████| 8/8 [00:11<00:00, 1.49s/it] [INFO] Requested to load WanVAE [INFO] 0 models unloaded. [INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB. [INFO] Prompt executed in 25.12 seconds [INFO] Model Krea2 prepared for dynamic VRAM loading. 6877MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB. 100%|████████████████████████████████████████████████████████████████████████████████████| 8/8 [00:12<00:00, 1.52s/it] [INFO] 0 models unloaded. [INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB. [INFO] Prompt executed in 12.60 secondsComfyUI now supports INT8. Is GGUF really dead?
https://huggingface.co/silveroxides/models
4070s, windows
I tested Ideogram4 INT8 with Turbo LoRA, generating 768x1280 images in 6 steps.
GGUF Q4K, 5.7G, 12s
FP8, 9G, 6s
INT8 9.3G 5skrea2
Q4KM, 7G, 10s
INT8 14G, 8sIs the gap narrowing because INT8 exceeds the VRAM, or is the BF16 of Ideogram 4 GGUF being handled incorrectly?
The "comfyanonymous" dev always talks about performance when making changes like this but not about quality. I have yet to see anyone produce data that shows any version of FP8 beating or equaling Q8_0 in perplexity. Performance is meaningless if the quality is inferior.
@SRStwo I always use Turbo models and treat it like a slot machine, but I totally agree with you. There definitely should be an option for higher quality.
I don't think it's over. GGUF was always slower but fidelity and compatibility is better. FP8 loses on quality time after time. FP4 is better but you better have
$$blackwell$$ . Int8 is good if you only want that size. GGUF can be sped up a few % with cublas too.. but I gotta make it a checkbox. Especially for stuff like text encoders, there really is no other sane option.I tried int8 ltx2.3 but all the models won't fit in my 32gb ram 12gb vram so it reloads from disk, making GGUF faster. Pretty sure there are other similar setups like mine, which still benefit from gguf.
Unless I'm doing something wrong with int8.
Reacted by tom-mReacted by tom-mI tested LTX INT8 (20G). Both the loading and generation speeds were faster than GGUF (17G). Shared GPU memory usage increased by pretty much the difference in their sizes.
768x1280x129, 6/3 steps, fast fp16, sage INT8 [INFO] got prompt [INFO] Model LTXAV prepared for dynamic VRAM loading. 20486MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB. 100%|████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:10<00:00, 1.77s/it] [INFO] 0 models unloaded. [INFO] Model LTXAV prepared for dynamic VRAM loading. 20486MB Staged. 0 patches attached. Force pre-loaded 608 weights: 3303 KB. 100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:17<00:00, 5.73s/it] [INFO] Requested to load AudioVAE [INFO] loaded completely; 693.46 MB loaded, full load: True [INFO] 0 models unloaded. [INFO] Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached. [INFO] Prompt executed in 46.15 seconds GGUF [INFO] got prompt [INFO] Model LTXAV prepared for dynamic VRAM loading. 16915MB Staged. 0 patches attached. Force pre-loaded 608 weights: 6567 KB. 100%|████████████████████████████████████████████████████████████████████████████████████| 6/6 [00:11<00:00, 1.94s/it] [INFO] 0 models unloaded. [INFO] Model LTXAV prepared for dynamic VRAM loading. 16915MB Staged. 0 patches attached. Force pre-loaded 608 weights: 6567 KB. 100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:22<00:00, 7.51s/it] [INFO] Requested to load AudioVAE [INFO] loaded completely; 693.46 MB loaded, full load: True [INFO] 0 models unloaded. [INFO] Model VideoVAE prepared for dynamic VRAM loading. 1384MB Staged. 0 patches attached. [INFO] Prompt executed in 52.20 seconds
Looking forward to INT4.
INT8 has higher perplexity than Q8_0 GGUF, as I recall. I don't recall how much of a difference there was, though. I think INT8 was better than FP8. What I'd like to see is Q8_1 GGUF to see if it's worthwhile to get the lowest possible perplexity (when compared with 16-bit) without the size (particularly VRAM use) of 16-bit. However, although that format was described early on, it looks like no one bothered to do anything with it for consumer image/video models.
There is a Wan checkpoint maker who adamantly evangelizes FP8 over "outdated" GGUF and yet his/her most recent release is much larger, apparently because even more 16-bit data has been put into it to try to compensate for the perplexity of FP8, which I find a bit droll.
Reacted by gdavidson75I thought you could quantize to Q8_1 if you want to. You need it's numerical ID.
I thought you could quantize to Q8_1 if you want to. You need it's numerical ID.
I've looked everywhere and I've never seen any image/video model offered in anything above Q8_0. Even if it's possible no one has done it.
Someone on Reddit said there is also a problem with the nomenclature because apparently something unrelated has already taken the Q8_1 name. I think they suggested using Q8_1_L or Q8_0_L instead; I don't recall exactly which was suggested. Personally, I think using the Q8_1 name is okay as long as it's in the context of that being offered alongside the various others (Q8_0, Q6_K, etc.)
The name doesn't matter so much, there are only so many tensor quant types in the header where it lists them. I remember trying a model like this before and not seeing much in the way of quality improvements. Pretty easy to replicate if you feed it the correct numerical value. Just because nobody released it, doesn't mean you can't quantize yourself.
The name doesn't matter so much, there are only so many tensor quant types in the header where it lists them. I remember trying a model like this before and not seeing much in the way of quality improvements. Pretty easy to replicate if you feed it the correct numerical value. Just because nobody released it, doesn't mean you can't quantize yourself.
I'd be interested to see how its perplexity compares (to FP32, BF16/FP16, Q8_0, INT8, FP8, and one of the Q6 variants). The idea is to get the closest possible quality to 16-bit without the VRAM usage. I know Q8_0 is supposed to be close but even closer might have a use case I think.
I don't think I ever measured perplexity on an image model, only on LLM.
I don't think I ever measured perplexity on an image model, only on LLM.
Where I can see it being particularly important is with video models like Wan 2.2. That can be run with 12 GB of VRAM via Q8_0 but the quality doesn't seem to be quite at the level of 16-bit. If Q8_1 were to narrow the gap further it might be useful for low-VRAM systems.
6 remaining items
Thanks mate I'll give this a try.
For the record and for anyone following, here's my comfy versions
OS
linux
Python Version
3.13.11 | packaged by Anaconda, Inc. | (main, Dec 10 2025, 21:28:48) [GCC 14.3.0]
Embedded Python
false
Pytorch Version
2.12.0+cu130
Arguments
main.py --listen 0.0.0.0 --reserve-vram=0.05
RAM Total
31.28 GB
RAM Free
14.47 GB
Templates Version
0.11.6
Devices
Name
cuda:0 NVIDIA GeForce RTX 4070 : cudaMallocAsync
Type
cuda
VRAM Total
11.6 GB
VRAM Free
8.22 GB
Torch VRAM Total
96 MB
Torch VRAM Free
59.87 MBTE isn't being reset, but it's true that it's reading from the drive. Since LTX files are large, that makes sense.
These are the results from krea2. Memory usage is almost the same and there seems to be enough RAM, but due to a 4G difference in TE, BF16 is being read from the drive.
UNET INT8 13G
TE INT8 4.7G
TE BF16 8.6GTried to install the same versions as yours (pytorch, comfy) and used your workflow but TE still doesn't want to get pinned. Only difference now i think is just the OS. Weirdly though in the issue thread i linked, they're using Windows. Will think of other differences.
Thanks for sharing!
[INFO] Total VRAM 11873 MB, total RAM 32027 MB
[INFO] pytorch version: 2.13.0+cu130
[INFO] Enabled fp16 accumulation.
[INFO] Set vram state to: NORMAL_VRAM
[INFO] Device: cuda:0 NVIDIA GeForce RTX 4070 : native
[INFO] Using async weight offloading with 2 streams
[INFO] Enabled pinned memory 28823.0
[INFO] Using sage attention
aimdo: /project/src-posix/cuda-funchooks.c:52:DEBUG:aimdo_setup_hooks: hooks succe
ssfully installed
aimdo: /project/src/control.c:247:INFO:comfy-aimdo inited for GPU: NVIDIA GeForce
RTX 4070 (VRAM: 11873 MB)
[INFO] DynamicVRAM support detected and enabled
[INFO] Python version: 3.13.11 | packaged by Anaconda, Inc. | (main, Dec 10 2025,
21:28:48) [GCC 14.3.0]
[INFO] ComfyUI version: 0.27.0
[INFO] comfy-aimdo version: 0.4.10
[INFO] comfy-kitchen version: 0.2.18
[INFO] comfyui-frontend-package version: 1.45.20
[INFO] comfyui-workflow-templates version: 0.11.6
[INFO] comfyui-embedded-docs version: 0.5.7
[INFO] comfy-kitchen version: 0.2.18
[INFO] comfy-aimdo version: 0.4.10TE isn't being reset, but it's true that it's reading from the drive. Since LTX files are large, that makes sense.
But your ltx (UNET INT8(20G)+TE INT8(13G).) is not being re-read from your disk right?
But your ltx (UNET INT8(20G)+TE INT8(13G).) is not being re-read from your disk right?
@kingp0dd It looks like TE INT8 is being read almost entirely from the drive, whereas TE GGUF barely reads from the drive at all, if ever. Compare the initial loading state with the state where both samplers are bypassed (TE+VAE only).
It looks like they're both being read from disk but since gguf is smaller, hence the disk activity is shorter? Does it seem like you're also experiencing issue wherein TE is not pinned and reloaded from disk
It looks like they're both being read from disk but since gguf is smaller, hence the disk activity is shorter?
That's possible.
To sum it up: I changed the prompt for all of them, ran them 3 times each, then connected 'Preview as Text' to the 'CLIP Text Encode' and ran just the CLIP.krea2 INT8 (13G), TE INT8 (4.7G), there was no disk activity, so I skipped it (or maybe it was just too fast to notice).
krea2 INT8 (13G), TE BF16 (8.6G)

@m8rr it seems you also experience the same disk reloading issue. BTW i also tried int4 ltx and int4 TE which should both fit in RAM, but they also reload from disk.
Right now, GGUF Q4KM is faster for me (and better quality than int4) but have to use --high-ram. Without --high-ram, comfy only uses ~85% of RAM.
@m8rr using new pull, my models seem to stay pinned now.
@kingp0dd The PR was officially merged and I tested it.
The previous issue with TE disk loading in the krea2 Unet INT8 + TE BF16 combination is gone. It's using all of the RAM instead of leaving it empty.LTX GGUF+GGUF, INT8+GGUF, and INT8+INT8 combinations didn't show any changes. With 32GB of RAM, it's probably still too large to make a difference.
Anyway, it's improved.
- @m8rr same with me. i replaced my ssd with a faster nvme, i realized that if disk transfer is taken out of equation (too fast), then int8 is as fast as GGUF Q4KM (so now i have quality > GGUF Q8 but at the same speed). Before, GGUF Q4KM was faster than int8 on my setup because my ssd was slow and since int8 is larger, it significantly affects the total gen time.…On Wed, Jul 29, 2026 at 10:08 AM m8rr ***@***.***> wrote: *m8rr* left a comment (city96/ComfyUI-GGUF#463) <#463 (comment)> @kingp0dd <https://github.com/kingp0dd> The PR was officially merged and I tested it. The previous issue with TE disk loading in the krea2 Unet INT8 + TE BF16 combination is gone. It's using all of the RAM instead of leaving it empty. LTX GGUF+GGUF, INT8+GGUF, and INT8+INT8 combinations didn't show any changes. With 32GB of RAM, it's probably still too large to make a difference. Anyway, it's improved. — Reply to this email directly, view it on GitHub <#463?email_source=notifications&email_token=ACGD6KSX3MSSEZ2JLH7DNQ35HFMAHA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKMJRGE4TMNBWGEZKM4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-5111964612>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/ACGD6KV5FBES3YCIDS2NP6L5HFMAHAVCNFSNUABFKJSXA33TNF2G64TZHM4DIMRXGY4DQNZQHNEXG43VMU5TINZSGMZDQNZTG422C5QC> . Triage notifications, keep track of coding agent tasks and review pull requests on the go with GitHub Mobile for iOS <https://github.com/notifications/mobile/ios/ACGD6KSBOZDT73CSWLVSCLT5HFMAHA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKMJRGE4TMNBWGEZKM4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJKTGN5XXIZLSL5UW64Y> and Android <https://github.com/notifications/mobile/android/ACGD6KSTFK7GO23BEYXFZWL5HFMAHA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKMJRGE4TMNBWGEZKM4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLTGN5XXIZLSL5QW4ZDSN5UWI>. Download it today! You are receiving this because you were mentioned.Message ID: ***@***.***>
I am still turning dynamic vram off. All it gives me is overhead. Whether I use GGUF or int8 or anything. This whole time, from all this effort, the only thing it provided was slowdown or compile breakage. Agree that it's finally not massively detrimental but it's no help either.
- In my case, dynamic vram is faster. I also wouldn't be able to run these big ltx models without it.…On Wed, Jul 29, 2026, 8:47 PM Forkoz ***@***.***> wrote: *Ph0rk0z* left a comment (city96/ComfyUI-GGUF#463) <#463 (comment)> I am still turning dynamic vram off. All it gives me is overhead. Whether I use GGUF or int8 or anything. This whole time, from all this effort, the only thing it provided was slowdown or compile breakage. Agree that it's finally not massively detrimental but it's no help either. — Reply to this email directly, view it on GitHub <#463?email_source=notifications&email_token=ACGD6KSX5L4PF7YMMWOWSK35HHW7TA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKMJRG44TINJSGEZ2M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-5117945213>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/ACGD6KVHX6EBHXHECKAJXGD5HHW7TAVCNFSNUABFKJSXA33TNF2G64TZHM4DIMRXGY4DQNZQHNEXG43VMU5TINZSGMZDQNZTG422C5QC> . Triage notifications, keep track of coding agent tasks and review pull requests on the go with GitHub Mobile for iOS <https://github.com/notifications/mobile/ios/ACGD6KQ33A4M6ASOMHGTAOL5HHW7TA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKMJRG44TINJSGEZ2M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJKTGN5XXIZLSL5UW64Y> and Android <https://github.com/notifications/mobile/android/ACGD6KVQUEBNKL2WTWAXLUL5HHW7TA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKMJRG44TINJSGEZ2M4TFMFZW63VHNVSW45DJN5XKKZLWMVXHJLTGN5XXIZLSL5QW4ZDSN5UWI>. Download it today! You are receiving this because you were mentioned.Message ID: ***@***.***>










Native formats are definitely faster, though my hard drive would beg to differ.