llama: disable lazy tensor loading by default on iGPUs - #28326
Conversation
| ggml_backend_dev_props props; | ||
| ggml_backend_dev_get_props(dev.dev, &props); | ||
| if (!props.caps.mmap_support) { | ||
| ml.lazy.mode = LLAMA_LAZY_MODE_OFF; |
There was a problem hiding this comment.
For my understanding, this corresponds to --load-mode none?
There was a problem hiding this comment.
probably be useful to establish a table with one axis as --load-mode and one as --lazy-mode, to see which combination does what
edit: in the end, 2 modes stay independent, so such table is not necessary anymore
There was a problem hiding this comment.
Should we unify the two modes so that they take all the same settings / have all the same behaviours? It is a bit hard to reason about still since they are almost-but-not-always the same right now.
| } | ||
| } | ||
|
|
||
| // resolve AUTO: LARGE where mmap is supported, else OFF (e.g. iGPUs); see #28160 |
note that it's not practically possible atm, only tensors marked as and currently only PLE tensors are being marked this way, so |
This comment was marked as duplicate.
This comment was marked as duplicate.
|
As discussed, I simplified it to only switch AUTO to mean OFF on iGPUs, no other change. |
|
@0cc4m @pwilkin |
This comment was marked as abuse.
This comment was marked as abuse.
|
A model that fits into your memory on an iGPU will now run as fast as it can. If you run a model that's larger than your available memory, you should reenable lazy-mode manually, yes. |
|
Ideally autofitting should be made capable of figuring this out. |
|
Is there a reason the 104GB Q4 GGUF isn't fitting a 124GB available memory slot (when |
* llama: add lazy mode auto, fix iGPU regression * revert changes except disabling lazy load on iGPUs in AUTO (cherry picked from commit f3f1a8f)
I figured that out after llama.cpp crashed my system with a previously working config, yes. |
* llama: add lazy mode auto, fix iGPU regression * revert changes except disabling lazy load on iGPUs in AUTO
As a general note, I really do recommend setting up a good oomkiller when dealing with local models. I can't remember how many times I crashed my system before I decided to set up one, but it really helps. |
Big +1. Some other inference engines set their oom score very high in main(), since it's common that e.g. Vulkan memory won't be accounted to the process. I've been meaning to send this as a PR honestly... |
Thanks for the advice. I will do this. |
* llama: add lazy mode auto, fix iGPU regression * revert changes except disabling lazy load on iGPUs in AUTO
* llama: add lazy mode auto, fix iGPU regression * revert changes except disabling lazy load on iGPUs in AUTO
* llama: add lazy mode auto, fix iGPU regression * revert changes except disabling lazy load on iGPUs in AUTO
* llama: add lazy mode auto, fix iGPU regression * revert changes except disabling lazy load on iGPUs in AUTO
…ing CLI, route-info - apu-cli route-info: NPU/GPU capability detection (runtime-only XRT dlopen probe, /dev/accel + amdxdna driver, /dev/kfd), four-tier .xclbin discovery, memory-estimate gate, honest route resolution with explicit fallback reasons - common: --tokenize/--prefill/--decode, --gpu-based/--cpu-based/--npu-based, --apu-xclbin, --apu-verbose; presets honor explicit -ngl; APU flags raise log threshold so NPU->GPU fallback is visible at default verbosity - llama-model: log reason when lazy loading auto-resolves OFF on iGPU/unified memory (upstream IGPU mmap_support=false mechanism, issue ggml-org#28326)
Overview
Refactor lazy mode
autoto mean "pick a probably good mode for your system", whereas the currentautobehaviour (lazy load tensors > 4GiB) moves to lazy modelarge, and lazy load all of these tensors becomes lazy modeall. This makes the meaning ofautoconsistent with how it works for load mode, and resolves the issue where we're back to mmap enabled by default on iGPUs that lose a lot of performance this way.Fixes #28160
Please let me know what you think, or if you have better ideas.
Requirements