Skip to content

llama: disable lazy tensor loading by default on iGPUs - #28326

Merged
0cc4m merged 2 commits into
masterfrom
0cc4m/lazy-mode-auto
Sep 8, 2026
Merged

0cc4m merged 2 commits into
masterfrom
0cc4m/lazy-mode-auto

Conversation

@0cc4m

@0cc4m 0cc4m commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Refactor lazy mode auto to mean "pick a probably good mode for your system", whereas the current auto behaviour (lazy load tensors > 4GiB) moves to lazy mode large, and lazy load all of these tensors becomes lazy mode all. This makes the meaning of auto consistent with how it works for load mode, and resolves the issue where we're back to mmap enabled by default on iGPUs that lose a lot of performance this way.

Fixes #28160

Please let me know what you think, or if you have better ideas.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, Claude wrote the code, I tested and reviewed.

@0cc4m
0cc4m requested review from a team, CISC and ggerganov as code owners September 3, 2026 13:29
@0cc4m 0cc4m changed the title llama: add lazy mode auto, fix iGPU regression llama: refactor lazy mode auto, fix iGPU regression Sep 3, 2026

@ORippler ORippler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't forget to update #9289 should this be merged

Comment thread src/llama-model.cpp
ggml_backend_dev_props props;
ggml_backend_dev_get_props(dev.dev, &props);
if (!props.caps.mmap_support) {
ml.lazy.mode = LLAMA_LAZY_MODE_OFF;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my understanding, this corresponds to --load-mode none?

@ngxson ngxson Sep 3, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

probably be useful to establish a table with one axis as --load-mode and one as --lazy-mode, to see which combination does what

edit: in the end, 2 modes stay independent, so such table is not necessary anymore

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we unify the two modes so that they take all the same settings / have all the same behaviours? It is a bit hard to reason about still since they are almost-but-not-always the same right now.

Comment thread src/llama-model.cpp Outdated
}
}

// resolve AUTO: LARGE where mmap is supported, else OFF (e.g. iGPUs); see #28160

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

may wish to expand #28160 to the full link

Comment thread common/arg.cpp Outdated
@ngxson

ngxson commented Sep 3, 2026 •

Copy link
Copy Markdown
Collaborator

lazy load all of these tensors becomes lazy mode all

note that it's not practically possible atm, only tensors marked as TENSOR_READ_LAZY can be lazy-read

and currently only PLE tensors are being marked this way, so all mode is redundant in practice

@ilariofebi

This comment was marked as duplicate.

@0cc4m

0cc4m commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

As discussed, I simplified it to only switch AUTO to mean OFF on iGPUs, no other change.

@0cc4m 0cc4m changed the title llama: refactor lazy mode auto, fix iGPU regression llama: don't enable lazy mode auto + mmap on iGPUs by default Sep 8, 2026
@0cc4m 0cc4m changed the title llama: don't enable lazy mode auto + mmap on iGPUs by default llama: disable lazy tensor loading by default on iGPUs Sep 8, 2026
@0cc4m
0cc4m merged commit f3f1a8f into master Sep 8, 2026
25 of 26 checks passed
@0cc4m
0cc4m deleted the 0cc4m/lazy-mode-auto branch September 8, 2026 16:05
@remeh

remeh commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

@0cc4m @pwilkin
It might be known but just in case it has been missed (and also for reference for people hitting this problem): when loading a Q4 quants of Qwen3.8-Flash-Next on a 128GB Strix Halo, with this PR in, it is now OOMing, with or without --no-mmap.
Obviously, when manually setting lazy-mode = on it loads correctly without OOMing, but just wanted to notice that the "default" case is now more prone to OOMs by default on some iGPUs hardware.

@oldgithubman

This comment was marked as abuse.

@0cc4m

0cc4m commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

A model that fits into your memory on an iGPU will now run as fast as it can. If you run a model that's larger than your available memory, you should reenable lazy-mode manually, yes.

@CISC

CISC commented Sep 9, 2026

Copy link
Copy Markdown
Member

Ideally autofitting should be made capable of figuring this out.

@remeh

remeh commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Is there a reason the 104GB Q4 GGUF isn't fitting a 124GB available memory slot (when lazy-mode is not set)? I don't think it is the kv-cache since it happens even with ctx-size = 8192.
Watching free -h, it seems that there is a first pass reserving ~80GB of memory, and after approximately 10-15s, it slowly creeps all the memory out until the OOM (again, when not setting lazy-mode). Probably not the good PR to ask about this, in all cases, it works forcing lazy-mode=on 👍

SteelPh0enix pushed a commit to SteelPh0enix/llama.cpp-qwen4exp that referenced this pull request Sep 9, 2026
* llama: add lazy mode auto, fix iGPU regression

* revert changes except disabling lazy load on iGPUs in AUTO

(cherry picked from commit f3f1a8f)
@oldgithubman

Copy link
Copy Markdown

A model that fits into your memory on an iGPU will now run as fast as it can. If you run a model that's larger than your available memory, you should reenable lazy-mode manually, yes.

I figured that out after llama.cpp crashed my system with a previously working config, yes.

x1250 pushed a commit to x1250/llama.cpp that referenced this pull request Sep 9, 2026
* llama: add lazy mode auto, fix iGPU regression

* revert changes except disabling lazy load on iGPUs in AUTO
@pwilkin

pwilkin commented Sep 9, 2026

Copy link
Copy Markdown
Member

I figured that out after llama.cpp crashed my system with a previously working config, yes.

As a general note, I really do recommend setting up a good oomkiller when dealing with local models. I can't remember how many times I crashed my system before I decided to set up one, but it really helps.

@cdanis

cdanis commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

As a general note, I really do recommend setting up a good oomkiller when dealing with local models. I can't remember how many times I crashed my system before I decided to set up one, but it really helps.

Big +1.

Some other inference engines set their oom score very high in main(), since it's common that e.g. Vulkan memory won't be accounted to the process. I've been meaning to send this as a PR honestly...

@oldgithubman

Copy link
Copy Markdown

I figured that out after llama.cpp crashed my system with a previously working config, yes.

As a general note, I really do recommend setting up a good oomkiller when dealing with local models. I can't remember how many times I crashed my system before I decided to set up one, but it really helps.

Thanks for the advice. I will do this.

zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* llama: add lazy mode auto, fix iGPU regression

* revert changes except disabling lazy load on iGPUs in AUTO
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* llama: add lazy mode auto, fix iGPU regression

* revert changes except disabling lazy load on iGPUs in AUTO
fencerJP pushed a commit to fencerJP/llama-apu that referenced this pull request Sep 16, 2026
* llama: add lazy mode auto, fix iGPU regression

* revert changes except disabling lazy load on iGPUs in AUTO
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* llama: add lazy mode auto, fix iGPU regression

* revert changes except disabling lazy load on iGPUs in AUTO
fencerJP added a commit to fencerJP/llama-apu that referenced this pull request Sep 24, 2026
…ing CLI, route-info

- apu-cli route-info: NPU/GPU capability detection (runtime-only XRT dlopen
  probe, /dev/accel + amdxdna driver, /dev/kfd), four-tier .xclbin discovery,
  memory-estimate gate, honest route resolution with explicit fallback reasons
- common: --tokenize/--prefill/--decode, --gpu-based/--cpu-based/--npu-based,
  --apu-xclbin, --apu-verbose; presets honor explicit -ngl; APU flags raise
  log threshold so NPU->GPU fallback is visible at default verbosity
- llama-model: log reason when lazy loading auto-resolves OFF on iGPU/unified
  memory (upstream IGPU mmap_support=false mechanism, issue ggml-org#28326)
@ORippler ORippler mentioned this pull request Sep 29, 2026
1 task
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Regression: --lazy-mode auto halves pp512 for qwen4exp on Vulkan (AMD iGPU)

10 participants