Skip to content

Avoid a second full-size copy of each tensor with direct-io - #29749

Merged
ORippler merged 1 commit into
ggml-org:masterfrom
praneshgo:pgonegandla/dio-staged-read
Oct 1, 2026
Merged

ORippler merged 1 commit into
ggml-org:masterfrom
praneshgo:pgonegandla/dio-staged-read

Conversation

@praneshgo

@praneshgo praneshgo commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Overview

On Linux with -lm dio, read_aligned_chunk allocated an aligned temporary buffer as large as the whole tensor. Each host tensor was briefly held twice in memory. This PR fixes it by making the read go through a buffer of at most 64 MB, reused for each chunk of the tensor.

DGX Spark (GB10), CUDA, -lm dio:

model / setup peak RSS (GB) min free RAM (GB) load time (s) pp512 (t/s) tg128 (t/s) perplexity
Qwen3.8 UD-IQ4_XS, -lzm off 54.61 → 28.13 4.0 → 29.8 34.0 → 24.3 956.5 → 957.8 30.38 → 30.40 3.6031 → 3.6031
Qwen3.8 UD-IQ4_XS, -lzm on 1.60 → 1.31 56.4 → 56.9 10.2 → 10.5 913.3 → 909.0 29.77 → 29.80 3.6031 → 3.6031
MoE Qwen3-30B-A3B Q4_K_M, -ngl 99 0.72 → 0.72 100.5 → 100.4 4.3 → 4.0 3041.4 → 3032.2 96.21 → 96.20 4.1900 → 4.1900
MoE Qwen3-30B-A3B Q4_K_M, -ngl 99 -ncmoe 99 17.07 → 17.07 100.8 → 100.7 19.3 → 16.1 1303.6 → 1297.9 22.54 → 22.23 4.1900 → 4.1900
dense Llama-3.1-8B Q4_K_M, -ngl 99 0.87 → 0.84 112.5 → 112.2 2.5 → 2.6 3220.8 → 3212.9 47.09 → 47.08 4.4046 → 4.4046
dense Llama-3.1-8B Q4_K_M, -ngl 0 5.29 → 4.96 113.7 → 113.8 6.8 → 6.7 1775.6 → 1783.4 13.32 → 13.45 4.4046 → 4.4046

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: AI was used to partially assist while making the code edits and to understand the contexts of the code base.

@ORippler ORippler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we have perf data (mode lloading) for non-PLE checkpoints with non-cpu-offloading? Both dense Llama-3.1-8B Q4_K_M, -ngl 0 and MoE Qwen3-30B-A3B Q4_K_M, -ngl 99 -ncmoe 99 are not pure CUDA runs

Comment thread src/llama-mmap.cpp
Comment on lines +344 to +354
for (size_t done = 0; done < bytes_to_read; ) {
const size_t n = std::min(buffer_size, bytes_to_read - done);
read_raw_unsafe(buffer.get(), n);

const size_t count = std::min(n - skip, size - copied);
memcpy(reinterpret_cast<char *>(dest) + copied, reinterpret_cast<char *>(buffer.get()) + skip, count);

copied += count;
skip = 0;
done += n;
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would it make sense to pipeline here? I.e. have double-buffering and memcpy stage 1 while stage 0 reads the next bytes. Or is memcpy much cheaper than reading

@praneshgo

Copy link
Copy Markdown
Contributor Author

Do we have perf data (mode lloading) for non-PLE checkpoints with non-cpu-offloading? Both dense Llama-3.1-8B Q4_K_M, -ngl 0 and MoE Qwen3-30B-A3B Q4_K_M, -ngl 99 -ncmoe 99 are not pure CUDA runs

Thanks for pointing that out, I have now added the non-CPU offload configs to the PR, which I earlier missed.

@ORippler ORippler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can target async/pipelined reads in a follow-up PR

@ORippler
ORippler marked this pull request as ready for review September 30, 2026 15:19
@ORippler
ORippler requested a review from ggerganov as a code owner September 30, 2026 15:19
@ggerganov ggerganov self-assigned this Sep 30, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 1, 2026
Stage direct I/O reads through a bounded 64MB buffer so a large
tensor is not held in memory twice while it loads.
Port of upstream PR ggml-org#29749.

Assisted-by: Qwen Code
@ORippler
ORippler merged commit 32dd62e into ggml-org:master Oct 1, 2026
17 of 18 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…-io (ggml-org#29749)

Assisted-by: Claude

Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…-io (ggml-org#29749)

Assisted-by: Claude

Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants