Repository navigation
Avoid a second full-size copy of each tensor with direct-io - #29749
Conversation
Assisted-by: Claude
ORippler
left a comment
There was a problem hiding this comment.
Do we have perf data (mode lloading) for non-PLE checkpoints with non-cpu-offloading? Both dense Llama-3.1-8B Q4_K_M, -ngl 0 and MoE Qwen3-30B-A3B Q4_K_M, -ngl 99 -ncmoe 99 are not pure CUDA runs
| for (size_t done = 0; done < bytes_to_read; ) { | ||
| const size_t n = std::min(buffer_size, bytes_to_read - done); | ||
| read_raw_unsafe(buffer.get(), n); | ||
|
|
||
| const size_t count = std::min(n - skip, size - copied); | ||
| memcpy(reinterpret_cast<char *>(dest) + copied, reinterpret_cast<char *>(buffer.get()) + skip, count); | ||
|
|
||
| copied += count; | ||
| skip = 0; | ||
| done += n; | ||
| } |
There was a problem hiding this comment.
Would it make sense to pipeline here? I.e. have double-buffering and memcpy stage 1 while stage 0 reads the next bytes. Or is memcpy much cheaper than reading
Thanks for pointing that out, I have now added the non-CPU offload configs to the PR, which I earlier missed. |
ORippler
left a comment
There was a problem hiding this comment.
Can target async/pipelined reads in a follow-up PR
Stage direct I/O reads through a bounded 64MB buffer so a large tensor is not held in memory twice while it loads. Port of upstream PR ggml-org#29749. Assisted-by: Qwen Code
…-io (ggml-org#29749) Assisted-by: Claude Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
…-io (ggml-org#29749) Assisted-by: Claude Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
Overview
On Linux with -lm dio, read_aligned_chunk allocated an aligned temporary buffer as large as the whole tensor. Each host tensor was briefly held twice in memory. This PR fixes it by making the read go through a buffer of at most 64 MB, reused for each chunk of the tensor.
DGX Spark (GB10), CUDA, -lm dio:
Requirements