Skip to content

Windows unbuffered model load - #26014

Open
JTischbein wants to merge 3 commits into
ggml-org:masterfrom
JTischbein:windows_unbuffered_model_load
Open

JTischbein wants to merge 3 commits into
ggml-org:masterfrom
JTischbein:windows_unbuffered_model_load

Conversation

@JTischbein

@JTischbein JTischbein commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Testing

As the DirectIO change on Linux lead to many new issues with several backends and devices, I'd appreciate it if people could test this PR on different devices and let me know if it works.

Overview

Linux supported unbuffered reading from disk via DirectIO for fast model loading. On my machine with PCIe5.0 NVMe drive this path increases the model loading performance from 1-3GB/s to >9GB/s.

For Windows this change is needed, as using mmap of large models on setups with small CPU RAM lead to Sysmem pressure and eviction of runtime pages.

Additional information

This PR introduces the unbuffered read file API of Windows. #18012 is the merged Linux PR.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: I have used AI for a first implementation. I reviewed, edited and tested the code on my own.

@ggml-gh-bot

ggml-gh-bot Bot commented Jul 22, 2026

Copy link
Copy Markdown

Hi @JTischbein, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 2 open PRs.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@JTischbein
JTischbein marked this pull request as ready for review July 23, 2026 13:46
@JTischbein
JTischbein requested a review from ggerganov as a code owner July 23, 2026 13:46
Comment thread src/llama-mmap.cpp
LocalFree(lpMsgBuf);
}

while (!ret.empty() && (ret.back() == '\r' || ret.back() == '\n')) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you elaborate why this is needed?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The returned message from FormatMessageA often contains a line break at the end. With this change we avoid double line breaks when logging the error message

Comment thread src/llama-mmap.cpp

impl(const char * fname, const char * mode, [[maybe_unused]] const bool use_direct_io = false) {
fp = ggml_fopen(fname, mode);
impl(const char * fname, const char * mode, const bool use_direct_io = false) : fname(fname) {

@stevenhoving stevenhoving Jul 23, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My 2 cents on this mr, is that it actually needs to be decomposed into separate classes. With an interface and multiple implementation classes. This allow you to keep things simple and seperate (mmap, direct_io, win_unbuffered, etc).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you on that, for now I wanted to keep the changes minimal to validate whether unbuffered IO is welcome on WIndows and works on all systems.

Let me try to get a new version together

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would prefer to make a separate PR for the abstraction of the paths. The required changes are quite large

@taronaeo taronaeo added help wanted Needs help from the community windows Issues specific to Windows labels Jul 27, 2026

@ORippler ORippler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for bringing this feature to Windows as well!

  • Regarding the refactor: nice to have, but beyond the scope of this PR imo - this PR aims to bring Windows to performance parity with Linux for the direct-io path
  • What would be nice is some tests for file loading (both for Linux and Windows). Not sure how difficult they are to spin up and if they should be part of this PR.
  • Do you have some perf numbers for dGPU systems you could share?

Comment thread src/llama-mmap.cpp
Comment on lines +315 to +330
void * raw_buffer = _aligned_malloc(bytes_to_read, alignment);
if (raw_buffer == nullptr) {
LLAMA_LOG_WARN("%s: Falling back to buffered I/O due to %s\n",
__func__, GetErrorMessageWin32(ERROR_NOT_ENOUGH_MEMORY).c_str());
CloseHandle(fp_win32_direct);
fp_win32_direct = INVALID_HANDLE_VALUE;
alignment = 1;
seek(offset, SEEK_SET);
read_raw_unsafe(dest, len);
return;
}

struct aligned_buffer_deleter {
void operator()(void * p) const { _aligned_free(p); }
};
std::unique_ptr<void, aligned_buffer_deleter> buffer(raw_buffer);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no need to materialize raw_buffer, just pass it into the unique_ptr directly

@JTischbein JTischbein Aug 1, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changed with the upcoming commit, thanks!

Comment thread src/llama-mmap.cpp
}
const size_t bytes_to_read = (offset_from_alignment + len + alignment - 1) & ~(alignment - 1);

void * raw_buffer = _aligned_malloc(bytes_to_read, alignment);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I guess adding a chunked read is not advisable from a perf perspective? LM Head may be a couple of 100 MB in size (500 MB for Qwen3.6-35BA3B if we have 8 BPW)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually we sometimes have regressing performance with larger read size, in my tests above 512MB (1GB compared to 512MB loads 16% slower). I will add chunked reads in the upcoming PR

@stevenhoving stevenhoving Aug 7, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can remember a 'trick' to deal with this more efficiently.

  • first check if the buffer given is already aligned.
  • If not, we read unaligned until we are at the aligned boundary.
  • Handle the left over bytes as an aligned read.

That would avoid the memory allocation

@nibor1896

nibor1896 commented Aug 4, 2026 •

Copy link
Copy Markdown

Just FYI, no critic at all or anything like that:

PR#26542 covers the "OVERLAPPED" and "Handle/Thread" part of this problem - both PRs combined, should solve this problem.

Shared handle at depth 8: 1.01x
One handle per thread: 2.22x

Yours
Robin

@ggerganov

Copy link
Copy Markdown
Member

What would be nice is some tests for file loading (both for Linux and Windows). Not sure how difficult they are to spin up and if they should be part of this PR.

Yes, it shouldn't be difficult to unit-test the llama_file and llama_mmap classes:

  • Generate some large dummy file with known bytes
  • Read it in different models
  • Validate read contents are expected
  • Track performance

Then we can run these tests on multiple devices on master and compare the results after the changes. This way there is no second-guessing.

@JTischbein
JTischbein force-pushed the windows_unbuffered_model_load branch from 2cd2f5f to a2923f6 Compare September 3, 2026 13:48
@JTischbein
JTischbein requested a review from CISC as a code owner September 3, 2026 13:48
@github-actions github-actions Bot added the testing Everything test related label Sep 3, 2026
@ORippler

ORippler commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

We should add buffered fall-back path similar to #29749 for Windows also

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

help wanted Needs help from the community testing Everything test related windows Issues specific to Windows

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants