(not quite the same as PR #26014, but the full path to resolution:
26014 does not contain "OVERLAPPED" and no "handle / thread")
Problem:
On Windows, llama.cpp cannot do "read from byte X" - only "read from
wherever you currently are". When several threads want to read from the
same file, they get in each other's way: they keep moving each other's
position.
The code implied it could read unbuffered: it takes the flag, then does
nothing with it, answers "yes" when asked, and writes "read unbuffered"
to the log.
What I built together with Claude:
Exactly what was already promised. llama.cpp can now jump to any point
(byte X) and read unbuffered, and every thread gets its own handle, so
they no longer interfere.
The other six pieces are not features, they are dependencies this needs.
Shared handle at depth 8: 1.01x
One handle per thread: 2.22x
Built with Claude Code (AI)
Text written by me (Robin - Human)
(not quite the same as PR #26014, but the full path to resolution:
26014 does not contain "OVERLAPPED" and no "handle / thread")
Problem:
On Windows, llama.cpp cannot do "read from byte X" - only "read from
wherever you currently are". When several threads want to read from the
same file, they get in each other's way: they keep moving each other's
position.
The code implied it could read unbuffered: it takes the flag, then does
nothing with it, answers "yes" when asked, and writes "read unbuffered"
to the log.
What I built together with Claude:
Exactly what was already promised. llama.cpp can now jump to any point
(byte X) and read unbuffered, and every thread gets its own handle, so
they no longer interfere.
The other six pieces are not features, they are dependencies this needs.
Shared handle at depth 8: 1.01x
One handle per thread: 2.22x
Built with Claude Code (AI)
Text written by me (Robin - Human)