Repository navigation
Conversation
|
Great, I'm running my repros on it + quick windows build/test. |
ServeurpersoCom
left a comment
There was a problem hiding this comment.
Built and tested on Windows, no regression, and on Linux I confirmed this fixes real corruption on top of the CI annoyance: on master two servers sharing one LLAMA_CACHE produce a blob exactly twice the right size, while with this PR the sha256 matches its filename again and the full server suite passes with 4 xdist workers.
Nits, none blocking: remove(path_progress) comes after the early return -1 so a failed download leaves the lock file behind, and the std::rename just below still fails on Windows when the destination exists, same for the etag rewrite in write_file. I have the cross-platform fix ready, std::filesystem::rename on both sites, it went away with the revert in #28555, happy to push it here or as its own PR, whichever you prefer.
|
For cached downloads, I think the cache "module" should define the locking scheme, we should match the |
|
That feels like a lot of complexity (and potentially error-prone) just to share download progress. Couldn't we simply wait to acquire the appropriate lock (generic or provided by the cache backend), then check whether the download completed? |
|
Agreed the lock belongs to the cache module with the HF .locks layout. Progress is not optional for us though, a server download can run for tens of minutes and the WebUI has to show something. |
|
Hmm ok right, this case is rare in general so probably progress reporting is an unnecessary complexity part. will remove that logic and simply replace with a log message "file is downloaded by another process, waiting..." |
|
I retested by starting two llama-server at the same time with the same LLAMA_CACHE and the same -hf model. On master one of them dies with unable to rename and the blob that survives is exactly twice the right size, which is easy to check since HF names blobs by their hash: the sha256sum of the file should equal its filename, and on master it does not. With your PR it matches again and one process waits for the other. Two different models never contend, they lock different paths. Two things left: Across six runs I still lose a server twice, on get_repo_commit: error: failed to write file: .../refs/main. That is safe_write_file in hf-cache.cpp going through a shared path + ".tmp", the same bug one level down, and a per-process temp name is enough since there is nothing to resume and both write the same content. And the std::rename just below still fails on Windows when the destination exists, same for the etag rewrite in write_file, where std::filesystem::rename has the POSIX semantics everywhere. |
|
@angt can you have a look? |
|
per discussion via DM, let's push your changes directly here @angt . I'll temporary move this PR to draft |
Overview
Ref discussion: #28555 (comment)
Requirements