Repository navigation
ci : add self-hosted webgpu to hf-jobs - #28712
Conversation
|
To be fair I've seen it timeout on our selfhosted t4 as well in the past https://github.com/ggml-org/llama.cpp/actions/runs/34319079872/job/102361526168, I don't know if it's just slow or if certain tests get stuck. |
|
That may just be shader caching. The nvidia driver is not fast with compiles, a run can be significantly slower if it loads all or many for the first time. |
|
Yeah, the lack of Vulkan shader cache makes it slow. The ephemeral HF jobs runners don't have a persistent storage like the self-hosted T4 runners, so every new run will start with a cold cache. Hopefully the |
Ah, is this cache stored somewhere trivial so we can put it in buckets? Either way |
|
It should be stored in I am preparing a PR that will extract the |
|
I think it should be fine to use -j for vulkan testing, and several months ago I made it so the compiles happen outside of any locks, so it should help with compile time. Not sure about other backends. There is an env var to override the shader cache location, but if there is no persistent storage I guess that won't help. |
The persistent storage issue may have been fixed, so it can probably be enabled if need be. |
|
Let's merge #28740, then rebase this PR on |
With -j 4 a test-backend-ops on my 470 finishes in around 5 minutes versus the 11 minutes or so for a regular run. No issues so far. |
9e33811 to
a165aba
Compare
|
Hm, looks like the |
|
Hm, and another Vulkan job also deadlocked, even without |
Yep, got killed after 2 hour limit: |
|
Yeah, I suspect some weird regression in Vulkan. Here is a 3rd job that got stuck earlier today: https://github.com/ggml-org/llama.cpp/actions/runs/34686526464/job/103534353964 |
|
I'm trying to debug the original device lost reported in #28740. I haven't seen any failures where it just deadlocks. It can crash after the device lost, not sure how that would show up in CI. |
a165aba to
56c11dd
Compare
eaf7192 to
5ad9504
Compare
5ad9504 to
3d882af
Compare
There was a problem hiding this comment.
IMO we can merge this. The runtimes are bit large (almost 1h), but the jobs don't utilize the ccache yet since we create is only on the master branch. So the runtime will be less after we merge.
Note that the HF runners are using the 580 NVIDIA driver. Not sure if there is a simple way to upgrade them to the new 615 driver that we recently installed on the old Azure runners.
Alternatively, we can split the work like this:
- Keep the existing Azure Tesla T4 runners (with the new 615 driver) running the
gpu-vulkan-nvidia-cmandgpu-vulkan-nvidia-cm2jobs - Move only the
gpu-webgpu-nvidiajob to the HF runners
This way we lift some of the pressure from the Azure runners (by moving the webgpu jobs to HF) + keep running the Vulkan CI jobs on the new driver.
Eventually, I would like to retire the Azure runners since the sponsorship will end sometime in 2027. But we can still keep them running for a little while until that happens.
|
Think there's much point in running them on |
|
https://github.com/ggml-org/llama.cpp/actions/runs/34751007440?pr=28712 |
|
Ok, so definitely slower, but most of that is build time, so should not be too bad once cache is in play. |
So, do this for now and comment out the new cm jobs for future use? |
Yes, let's do it like this. |
|
@ggerganov I think maybe Azure runners need a restart, seems to be stuck/slow, I just cancelled a bunch of queued jobs from yesterday. |
|
Yeah, there was a job in the morning stuck for 4h. I am not sure if this is somehow related to the issue that @jeffbolznv has identified: ggml-org/ggml#1626 (comment). To summarize, so far we observe:
|
|
Just found this one stuck for 5hrs+ |
Is the assert failure not being handled by the CI scripts? |
Usually is, no idea why it got stuck. |
Isn't there some kind of timeout we can set? There's no reason why that job should run longer than an hour. |
Apparently one can set |
* add self-hosted vulkan and webgpu to hf-jobs * try t4-medium * cont : adjust cpu backend threads * try t4-small again * restore cm jobs --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* add self-hosted vulkan and webgpu to hf-jobs * try t4-medium * cont : adjust cpu backend threads * try t4-small again * restore cm jobs --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* add self-hosted vulkan and webgpu to hf-jobs * try t4-medium * cont : adjust cpu backend threads * try t4-small again * restore cm jobs --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Overview
cont #28258
Additional information
Disabled CM hf-jobs for now.
Requirements