Skip to content

ci : add self-hosted webgpu to hf-jobs - #28712

Merged
CISC merged 5 commits into
masterfrom
cisc/ci-even-more-hf-jobs
Sep 16, 2026
Merged

CISC merged 5 commits into
masterfrom
cisc/ci-even-more-hf-jobs

Conversation

@CISC

@CISC CISC commented Sep 10, 2026 •

Copy link
Copy Markdown
Member

Overview

cont #28258

Additional information

Turns out it was simpler than initially thought, Docker Vulkan passthrough works OOTB, thanks to @olliewalsh for pointing this out.

Disabled CM hf-jobs for now.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Təwə
  • Language disclosure: Malay (Rejang script)

@CISC
CISC requested a review from a team as a code owner September 10, 2026 18:24
@github-actions github-actions Bot added the devops improvements to build systems and github actions label Sep 10, 2026
@netrunnereve

Copy link
Copy Markdown
Contributor

To be fair I've seen it timeout on our selfhosted t4 as well in the past https://github.com/ggml-org/llama.cpp/actions/runs/34319079872/job/102361526168, I don't know if it's just slow or if certain tests get stuck.

@0cc4m

0cc4m commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

That may just be shader caching. The nvidia driver is not fast with compiles, a run can be significantly slower if it loads all or many for the first time.

@ggerganov

Copy link
Copy Markdown
Member

Yeah, the lack of Vulkan shader cache makes it slow. The ephemeral HF jobs runners don't have a persistent storage like the self-hosted T4 runners, so every new run will start with a cold cache.

Hopefully the test-backend-ops -j N option would result in a significant improvement in this use case - IIUC the shader compilation is CPU-bound.

@CISC

CISC commented Sep 11, 2026

Copy link
Copy Markdown
Member Author

Yeah, the lack of Vulkan shader cache makes it slow. The ephemeral HF jobs runners don't have a persistent storage like the self-hosted T4 runners, so every new run will start with a cold cache.

Ah, is this cache stored somewhere trivial so we can put it in buckets? Either way t4-medium should help a lot when the -j option lands.

@ggerganov

Copy link
Copy Markdown
Member

It should be stored in $HOME/.cache/nvidia/GLCache/ per Jeff's comment. I think depending on the outcome of enabling -j, we can consider caching or not the shaders.

I am preparing a PR that will extract the test-backend-ops runs into a separate CI workflow so we can control it more precisely.

@jeffbolznv

Copy link
Copy Markdown
Contributor

I think it should be fine to use -j for vulkan testing, and several months ago I made it so the compiles happen outside of any locks, so it should help with compile time. Not sure about other backends.

There is an env var to override the shader cache location, but if there is no persistent storage I guess that won't help.

@CISC

CISC commented Sep 11, 2026

Copy link
Copy Markdown
Member Author

There is an env var to override the shader cache location, but if there is no persistent storage I guess that won't help.

The persistent storage issue may have been fixed, so it can probably be enabled if need be.

@ggerganov

Copy link
Copy Markdown
Member

Let's merge #28740, then rebase this PR on master and see if the -j helped.

@netrunnereve

Copy link
Copy Markdown
Contributor

I think it should be fine to use -j for vulkan testing, and several months ago I made it so the compiles happen outside of any locks, so it should help with compile time. Not sure about other backends.

With -j 4 a test-backend-ops on my 470 finishes in around 5 minutes versus the 11 minutes or so for a regular run. No issues so far.

@ggerganov
ggerganov force-pushed the cisc/ci-even-more-hf-jobs branch from 9e33811 to a165aba Compare September 12, 2026 13:45
@ggerganov

Copy link
Copy Markdown
Member

Hm, looks like the gpu-vulkan-nvidia-cm job deadlocked now with -j 7: https://github.com/ggml-org/llama.cpp/actions/runs/34697303026/job/103562882003?pr=28712#step:6:24943

@ggerganov

Copy link
Copy Markdown
Member

Hm, and another Vulkan job also deadlocked, even without -j: https://github.com/ggml-org/ggml/actions/runs/34689579337/job/103542363483?pr=1626

@CISC

CISC commented Sep 12, 2026

Copy link
Copy Markdown
Member Author

Hm, looks like the gpu-vulkan-nvidia-cm job deadlocked now with -j 7: https://github.com/ggml-org/llama.cpp/actions/runs/34697303026/job/103562882003?pr=28712#step:6:24943

Yep, got killed after 2 hour limit:

2026-09-12 13:47:56Z: Running job: gpu-vulkan-nvidia-cm
2026-09-12 15:58:46Z: Job gpu-vulkan-nvidia-cm completed with result: Canceled

@ggerganov

Copy link
Copy Markdown
Member

Yeah, I suspect some weird regression in Vulkan. Here is a 3rd job that got stuck earlier today: https://github.com/ggml-org/llama.cpp/actions/runs/34686526464/job/103534353964

@jeffbolznv

Copy link
Copy Markdown
Contributor

I'm trying to debug the original device lost reported in #28740. I haven't seen any failures where it just deadlocks. It can crash after the device lost, not sure how that would show up in CI.

@ggerganov
ggerganov force-pushed the cisc/ci-even-more-hf-jobs branch from a165aba to 56c11dd Compare September 13, 2026 06:24
@ggerganov
ggerganov self-requested a review as a code owner September 13, 2026 07:20
@github-actions github-actions Bot added the testing Everything test related label Sep 13, 2026
@ggerganov
ggerganov force-pushed the cisc/ci-even-more-hf-jobs branch from eaf7192 to 5ad9504 Compare September 13, 2026 08:56
@ggerganov
ggerganov force-pushed the cisc/ci-even-more-hf-jobs branch from 5ad9504 to 3d882af Compare September 13, 2026 10:06

@ggerganov ggerganov left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

IMO we can merge this. The runtimes are bit large (almost 1h), but the jobs don't utilize the ccache yet since we create is only on the master branch. So the runtime will be less after we merge.

Note that the HF runners are using the 580 NVIDIA driver. Not sure if there is a simple way to upgrade them to the new 615 driver that we recently installed on the old Azure runners.

Alternatively, we can split the work like this:

  • Keep the existing Azure Tesla T4 runners (with the new 615 driver) running the gpu-vulkan-nvidia-cm and gpu-vulkan-nvidia-cm2 jobs
  • Move only the gpu-webgpu-nvidia job to the HF runners

This way we lift some of the pressure from the Azure runners (by moving the webgpu jobs to HF) + keep running the Vulkan CI jobs on the new driver.

Eventually, I would like to retire the Azure runners since the sponsorship will end sometime in 2027. But we can still keep them running for a little while until that happens.

@CISC

CISC commented Sep 13, 2026

Copy link
Copy Markdown
Member Author

Think there's much point in running them on t4-medium, or should we go back to t4-small?

@ggerganov

Copy link
Copy Markdown
Member

t4-small is probably OK. We can compare it against the last t4-medium run:

https://github.com/ggml-org/llama.cpp/actions/runs/34751007440?pr=28712

@CISC

CISC commented Sep 15, 2026

Copy link
Copy Markdown
Member Author

Ok, so definitely slower, but most of that is build time, so should not be too bad once cache is in play.

@CISC

CISC commented Sep 15, 2026

Copy link
Copy Markdown
Member Author
  • Keep the existing Azure Tesla T4 runners (with the new 615 driver) running the gpu-vulkan-nvidia-cm and gpu-vulkan-nvidia-cm2 jobs
  • Move only the gpu-webgpu-nvidia job to the HF runners

This way we lift some of the pressure from the Azure runners (by moving the webgpu jobs to HF) + keep running the Vulkan CI jobs on the new driver.

Eventually, I would like to retire the Azure runners since the sponsorship will end sometime in 2027. But we can still keep them running for a little while until that happens.

So, do this for now and comment out the new cm jobs for future use?

@ggerganov

Copy link
Copy Markdown
Member

So, do this for now and comment out the new cm jobs for future use?

Yes, let's do it like this.

@CISC CISC changed the title ci : add self-hosted vulkan and webgpu to hf-jobs ci : add self-hosted webgpu to hf-jobs Sep 15, 2026
@CISC

CISC commented Sep 15, 2026

Copy link
Copy Markdown
Member Author

@ggerganov I think maybe Azure runners need a restart, seems to be stuck/slow, I just cancelled a bunch of queued jobs from yesterday.

@ggerganov

Copy link
Copy Markdown
Member

Yeah, there was a job in the morning stuck for 4h. I am not sure if this is somehow related to the issue that @jeffbolznv has identified: ggml-org/ggml#1626 (comment).

To summarize, so far we observe:

  • Intermittent ARGSORT failures
  • Some jobs getting stuck (I think mainly CM2, rather than CM, though not 100% sure)
  • Random crash in one of the whisper.cpp vulkan jobs

@CISC

CISC commented Sep 15, 2026

Copy link
Copy Markdown
Member Author

@jeffbolznv

Copy link
Copy Markdown
Contributor

Just found this one stuck for 5hrs+ https://github.com/ggml-org/llama.cpp/actions/runs/34939723411/job/104285387342

Is the assert failure not being handled by the CI scripts?

@CISC

CISC commented Sep 15, 2026

Copy link
Copy Markdown
Member Author

Just found this one stuck for 5hrs+ https://github.com/ggml-org/llama.cpp/actions/runs/34939723411/job/104285387342

Is the assert failure not being handled by the CI scripts?

Usually is, no idea why it got stuck.

@netrunnereve

Copy link
Copy Markdown
Contributor

Usually is, no idea why it got stuck.

Isn't there some kind of timeout we can set? There's no reason why that job should run longer than an hour.

@CISC

CISC commented Sep 15, 2026

Copy link
Copy Markdown
Member Author

Usually is, no idea why it got stuck.

Isn't there some kind of timeout we can set? There's no reason why that job should run longer than an hour.

Apparently one can set timeout-minutes on a job, it defaults to 6hrs though so I guess it would have timed out on its own shortly afterwards if I hadn't cancelled it.

@CISC
CISC merged commit 583926e into master Sep 16, 2026
26 checks passed
@CISC
CISC deleted the cisc/ci-even-more-hf-jobs branch September 16, 2026 06:24
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
* add self-hosted vulkan and webgpu to hf-jobs

* try t4-medium

* cont : adjust cpu backend threads

* try t4-small again

* restore cm jobs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* add self-hosted vulkan and webgpu to hf-jobs

* try t4-medium

* cont : adjust cpu backend threads

* try t4-small again

* restore cm jobs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* add self-hosted vulkan and webgpu to hf-jobs

* try t4-medium

* cont : adjust cpu backend threads

* try t4-small again

* restore cm jobs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devops improvements to build systems and github actions testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants