Skip to content

ci : run test-backend-ops as a dedicated ci/run.sh test - #28740

Merged
ggerganov merged 5 commits into
masterfrom
gg/ci-workflow-backend-ops
Sep 11, 2026
Merged

ggerganov merged 5 commits into
masterfrom
gg/ci-workflow-backend-ops

Conversation

@ggerganov

@ggerganov ggerganov commented Sep 11, 2026 •

Copy link
Copy Markdown
Member

Overview

Run test-backend-ops as a dedicated test in ci/run.sh outside ctest. When GG_BUILD_HIGH_PERF is set it keeps the existing CPU-only invocation (-b CPU); otherwise it runs all available backends without a backend filter. The test is parallelized with -j $(nproc).

Keep test-backend-ops as a built target but not registered with ctest to avoid duplicate execution, and remove the stale test-backend-ops exclusions from the OpenVINO, Vulkan and WebGPU workflows.

Enable GG_BUILD_HIGH_PERF and LLAMA_ARG_THREADS on the Graviton4 KleidiAI job and switch it to the standard self-hosted results/mnt paths.

Additional information

Adds TODO markers for decoupling tests from libllama.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

@github-actions github-actions Bot added testing Everything test related devops improvements to build systems and github actions labels Sep 11, 2026
@ggerganov ggerganov changed the title ci : run test-backend-ops as a dedicated gg test ci : run test-backend-ops as a dedicated ci/run.sh test Sep 11, 2026
Run test-backend-ops as a separate gg test in ci/run.sh so it is executed outside ctest. With GG_BUILD_HIGH_PERF it keeps the existing CPU-only invocation (-b CPU); otherwise it runs all available backends without a backend filter.

Remove the dedicated backend-ops workflow and keep test-backend-ops as a built target that is not registered with ctest to avoid duplicate runs.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
Move the test-backend-ops gg test before test-llama-archs.

Enable GG_BUILD_HIGH_PERF and LLAMA_ARG_THREADS on the Graviton4 KleidiAI job and use the standard self-hosted results/mnt paths.

Add TODO markers for decoupling tests from libllama.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
Pass -j $(nproc) to test-backend-ops in both high-perf and all-backend modes.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
@ggerganov
ggerganov force-pushed the gg/ci-workflow-backend-ops branch from 55a7cdc to 6c60722 Compare September 11, 2026 14:37
@ggerganov

Copy link
Copy Markdown
Member Author

@ggml-org/amd The test-backend-ops -j seems to crash when ROCm is enabled: https://github.com/ggml-org/llama.cpp/actions/runs/34576278519/job/103297889044?pr=28740#step:3:4865

@ggerganov
ggerganov marked this pull request as ready for review September 11, 2026 14:38
@ggerganov
ggerganov requested a review from a team as a code owner September 11, 2026 14:38
@ggerganov

Copy link
Copy Markdown
Member Author

The test-backend-ops -j N seems to still not work correctly with MoltenVK:

https://github.com/ggml-org/llama.cpp/actions/runs/34611260059/job/103302413736?pr=28740#step:3:5897

fyi @ggml-org/ggml-vulkan

@jeffbolznv

Copy link
Copy Markdown
Contributor

Likely a moltenvk bug. We could try to workaround this by adding locks in a few places, or should we just try to run that CI job without -j?

@ggerganov

Copy link
Copy Markdown
Member Author

For now I've disabled -j for this job: 6c5fd9b. We can revisit in the future.

@ggerganov
ggerganov merged commit b78a39a into master Sep 11, 2026
28 of 32 checks passed
@ggerganov
ggerganov deleted the gg/ci-workflow-backend-ops branch September 11, 2026 19:01
@harkgill-amd

Copy link
Copy Markdown
Contributor

@ggml-org/amd The test-backend-ops -j seems to crash when ROCm is enabled: https://github.com/ggml-org/llama.cpp/actions/runs/34576278519/job/103297889044?pr=28740#step:3:4865

The stream-capture discrepancy is something I'll look into further but #28782 is the minimal fix on the llama.cpp end to unblock test-backend-ops -j N for ROCm.

@ggerganov

Copy link
Copy Markdown
Member Author

@jeffbolznv The parallel test-backend-ops with Vulkan on the DGX Spark also had a failure https://github.com/ggml-org/llama.cpp/actions/runs/34673623216/job/103499459706#step:3:21413

@jeffbolznv

Copy link
Copy Markdown
Contributor

I think I have a good lead on this, but it's very intermittent and I need time to make sure my fixes really work.

@jeffbolznv

Copy link
Copy Markdown
Contributor

#28830 should fix the vulkan failure.

pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* ci : run test-backend-ops as a dedicated gg test

Run test-backend-ops as a separate gg test in ci/run.sh so it is executed outside ctest. With GG_BUILD_HIGH_PERF it keeps the existing CPU-only invocation (-b CPU); otherwise it runs all available backends without a backend filter.

Remove the dedicated backend-ops workflow and keep test-backend-ops as a built target that is not registered with ctest to avoid duplicate runs.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops earlier and enable high-perf on kleidiai

Move the test-backend-ops gg test before test-llama-archs.

Enable GG_BUILD_HIGH_PERF and LLAMA_ARG_THREADS on the Graviton4 KleidiAI job and use the standard self-hosted results/mnt paths.

Add TODO markers for decoupling tests from libllama.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops in parallel

Pass -j $(nproc) to test-backend-ops in both high-perf and all-backend modes.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : disable parallel tests for ROCm

* cont : disable parallel tests with MoltenVK
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
* ci : run test-backend-ops as a dedicated gg test

Run test-backend-ops as a separate gg test in ci/run.sh so it is executed outside ctest. With GG_BUILD_HIGH_PERF it keeps the existing CPU-only invocation (-b CPU); otherwise it runs all available backends without a backend filter.

Remove the dedicated backend-ops workflow and keep test-backend-ops as a built target that is not registered with ctest to avoid duplicate runs.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops earlier and enable high-perf on kleidiai

Move the test-backend-ops gg test before test-llama-archs.

Enable GG_BUILD_HIGH_PERF and LLAMA_ARG_THREADS on the Graviton4 KleidiAI job and use the standard self-hosted results/mnt paths.

Add TODO markers for decoupling tests from libllama.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops in parallel

Pass -j $(nproc) to test-backend-ops in both high-perf and all-backend modes.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : disable parallel tests for ROCm

* cont : disable parallel tests with MoltenVK
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* ci : run test-backend-ops as a dedicated gg test

Run test-backend-ops as a separate gg test in ci/run.sh so it is executed outside ctest. With GG_BUILD_HIGH_PERF it keeps the existing CPU-only invocation (-b CPU); otherwise it runs all available backends without a backend filter.

Remove the dedicated backend-ops workflow and keep test-backend-ops as a built target that is not registered with ctest to avoid duplicate runs.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops earlier and enable high-perf on kleidiai

Move the test-backend-ops gg test before test-llama-archs.

Enable GG_BUILD_HIGH_PERF and LLAMA_ARG_THREADS on the Graviton4 KleidiAI job and use the standard self-hosted results/mnt paths.

Add TODO markers for decoupling tests from libllama.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops in parallel

Pass -j $(nproc) to test-backend-ops in both high-perf and all-backend modes.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : disable parallel tests for ROCm

* cont : disable parallel tests with MoltenVK
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* ci : run test-backend-ops as a dedicated gg test

Run test-backend-ops as a separate gg test in ci/run.sh so it is executed outside ctest. With GG_BUILD_HIGH_PERF it keeps the existing CPU-only invocation (-b CPU); otherwise it runs all available backends without a backend filter.

Remove the dedicated backend-ops workflow and keep test-backend-ops as a built target that is not registered with ctest to avoid duplicate runs.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops earlier and enable high-perf on kleidiai

Move the test-backend-ops gg test before test-llama-archs.

Enable GG_BUILD_HIGH_PERF and LLAMA_ARG_THREADS on the Graviton4 KleidiAI job and use the standard self-hosted results/mnt paths.

Add TODO markers for decoupling tests from libllama.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : run test-backend-ops in parallel

Pass -j $(nproc) to test-backend-ops in both high-perf and all-backend modes.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* ci : disable parallel tests for ROCm

* cont : disable parallel tests with MoltenVK
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devops improvements to build systems and github actions testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants