Skip to content

ci: fix Models Backend by shortening the hrm_text fixture - #29744

Merged
CISC merged 1 commit into
ggml-org:masterfrom
ServeurpersoCom:ci-hrm-text-fixture
Sep 30, 2026
Merged

CISC merged 1 commit into
ggml-org:masterfrom
ServeurpersoCom:ci-hrm-text-fixture

Conversation

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Overview

Fix CI:

The new Models Backend check is red on every run because of hrm_text: it lands just above the 1e-4 bound on Vulkan (1.10e-04) and WebGPU (1.53e-04), same numbers every time.

The bound is fine, the hrm_text test model is just too deep for it. HRM reuses its layers in cycles, and our fixture runs them through 8 passes where most archs have 2. Each pass adds a bit of fp16 rounding, so it ends up about 10x noisier than any other arch. With f16 disabled it drops to 3.5e-10, so nothing is wrong in the model itself.

This PR makes the fixture a bit shorter: 6 passes instead of 8. It still goes through every branch of the cycle loop, and the error roughly halves.

NMSE over 5 seeds on an RTX PRO 6000:

  • CUDA: 1.0e-05 -> 4.5e-06
  • Vulkan: 6.0e-05 -> 2.6e-05

Additional information

Follow-up #29651 cc @CISC

Requirements

The fixture recycles its two blocks over 8 cache slots, so the fp16
error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU
jobs of Models Backend. Two l-cycles keep every branch of the cycle
loop and halve the error.
@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

Measured on an RTX PRO 6000, where the values differ from the CI GPUs, but the error scales with the depth so the CI should drop by the same ratio. The Models Backend run on this PR will confirm.

@CISC

CISC commented Sep 30, 2026

Copy link
Copy Markdown
Member

cc/ @jeffbolznv

@github-actions github-actions Bot added the testing Everything test related label Sep 30, 2026
@CISC
CISC merged commit 3b3d022 into ggml-org:master Sep 30, 2026
17 of 18 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…l-org#29744)

The fixture recycles its two blocks over 8 cache slots, so the fp16
error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU
jobs of Models Backend. Two l-cycles keep every branch of the cycle
loop and halve the error.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…l-org#29744)

The fixture recycles its two blocks over 8 cache slots, so the fp16
error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU
jobs of Models Backend. Two l-cycles keep every branch of the cycle
loop and halve the error.
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…l-org#29744)

The fixture recycles its two blocks over 8 cache slots, so the fp16
error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU
jobs of Models Backend. Two l-cycles keep every branch of the cycle
loop and halve the error.

(cherry picked from commit 3b3d022)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants