Repository navigation
Tests: extend test-quantize-fns to test nrc=2 (i8mm) kernels - #16234
Conversation
|
Hi @ggerganov , could you trigger the CI again? |
|
@ggerganov thanks for running the CI! Let me know if I can make any more changes |
|
|
||
| float result = INFINITY; | ||
| qfns_cpu->vec_dot(test_size, &result, 0, tmp_q1.data(), 0, tmp_q2.data(), 0, 1); | ||
| qfns_cpu->vec_dot(test_size, &result, 0, tmp_q1.data(), 0, tmp_q2.data(), 0, nrc); |
There was a problem hiding this comment.
I think we need to expand the input data depending on nrc?
There was a problem hiding this comment.
Sure. Can you tell me an appropriate test_size for nrc=2 ? Would it just be double?
|
I went through the logs of some of the failed tests and it looks like they may be unrelated to the modified CI in this PR. @ggerganov or the team, can you help me with this/ tell me if I'm missing something? Thank you! |
|
@Rohanjames1997 , how about you rebase the PR to the main? that will trigger CI and you can check if the current failures have been fixed on the mainline already. |
|
Some CI failed again due to the Cloudflare/Github outage yesterday. |
|
the |
|
what's the next step to get this PR merged? |
|
Fixed the Buffers are now sized with All 4 elements of the 2x2 output matrix are validated against float references. |
|
/bot review |
Automated code reviewReview of PR #16234 (commit f48b0ba) - extends Verified correct
Blocking (defeats the stated PR goal) (point 1) The PR is titled "Extend CI for i8mm kernels", but the diff only touches the test file - no workflow is added or changed to build and run
Note that simply forcing Will slow the review (point 2) Nits (point 3) In the (point 4) Summary: the test logic itself is correct and well designed (it mirrors the real This review was generated automatically by pi coding agent using |
|
Addressing the "Blocking" comment:
True. This PR future-proofs the CI. |
|
Since we are moving to HuggingFace runners, I was wondering if HuggingFace Jobs has any ARM CPU that supports $ hf jobs hardware
Hint: A new version of huggingface_hub (1.30.0) is available! You are using version 1.23.0.
To update, run: hf update
NAME PRETTY NAME CPU RAM STORAGE ACCELERATOR COST/MIN COST/HOUR
--------------- ---------------------- -------- ------- -------- ------------------------ -------- ---------
cpu-basic CPU Basic 2 vCPU 16 GB 50 GB $0.0002 $0.01
cpu-upgrade CPU Upgrade 8 vCPU 32 GB 50 GB $0.0005 $0.03
cpu-performance CPU Performance 32 vCPU 256 GB 1024 GB $0.0317 $1.90
cpu-xl CPU XL 16 vCPU 124 GB 1000 GB $0.0167 $1.00
t4-small Nvidia T4 - small 4 vCPU 15 GB 50 GB 1x T4 (16 GB) $0.0067 $0.40
t4-medium Nvidia T4 - medium 8 vCPU 30 GB 100 GB 1x T4 (16 GB) $0.0100 $0.60
a10g-small Nvidia A10G - small 4 vCPU 15 GB 110 GB 1x A10G (24 GB) $0.0167 $1.00
a10g-large Nvidia A10G - large 12 vCPU 46 GB 200 GB 1x A10G (24 GB) $0.0250 $1.50
a10g-largex2 2x Nvidia A10G - large 24 vCPU 92 GB 1000 GB 2x A10G (48 GB) $0.0500 $3.00
a10g-largex4 4x Nvidia A10G - large 48 vCPU 184 GB 2000 GB 4x A10G (96 GB) $0.0833 $5.00
a100-large Nvidia A100 - large 12 vCPU 142 GB 1000 GB 1x A100 (80 GB) $0.0417 $2.50
a100x4 4x Nvidia A100 48 vCPU 568 GB 4000 GB 4x A100 (320 GB) $0.1667 $10.00
a100x8 8x Nvidia A100 96 vCPU 1136 GB 8000 GB 8x A100 (640 GB) $0.3333 $20.00
h200 Nvidia H200 23 vCPU 256 GB 3000 GB 1x H200 (141 GB) $0.0833 $5.00
h200x2 Nvidia H200 46 vCPU 512 GB 6000 GB 2x H200 (282 GB) $0.1667 $10.00
h200x4 Nvidia H200 92 vCPU 1024 GB 12000 GB 4x H200 (564 GB) $0.3333 $20.00
h200x8 Nvidia H200 184 vCPU 2048 GB 24000 GB 8x H200 (1128 GB) $0.6667 $40.00
rtx-pro-6000 Nvidia RTX PRO 6000 23 vCPU 256 GB 475 GB 1x RTX PRO 6000 (96 GB) $0.0458 $2.75
rtx-pro-6000x2 Nvidia RTX PRO 6000 46 vCPU 512 GB 950 GB 2x RTX PRO 6000 (192 GB) $0.0917 $5.50
rtx-pro-6000x4 Nvidia RTX PRO 6000 92 vCPU 1024 GB 1900 GB 4x RTX PRO 6000 (384 GB) $0.1833 $11.00
rtx-pro-6000x8 Nvidia RTX PRO 6000 184 vCPU 2048 GB 3800 GB 8x RTX PRO 6000 (768 GB) $0.3667 $22.00
l4x1 1x Nvidia L4 8 vCPU 30 GB 400 GB 1x L4 (24 GB) $0.0133 $0.80
l4x4 4x Nvidia L4 48 vCPU 186 GB 3200 GB 4x L4 (96 GB) $0.0633 $3.80
l40sx1 1x Nvidia L40S 8 vCPU 62 GB 380 GB 1x L40S (48 GB) $0.0300 $1.80
l40sx4 4x Nvidia L40S 48 vCPU 382 GB 3200 GB 4x L40S (192 GB) $0.1383 $8.30
l40sx8 8x Nvidia L40S 192 vCPU 1534 GB 6500 GB 8x L40S (384 GB) $0.3917 $23.50
Hint: Use `hf jobs run --flavor <name> ...` to request a specific hardware flavor. |
No ARM runners, not sure if/when there will be any. |
|
I think we are good to merge? We will still need to find an ARM CI runner that has support for |
Any idea on what does? |
|
Did a quick google search, looks like SMMLA instruction support is what we need to look out for. For AWS, it seems like Graviton 3 through 5 supports this instruction. If I recall correctly, we have an existing KleidiAI self-hosted CI runner that uses AWS right? Let me check if it supports the instruction. |
|
Yes, we do have Graviton 4 available as an Arm-hosted runner for ggml-org/llama.cpp. |
Yay OK let's merge this and work making a CI run |
|
Merging in a few hours if no further comments. CI failures appear to not be related. |
|
@zhiyuan8 would be good to enable these in the Snapdragon CI |
…g#16234) * Test for nrc=2 as well | i8mm kernels * Trigger only on supported HW * Remove trailing whitespace * Address review comment * test: properly prepare nrc=2 inputs with independent data per row * tests : make nrc=2 dot product inputs distinct Assisted-by: Kiro * tests : use non-trivial strides in nrc=2 dot product test * tests : fail nrc=2 dot product test on non-finite errors
…g#16234) * Test for nrc=2 as well | i8mm kernels * Trigger only on supported HW * Remove trailing whitespace * Address review comment * test: properly prepare nrc=2 inputs with independent data per row * tests : make nrc=2 dot product inputs distinct Assisted-by: Kiro * tests : use non-trivial strides in nrc=2 dot product test * tests : fail nrc=2 dot product test on non-finite errors
…g#16234) * Test for nrc=2 as well | i8mm kernels * Trigger only on supported HW * Remove trailing whitespace * Address review comment * test: properly prepare nrc=2 inputs with independent data per row * tests : make nrc=2 dot product inputs distinct Assisted-by: Kiro * tests : use non-trivial strides in nrc=2 dot product test * tests : fail nrc=2 dot product test on non-finite errors
…g#16234) * Test for nrc=2 as well | i8mm kernels * Trigger only on supported HW * Remove trailing whitespace * Address review comment * test: properly prepare nrc=2 inputs with independent data per row * tests : make nrc=2 dot product inputs distinct Assisted-by: Kiro * tests : use non-trivial strides in nrc=2 dot product test * tests : fail nrc=2 dot product test on non-finite errors
The i8mm kernels require
nrc == 2to actually get triggered.But currently, the CI's test-quantize-fns.cpp tests only the scenario where
nrc=1.