Repository navigation
Conversation
|
Hi @4ndrearossetti, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Hello,
So, it seems it's three infrastructure-related faillures, rather than anything from my code. Thank you very much! |
JohannesGaessler
left a comment
There was a problem hiding this comment.
I just realized: please add a corresponding test case to test-backend-ops.cpp as well.
c18aa07 to
0458c7c
Compare
|
Hello, |
|
A permutation is not enough for the tests. Please add the option to calculate a sum over a non-contiguous view. |
|
Hello Johannes, That's right. I added the slice option following the That should be all. Let me know if I missed anything. Thank you! |
Overview
test-backend-opsincludesSUM(type=f32,ne=[33,256,1,1],permute=[1,0,2,3]), which CUDA rejects and which crashes the CPU reference if enabled. Two small changes:supports_oprequiredggml_is_contiguous_rows, but the kernel only requiresggml_is_contiguously_allocated(asserted inggml_cuda_op_sumsince llama/ggml: add LLM training support #10544). The flat CUB reduction is valid for any gap-free view, since summation is order-independent. Thecontiguous_rowscheck was added in metal: optimiseGGML_OP_SUM#16559 to keep CI green after that PR introduced permuted SUM test cases, it wasn't based on what the CUDA kernel can actually handle. This change makessupports_opmatch the condition the kernel itself asserts. The kernel is unchanged.ggml_compute_forward_sum_f32assertednb[0] == sizeof(float), so it aborted on non-contiguous rows, reached as soon as the CPU acts as reference (or fallback) for this case. Added a strided path keepingggml_floataccumulation; contiguous fast path unchanged. f16/bf16 variants left as-is.Why CPU and CUDA together : the CPU fix is required for the CUDA case to be testable, since the reference crashes once the case is enabled.
Additional information
Testing (RTX 3070 Ti Laptop, CUDA 13.3):
test-backend-ops -o SUM: 8/8, permuted case now passingtest-backend-ops: 16094/16094perf -o SUM: unchangedRef #14909, #16559.
Requirements