Repository navigation
vulkan : fix TOP_K for +inf/NaN inputs and k = 1 on negative values - #30107
Merged
Merged
Conversation
The bucket search in topk_nary_search.comp started from the range [0, 0xFF800000), which ends just below the ordered-uint mapping of +inf, so +inf and NaN were never counted. A workgroup block with fewer than k countable values left the ballot empty and the shader read uninitialized shared state (hang/device lost on NVIDIA, wrong indices on AMD), and a few +inf in a block were selected without being counted, dropping real top values. Map NaN to -inf on input, start from [0, 0xFFFFFFFF) so every value is counted, and clamp the top bucket's end (2^32) instead of wrapping to 0. The k = 1 path compared float bits as signed integers, which orders negative values backwards; compare floats instead. Add test_top_k_inf to test-backend-ops: negative values, fewer than k +inf and many -inf, for k = 1, 10, 40. Assisted-by: Claude Opus 5.5
|
Hi @gianni-cor, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
gianni-cor
marked this pull request as ready for review
October 7, 2026 16:02
jeffbolznv
approved these changes
Oct 7, 2026
0cc4m
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Fixes two bugs in the Vulkan
TOP_Kshader (topk_nary_search.comp).+inf / NaN inputs. The shader finds the k-th largest value of each workgroup's block by mapping the floats to ordered uints and counting them into buckets, narrowing the range each step. The initial range was
[0, 0xFF800000), which ends just below the mapping of +inf, so +inf and NaN were never counted. When a block has fewer than k countable values, no bucket reaches the limit, the ballot is empty, and the shader continues with uninitialized shared state (sh_min_idx,sh_total): on NVIDIA the dispatch hangs until the driver resets the device, on AMD it returns wrong indices. With only a few +inf in a block, they still pass the final>= range_minselection without having been counted, so real top values are dropped.k = 1 on negative values. The k = 1 fast path compared the float bit patterns as signed integers, which orders negative values backwards (-2.0 beats -1.0), so the result is wrong whenever a block's maximum is negative.
Fix.
[0, 0xFFFFFFFF), so every value, +inf included, falls in a bucket and the search always finds one.range_min + (SUBGROUP_SIZE << shift)= 2^32, is clamped instead of wrapping to 0 (previously only reachable with finite values above ~2^113; +inf now lands in that bucket).Additional information
How it shows up. In a downstream application built on llama.cpp b11018, on an RTX 5090 (NVIDIA Vulkan), loading an MTP model with
parallel: 1and thenparallel: 2in one process lost the device on the second load. The faulting submission was theTOP_Knode of the MTP draft context's backend sampler, run while that context was being created, most likely its warmup decode on logits computed from uninitialized (NaN) data. With this fix the crash no longer reproduces.Standalone repro (
ggml_top_kon the Vulkan backend, 248320 columns, k = 10):Tests. New
test_top_k_infcase intest-backend-ops: rows of distinct negative values with fewer than k +inf (none for k = 1, so the expected indicesare unique) and many -inf, for k = 1, 10, 40. On master 9881906, oe Radeon 8060S:
TOP_Kcases pass; the 6 new cases fail (k = 1 = 10/40 on the +inf);NaN is not in the test because the CPU reference (
std::partial_sortwith>) has no defined result for it. CUDA is not affected: it has its own implementation and passes the new cases.this ports tetherto#350 to master
Requirements