CUDA: bitonic argsort handles rows wider than one block (#28957)
Without CUB (HIP, MUSA) argsort ran the bitonic kernel with one thread per padded column, so any row above 1024 entries launched an invalid block configuration. Each thread now owns several columns, every stage of the network runs all owned columns before the barrier, and the block is capped at 1024 threads. Shared memory becomes the only bound, which supports_op checks against the device instead of a fixed 1024. Rows up to 1024 run the same work as before. Bit-exact with the CUB path on rows of 2048.
P
Pascal committed
6a2743f028f78bfb88a7189607b49bde30df3769
Parent: 748d422
Committed by GitHub <noreply@github.com>
on 9/29/2026, 6:09:10 PM