SIGN IN SIGN UP

CUDA: bitonic argsort handles rows wider than one block (#28957)

Without CUB (HIP, MUSA) argsort ran the bitonic kernel with one thread
per padded column, so any row above 1024 entries launched an invalid
block configuration. Each thread now owns several columns, every stage
of the network runs all owned columns before the barrier, and the block
is capped at 1024 threads. Shared memory becomes the only bound, which
supports_op checks against the device instead of a fixed 1024.

Rows up to 1024 run the same work as before. Bit-exact with the CUB
path on rows of 2048.
P
Pascal committed
6a2743f028f78bfb88a7189607b49bde30df3769
Parent: 748d422
Committed by GitHub <noreply@github.com> on 9/29/2026, 6:09:10 PM