SIGN IN SIGN UP

[Performance Optimization] optimize cummax/cummin GPU scan kernels (#79729)

* align cummax/min kernel `KernelScanInnerWithIndices` with torch: >

1. change static share memory and block shape(16,32) to dynamic allocation
2. Parallel reduction with Sklansky replaces Brent-Kung(up-sweep+down-sweep)

* Add cub::DeviceScan::InclusiveScan for 1D tensor at cummax/min
O
omoYang committed
3ae83294a52c8dfb18987c576bbc2d1cc71e9cfe
Parent: f4e6f68
Committed by GitHub <noreply@github.com> on 9/10/2026, 3:25:58 AM