SIGN IN SIGN UP

cuda: guard the iq4_nl dequantize row kernel against short rows (#29683)

dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
R
R0CKSTAR committed
f872b591121761ac7b2af18283bd99bdc092a63a
Parent: a4d880f
Committed by GitHub <noreply@github.com> on 9/30/2026, 8:33:22 PM