hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)
* hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA * hex-cpy: various fixes on top of the concat optimizations Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA. While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA. Added missing dma_queue_flush() calls. Added additional guards for conditions we don't support. --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
Y
Yiwei Shao committed
dcd387a412ca54e172a8d60eb71ef6753850c8ca
Parent: d775ebf
Committed by GitHub <noreply@github.com>
on 10/1/2026, 3:38:37 PM