[Operator Mechanism] Support FP32 output for FP16/BF16 paddle.bmm on CUDA (#79552)
* [Python API] Move bmm to handwritten implementation * [Operator Mechanism] Support bmm out_dtype for BF16 CUDA * [Operator Mechanism] Support bmm out_dtype for FP16 CUDA * [CodeStyle] Fix bmm out_dtype pre-commit issues * [API Compatibility] Preserve paddle.tensor.linalg.bmm * [Test] Improve bmm Python API coverage * [CodeStyle] Centralize bmm validation in InferMeta * [Python API] Restore bmm legacy static support * [Python API] Validate bmm out_dtype device support * [Python API] Reject bmm out_dtype on ROCm * [API Compatibility] Align bmm positional signature * [API Compatibility] Preserve bmm positional name
F
feixi committed
f9e513c04ff545db64630389289119766e26ffed
Parent: 1a067f0
Committed by GitHub <noreply@github.com>
on 8/10/2026, 6:14:20 AM