[Inductor] Preserve input dtype in the CPU bmm mul+sum decomposition (#191324)
## Summary Fixes #191308. On CPU, `torch.compile(backend="inductor")` changed the output dtype of same-dtype integer `bmm`/`matmul` from the input dtype (e.g. `int8`) to `int64`, while eager, `eager` backend, and `aot_eager` all preserve the input dtype. The CPU-specific decomposition in `tuned_bmm` (taken when `mat1` has M==1 or `mat2` has N==1) lowers bmm to a broadcasted mul followed by a sum. For integer inputs the reduction accumulates at the compute type (`int64`), and the result was returned as-is, leaking the accumulator dtype into the output. The reduction result is now cast back to the input dtype. For float inputs the cast is a no-op, and the int64 accumulator already matches or exceeds eager's accumulation precision, so values are unchanged. The dot-shaped CUDA/XPU decomposition above it is restricted to float dtypes and unaffected. ## Test plan - New regression test `test_bmm_cpu_decompose_preserves_integer_output_dtype` in `test/inductor/test_torchinductor.py`: bmm with M==1 across int8/int32/int64/float32 asserts compiled dtype equals eager dtype and values match. - Verified on this machine (arm64 macOS, torch 2.13.0) by patching the installed wheel's `torch/_inductor/kernel/bmm.py` in place and running the issue repro plus the four-dtype matrix: before, `compiled.dtype == torch.int64` vs eager `int8`; after, all four dtypes match eager and `(compiled == eager).all()` holds. The in-tree test file imports symbols newer than the 2.13.0 wheel, so it was validated by its standalone equivalent rather than run from the test module itself; a full source build was not run on macOS. Pull Request resolved: https://github.com/pytorch/pytorch/pull/191324 Approved by: https://github.com/jansel
Y
Yufeng He committed
198f1763922971ef1eeb464a8e243ac360c6c63e
Parent: 67148e2
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com>
on 10/1/2026, 3:55:01 AM