[Inductor] Don't flatten Blackwell WS template loops when Triton's Meta WS knob is on (#199061)
With `TRITON_USE_META_WS=1` but `config.triton.enable_template_autows` off, max-autotune on the Blackwell WS persistent TMA template fails with `NoValidChoicesError`: every choice hits Triton's "Unexpected Producer Found" assert. Inductor only turns `FLATTEN` off for autoWS configs, but Triton's Meta WS knob is global, so it also warp-specializes the regular `FLATTEN=True` configs, and Meta WS doesn't support flattened loops. This change also turns `FLATTEN` off when the global knob is on. The heuristic is shared, so this also applies to the `addmm` and `scaled_mm` Blackwell TMA templates. Runs without the knob, including OSS Triton, are unchanged. Test Plan: New `test_global_meta_ws_disables_flatten` fails before the change and passes after. On B200 with `TRITON_USE_META_WS=1`, a max-autotune mm restricted to the WS template goes from `NoValidChoicesError` to all 10 choices compiling and matching eager exactly. ``` python test/inductor/test_max_autotune_blackwell.py ``` This PR was authored with Claude Code. Differential Revision: [D122477722](https://our.internmc.facebook.com/intern/diff/D122477722) Pull Request resolved: https://github.com/pytorch/pytorch/pull/199061 Approved by: https://github.com/njriasan
J
Janani Sriram committed
1d40e77c093f371f28fa9caa18ba2ff3bd1c8eff
Parent: 582f09a
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com>
on 10/1/2026, 6:14:43 AM