SIGN IN SIGN UP

[Inductor] Don't flatten Blackwell WS template loops when Triton's Meta WS knob is on (#199061)

With `TRITON_USE_META_WS=1` but `config.triton.enable_template_autows` off,
max-autotune on the Blackwell WS persistent TMA template fails with
`NoValidChoicesError`: every choice hits Triton's "Unexpected Producer Found"
assert. Inductor only turns `FLATTEN` off for autoWS configs, but Triton's Meta
WS knob is global, so it also warp-specializes the regular `FLATTEN=True`
configs, and Meta WS doesn't support flattened loops.

This change also turns `FLATTEN` off when the global knob is on. The heuristic
is shared, so this also applies to the `addmm` and `scaled_mm` Blackwell TMA
templates. Runs without the knob, including OSS Triton, are unchanged.

Test Plan:
New `test_global_meta_ws_disables_flatten` fails before the change and passes
after. On B200 with `TRITON_USE_META_WS=1`, a max-autotune mm restricted to the
WS template goes from `NoValidChoicesError` to all 10 choices compiling and
matching eager exactly.
```
python test/inductor/test_max_autotune_blackwell.py
```

This PR was authored with Claude Code.

Differential Revision: [D122477722](https://our.internmc.facebook.com/intern/diff/D122477722)
Pull Request resolved: https://github.com/pytorch/pytorch/pull/199061
Approved by: https://github.com/njriasan
J
Janani Sriram committed
1d40e77c093f371f28fa9caa18ba2ff3bd1c8eff
Parent: 582f09a
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com> on 10/1/2026, 6:14:43 AM