SIGN IN SIGN UP

[aotautograd] Cache graphs containing autocast/inference_mode and torch._refs.tensor nodes (#191409)

Several functions that Dynamo traces into FX graphs were missing from the
AOTAutogradCache allowlist in check_node_safe, so check_cacheable raised
BypassAOTAutogradCache on the first such node and disabled AOTAutograd caching
for the whole graph. Scuba shows these are the largest allowlist-gap bypasses in
prod. This adds to SAFE_TORCH_FUNCTIONS:

- _enter_inference_mode/_exit_inference_mode, traced for a torch.inference_mode
  region inside a compiled function. (_enter_autocast/_exit_autocast were
  already allowlisted by #192555.)
- torch._refs.tensor, which Dynamo emits as a raw call_function node for
  torch.tensor(data) when `data` contains a data-dependent/unspec scalar
  (torch/_dynamo/variables/torch.py) instead of decomposing it.

Their args (the inference_mode flag; the data list, whose scalars are graph
nodes or constants) are in the graph and hashed into the key.

This also fixes a soundness gap in the autocast caching from #192555.
torch.autocast(device) with no dtype traces to _enter_autocast(device, None,
...), which resolves its dtype at runtime from torch.get_autocast_dtype(device).
_record_runtime_state (from #181564) only recorded that dtype for devices where
autocast was already enabled, which it isn't when the autocast region is inside
the compiled function, so compiles under different ambient autocast dtypes
(torch.set_autocast_dtype) shared a key. The root cause is in the key, so the
fix is too: _record_runtime_state now records (enabled, dtype) for every
autocast device unconditionally. The alternative, bypassing the cache for
_enter_autocast with dtype=None, would drop caching for the common
`torch.autocast("cuda")` spelling. The extra key component costs essentially
nothing, since the default dtype only differs if the user calls
set_autocast_dtype; existing cache entries are invalidated once.

Fixes #191106.

Test Plan:

```
python -m pytest test/dynamo/test_aot_autograd_cache.py -k "autocast or inference_mode_in_graph or refs_tensor_in_graph"
```

Authored with Claude.

Pull Request resolved: https://github.com/pytorch/pytorch/pull/191409
Approved by: https://github.com/bobrenjc93
A
Aaron Orenstein committed
070e6e3ba911b52e12bfaaad9c30262c91dfb21d
Parent: 462c7e0
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com> on 10/1/2026, 2:02:40 AM