SIGN IN SIGN UP

[c10d] Batch DDP reducer bucket-to-grad copy-out with _foreach_copy_ (#190524)

## Summary

`DistributedDataParallel` has an opt-in `batched_grad_copy` flag. It reduces the
number of small copy kernels DDP launches per bucket. Today it only batches one
of the two copy directions.

- Copy-in (already batched): each parameter's gradient is copied into the
  bucket's flat buffer. With the flag on, these are deferred and flushed as a
  single `_foreach_copy_` plus one flat `div_` per bucket.
- Copy-out (this PR): after the allreduce, the reduced bucket is copied back into
  each parameter's `.grad`. This still ran one `grad.copy_(bucket_view)` per
  parameter in `finalize_bucket_dense`, which is the common path when
  `gradient_as_bucket_view=False` (the default with
  `optimizer.zero_grad(set_to_none=True)`).

This PR extends the same flag to the copy-out. For parameters whose grad is
already defined and globally used, the copy is deferred and the whole bucket is
written back with a single `at::_foreach_copy_`. Two cases keep their old
behavior, since they cannot be batched: an undefined grad is created with
`clone_obey_contract` (a fresh allocation), and a globally unused grad is left
untouched.

To stay correct, the batched copy-out is only taken when the grads are the
parameters' own `.grad` tensors: `gradient_as_bucket_view=False`, not
`optim_in_backward`, and not under distributed autograd (which owns grads through
an rpc context and writes them back per callback). Every other case falls back
to the existing per-parameter `copy_bucket_to_grad`.

For a bucket of N parameters this turns N copy dispatches into 1, removing N-1
kernel launches per bucket on CUDA.

## Benchmark

200-layer model (400 parameters), one bucket, `gradient_as_bucket_view=False`,
counting the reducer's copy-out dispatches:

| | copy-out dispatches |
|---|---:|
| `batched_grad_copy=False` | 400 per-parameter copies |
| `batched_grad_copy=True` (this PR) | 1 batched `_foreach_copy_` |

## Test Plan

```
python test/distributed/test_c10d_gloo.py -k batched_grad_copy
```

The existing `batched_grad_copy` tests already compare grads with and without
batching for `gradient_as_bucket_view=False`, `set_to_none`, unused parameters,
`create_graph`, comm hooks, and bucket-view aliasing, so they now exercise the
batched copy-out. A new test, `test_batched_grad_copy_batches_copy_out`, asserts
the per-parameter `copy_bucket_to_grad` record is gone and a single
`copy_bucket_to_grad_batched` flush runs when the flag is on.

`lintrunner` is clean on the changed files.

This change was authored with the help of an AI assistant. I have read the diff
and can explain every line.

Pull Request resolved: https://github.com/pytorch/pytorch/pull/190524
Approved by: https://github.com/d4l3k

Co-authored-by: Tristan Rice <rice@fn.lc>
S
Sidharth Rajmohan committed
8fa740aad0437568cbe424b1f1759b09ecc242ca
Parent: bb63328
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com> on 9/30/2026, 10:49:27 PM