[c10d] Batch DDP reducer bucket-to-grad copy-out with _foreach_copy_ (#190524)
## Summary `DistributedDataParallel` has an opt-in `batched_grad_copy` flag. It reduces the number of small copy kernels DDP launches per bucket. Today it only batches one of the two copy directions. - Copy-in (already batched): each parameter's gradient is copied into the bucket's flat buffer. With the flag on, these are deferred and flushed as a single `_foreach_copy_` plus one flat `div_` per bucket. - Copy-out (this PR): after the allreduce, the reduced bucket is copied back into each parameter's `.grad`. This still ran one `grad.copy_(bucket_view)` per parameter in `finalize_bucket_dense`, which is the common path when `gradient_as_bucket_view=False` (the default with `optimizer.zero_grad(set_to_none=True)`). This PR extends the same flag to the copy-out. For parameters whose grad is already defined and globally used, the copy is deferred and the whole bucket is written back with a single `at::_foreach_copy_`. Two cases keep their old behavior, since they cannot be batched: an undefined grad is created with `clone_obey_contract` (a fresh allocation), and a globally unused grad is left untouched. To stay correct, the batched copy-out is only taken when the grads are the parameters' own `.grad` tensors: `gradient_as_bucket_view=False`, not `optim_in_backward`, and not under distributed autograd (which owns grads through an rpc context and writes them back per callback). Every other case falls back to the existing per-parameter `copy_bucket_to_grad`. For a bucket of N parameters this turns N copy dispatches into 1, removing N-1 kernel launches per bucket on CUDA. ## Benchmark 200-layer model (400 parameters), one bucket, `gradient_as_bucket_view=False`, counting the reducer's copy-out dispatches: | | copy-out dispatches | |---|---:| | `batched_grad_copy=False` | 400 per-parameter copies | | `batched_grad_copy=True` (this PR) | 1 batched `_foreach_copy_` | ## Test Plan ``` python test/distributed/test_c10d_gloo.py -k batched_grad_copy ``` The existing `batched_grad_copy` tests already compare grads with and without batching for `gradient_as_bucket_view=False`, `set_to_none`, unused parameters, `create_graph`, comm hooks, and bucket-view aliasing, so they now exercise the batched copy-out. A new test, `test_batched_grad_copy_batches_copy_out`, asserts the per-parameter `copy_bucket_to_grad` record is gone and a single `copy_bucket_to_grad_batched` flush runs when the flag is on. `lintrunner` is clean on the changed files. This change was authored with the help of an AI assistant. I have read the diff and can explain every line. Pull Request resolved: https://github.com/pytorch/pytorch/pull/190524 Approved by: https://github.com/d4l3k Co-authored-by: Tristan Rice <rice@fn.lc>
S
Sidharth Rajmohan committed
8fa740aad0437568cbe424b1f1759b09ecc242ca
Parent: bb63328
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com>
on 9/30/2026, 10:49:27 PM