fix(loss): use autograd-aware all_reduce in BarlowTwinsLoss (#1980)
* fix(loss): use autograd-aware all_reduce in BarlowTwinsLoss The raw torch.distributed.all_reduce overwrote the cross-correlation matrix in place and was not registered in the autograd graph. Under DDP gradient averaging this scaled the backbone gradient down by 1/world_size relative to single-GPU Barlow Twins on the global batch. Swap it for the autograd-aware torch.distributed.nn.all_reduce, whose backward is itself an all_reduce, so the gradient is scaled back up by world_size. The forward value is unchanged. Fixes #1977. Same fix as SIGReg in #1923. * test(loss): remove mock-based gather_distributed tests Mocking all_reduce only tested the mock, not the real collective, and would break on any change to the reduction. Real multi-rank coverage is tracked in #1982. --------- Co-authored-by: Saud Kamran <saud.kamran92@gmail.com>
S
Saud Kamran committed
d8a260d502a7c12d9f4bdd88c6eabf9d58decaf7
Parent: 6b5e876
Committed by GitHub <noreply@github.com>
on 7/20/2026, 8:07:45 AM