[Elastic][Distributed] Add configurable randomized backoff after failed rendezvous state writes (#185930)
At large scale, concurrent rendezvous updates can cause repeated CAS conflicts. The executor currently retries without backoff, potentially increasing backend load and delaying rendezvous progress during elastic restarts. This patch adds a small randomized backoff (0–0.3s) when the existing sync() call reports a failed write. It preserves the current synchronization point and state-machine structure. Pull Request resolved: https://github.com/pytorch/pytorch/pull/185930 Approved by: https://github.com/d4l3k
A
alexqdh committed
bb63328c063ab885f62f554fd3a35603432ee7eb
Parent: 0242c7b
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com>
on 9/30/2026, 10:42:17 PM