SIGN IN SIGN UP

[ROCm] Enable HSA_ENABLE_IPC_MODE_LEGACY to work around driver handle freeing issue (#198921)

The amdgpu KFD driver has a bug in DMA-BUF IPC import. When a process opens another process's shared buffer on the same GPU, KFD reuses the owner's GEM handle instead of creating its own, and both processes delete that handle when they free. If another allocation reused the handle number in between, that buffer loses its handle. `hipIpcGetMemHandle` on it then fails with `invalid argument`, or returns a different process's buffer.

This shows up in CI whenever processes share GPU memory: RCCL (distributed tests) and PyTorch CUDA tensor sharing (e.g. `test_dataloader`). It is much more frequent on MI350 DPX runners, where two jobs share one physical GPU. We saw the RCCL error in 151 of 2,333 DPX distributed jobs vs 0 of 1,717 on whole-GPU runners, and the `test_dataloader` sharing error in 17 of 616 DPX runs vs 0 of 574. Most affected tests passed on rerun, which created flaky-test noise and retries, and in one job the test received wrong data.

In legacy mode, the importing process never gets a GEM handle, so the double delete cannot happen. This is a CI workaround until the driver is fixed.

Coauthored with Claude

Pull Request resolved: https://github.com/pytorch/pytorch/pull/198921
Approved by: https://github.com/pragupta, https://github.com/jeffdaily
Z
Z Liu committed
77dccac963ddc0a5d245e703178e5ffd033d4dc5
Parent: 22669c3
Committed by PyTorch MergeBot <pytorchmergebot@users.noreply.github.com> on 10/1/2026, 2:29:48 AM