[Distributed] Add P2P_ISSUED pipeline-parallel micro-step hook location (#79690)
* [Pipeline Parallel] Add P2P_ISSUED micro-step hook location Adds a new `PipelineParallelMicroStepLocations.P2P_ISSUED` callback location, raised after a micro-step's forward P2P has been *issued* and before its wait handles are consumed. Motivation: user code that wants to hide an independent side computation (e.g. an auxiliary-loss branch whose gradient does not reach the backbone) behind the pipeline's forward send/recv has no place to put it today. `FORWARD_END` looks like the right spot but cannot work: `isend` / `irecv` record an event on the calculation stream at issue time and have the pipeline group's comm stream wait on it, so anything enqueued before the issue is inside that event's reach and the NCCL kernel waits for it instead of running concurrently. Hooks must therefore run *after* the issue and *before* the wait -- the same "issue, compute, then wait" shape the schedule already uses for `WeightGradStore.pop()`. Changes: * `PipelineParallelMicroStepLocations.P2P_ISSUED` plus its entry in the hook registry and the two assertion messages. * `P2pHelper.send_forward` gains `overlap_p2p_comm=False`, mirroring the option `send_forward_recv_forward` already has: it forwards `wait_on_reqs=not overlap_p2p_comm` to `_p2p_helper` and returns the wait handles so the caller can defer the wait. Default keeps the previous blocking behaviour and return value. * `VPPFhenBInBalancedMemory` defers the forward-send wait in its warmup and steady loops and raises the new location in between. In the steady loop the hook is raised between `send_forward` and `recv_backward` on purpose: `recv_backward` blocks on a gradient that only exists once the downstream stage has finished its backward, so a hook placed after it could not overlap the forward send. * With `batch_p2p_comm=True` the NCCL kernel runs on the calculation stream and no overlap is possible; the location is still raised so hooks that rely on it for correctness are not silently skipped. No behaviour change for code that registers no hook at this location. Tests: * `test_pipeline_parallel_micro_step_hooks.py` (single process) covers the enum member, the hook registry, hook ordering and kwargs, the updated assertion messages, global hook registration, and every branch of `P2pHelper.send_forward`'s new `overlap_p2p_comm` plumbing (`wait_on_reqs` forwarding, handle return, the batched-p2p and last-stage no-op cases). * `test_pipeline_parallel_p2p_issued_hook.py` launches `hybrid_parallel_pp_p2p_issued_hook.py` on 4 ranks (the schedule asserts `pp_degree > 2`) for both `use_batch_p2p_comm=False` and `True`. It checks the location fires exactly once per micro-step of every model chunk, that a hook enqueuing real GPU kernels in the overlap window does not corrupt the tensors in flight (training loss still matches the `forward_only` schedule), and that `forward_only` does not raise the location. * [Pipeline Parallel] Deduplicate the P2P_ISSUED raise-then-wait logic Extract "raise P2P_ISSUED, then consume the deferred wait handles" into a module-level helper `_raise_p2p_issued_and_wait`, and drop the batch_p2p_comm if/else at the warmup site: `overlap_p2p_comm=True` is a no-op under batch p2p (the ops run on the calculation stream, so no wait handle is returned), which makes both branches identical. Behaviour is unchanged; the helper is directly unit-testable in a single process, so the new schedule logic is now covered without a >=3 GPU runner.
S
SUN Dong committed
16786f7f52fd33dde592f9605247f2f191a12b47
Parent: 0182ee8
Committed by GitHub <noreply@github.com>
on 8/25/2026, 9:31:10 AM