SIGN IN SIGN UP

[Distributed] Add P2P_ISSUED pipeline-parallel micro-step hook location (#79690)

* [Pipeline Parallel] Add P2P_ISSUED micro-step hook location

Adds a new `PipelineParallelMicroStepLocations.P2P_ISSUED` callback
location, raised after a micro-step's forward P2P has been *issued* and
before its wait handles are consumed.

Motivation: user code that wants to hide an independent side computation
(e.g. an auxiliary-loss branch whose gradient does not reach the
backbone) behind the pipeline's forward send/recv has no place to put it
today. `FORWARD_END` looks like the right spot but cannot work: `isend`
/ `irecv` record an event on the calculation stream at issue time and
have the pipeline group's comm stream wait on it, so anything enqueued
before the issue is inside that event's reach and the NCCL kernel waits
for it instead of running concurrently. Hooks must therefore run *after*
the issue and *before* the wait -- the same "issue, compute, then wait"
shape the schedule already uses for `WeightGradStore.pop()`.

Changes:

* `PipelineParallelMicroStepLocations.P2P_ISSUED` plus its entry in the
  hook registry and the two assertion messages.
* `P2pHelper.send_forward` gains `overlap_p2p_comm=False`, mirroring the
  option `send_forward_recv_forward` already has: it forwards
  `wait_on_reqs=not overlap_p2p_comm` to `_p2p_helper` and returns the
  wait handles so the caller can defer the wait. Default keeps the
  previous blocking behaviour and return value.
* `VPPFhenBInBalancedMemory` defers the forward-send wait in its warmup
  and steady loops and raises the new location in between. In the
  steady loop the hook is raised between `send_forward` and
  `recv_backward` on purpose: `recv_backward` blocks on a gradient that
  only exists once the downstream stage has finished its backward, so a
  hook placed after it could not overlap the forward send.
* With `batch_p2p_comm=True` the NCCL kernel runs on the calculation
  stream and no overlap is possible; the location is still raised so
  hooks that rely on it for correctness are not silently skipped.

No behaviour change for code that registers no hook at this location.

Tests:

* `test_pipeline_parallel_micro_step_hooks.py` (single process) covers
  the enum member, the hook registry, hook ordering and kwargs, the
  updated assertion messages, global hook registration, and every branch
  of `P2pHelper.send_forward`'s new `overlap_p2p_comm` plumbing
  (`wait_on_reqs` forwarding, handle return, the batched-p2p and
  last-stage no-op cases).
* `test_pipeline_parallel_p2p_issued_hook.py` launches
  `hybrid_parallel_pp_p2p_issued_hook.py` on 4 ranks (the schedule
  asserts `pp_degree > 2`) for both `use_batch_p2p_comm=False` and
  `True`. It checks the location fires exactly once per micro-step of
  every model chunk, that a hook enqueuing real GPU kernels in the
  overlap window does not corrupt the tensors in flight (training loss
  still matches the `forward_only` schedule), and that `forward_only`
  does not raise the location.

* [Pipeline Parallel] Deduplicate the P2P_ISSUED raise-then-wait logic

Extract "raise P2P_ISSUED, then consume the deferred wait handles" into a
module-level helper `_raise_p2p_issued_and_wait`, and drop the
batch_p2p_comm if/else at the warmup site: `overlap_p2p_comm=True` is a
no-op under batch p2p (the ops run on the calculation stream, so no wait
handle is returned), which makes both branches identical.

Behaviour is unchanged; the helper is directly unit-testable in a single
process, so the new schedule logic is now covered without a >=3 GPU runner.
S
SUN Dong committed
16786f7f52fd33dde592f9605247f2f191a12b47
Parent: 0182ee8
Committed by GitHub <noreply@github.com> on 8/25/2026, 9:31:10 AM