SIGN IN SIGN UP

[tests] Add Config to Exclude Modules from Leaf-Level Group Offloading (#14564)

* skip hunyuan framepack group offloading tests

* exclude only the SigLIP image encoder from leaf-level group offloading

`test_pipeline_level_group_offloading_inference` was skipped outright for
HunyuanVideoFramepack because `image_encoder` is a `SiglipVisionModel`, whose
attention pooling head wraps a `torch.nn.MultiheadAttention`. That hands
`self.out_proj.weight` to `torch.nn.functional.multi_head_attention_forward`
instead of calling `self.out_proj`, so the leaf-level onload hook on `out_proj`
never fires and its weights stay on the offload device.

Add a `group_offloading_leaf_level_exclude_modules` knob to the old-style
`PipelineTesterMixin` and the new-style `BasePipelineTesterConfig` (empty by
default, so no behavior change elsewhere), pass it through to
`enable_group_offload(exclude_modules=...)` in both implementations of the test,
and set it to `["image_encoder"]` for framepack instead of skipping. Block-level
offloading is unaffected — the whole head is onloaded as one unmatched module —
hence the level in the name.

The test now passes and covers leaf-level offloading of the transformer, VAE and
both text encoders. The VAE is coverage nothing else provided:
`test_group_offloading_inference` deliberately excludes `vae` and
`image_encoder`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: note how to handle a component that breaks leaf-level offloading

Records the rule behind `group_offloading_leaf_level_exclude_modules`: leaf-level
offloading hooks only the supported leaf types and onloads each on its own
forward, so any code that reads a leaf's `.weight` instead of calling the leaf
bypasses that hook.

Routes the fix by who owns the component. A diffusers model declares the gap with
`_supports_group_offloading = False` on the `ModelMixin` subclass, which both
offload mixins honor. A third-party component that can't be annotated goes in
`group_offloading_leaf_level_exclude_modules`, which keeps offload coverage for
every other component — where a hand-written skip would drop it for the whole
pipeline, the VAE included, since the component-scoped
`test_group_offloading_inference` deliberately excludes it.

`torch.nn.MultiheadAttention` is called out as the common instance rather than as
the definition, with `HunyuanDiTAttentionPool` as a case that fails the same way
with no MHA module involved, so the guidance still applies when a future
component fails for a different reason.

Also notes that a failure should be reproduced before a skip or exclusion is
added: of the five pipelines currently skipping the pipeline-level test, only
framepack and motif_video still fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Drop old style test info and simplify leaf level group offloading testing.md note.

---------

Co-authored-by: Sayak Paul <spsayakpaul@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
D
dg845 committed
0c8981bffe426fbd625ab8cc3512ffd5694d0996
Parent: 06e0f2a
Committed by GitHub <noreply@github.com> on 8/25/2026, 2:52:27 AM