SIGN IN SIGN UP

[Cosmos3] Mixed W8A8/W8A16 denoising for ModelOpt FP8 checkpoints (#14664)

* Add opt-in mixed W8A8/W8A16 denoising for Cosmos3 ModelOpt FP8 checkpoints.

Keep native ModelOpt GEMM on middle steps and dequant-linear W8A16 on the first/last steps so CFG cond/uncond share one precision per scheduler step.

* Read Cosmos3 mixed-precision schedule from the checkpoint policy.

Enable first/last W8A16 only when transformer/config.json declares diffusion_step_policy, so distilled FP8 stays native W8A8 instead of inheriting a hardcoded 3+3 window.

* Harden Cosmos3 mixed-precision loading and document official Hub fp8 schedules.

Read the checkpoint runtime policy from on-disk transformer/config.json when the live ModelOpt config omits it, fail closed on incomplete policies, and allow FP32 activations on the W8A16 path.

* Address review: simplify FP8 mixed docs, no-op non-ModelOpt backends, drop unit tests.

W8A8/W8A16 is explained without listing call-site overrides; TorchAO and other
quantizers keep their native forwards; focused tests are removed until Hub fp8
usage is clearer.

* Document FP8 mixed default vs none trade-offs without ModelOpt restore boilerplate.

Keep serialized restore in the ModelOpt guide and spell out that all-W8A8 is faster but can flicker on multi-step video.

* Apply style fixes

* Regenerate Cosmos modular auto docstrings for mixed-precision inputs.

---------

Co-authored-by: Sayak Paul <spsayakpaul@gmail.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Y
Yiming Zhao committed
759164b7ad116e091e9d3e222211c9aa27d835f6
Parent: 1f7be81
Committed by GitHub <noreply@github.com> on 9/14/2026, 8:16:31 PM