[Cosmos3] Mixed W8A8/W8A16 denoising for ModelOpt FP8 checkpoints (#14664)
* Add opt-in mixed W8A8/W8A16 denoising for Cosmos3 ModelOpt FP8 checkpoints. Keep native ModelOpt GEMM on middle steps and dequant-linear W8A16 on the first/last steps so CFG cond/uncond share one precision per scheduler step. * Read Cosmos3 mixed-precision schedule from the checkpoint policy. Enable first/last W8A16 only when transformer/config.json declares diffusion_step_policy, so distilled FP8 stays native W8A8 instead of inheriting a hardcoded 3+3 window. * Harden Cosmos3 mixed-precision loading and document official Hub fp8 schedules. Read the checkpoint runtime policy from on-disk transformer/config.json when the live ModelOpt config omits it, fail closed on incomplete policies, and allow FP32 activations on the W8A16 path. * Address review: simplify FP8 mixed docs, no-op non-ModelOpt backends, drop unit tests. W8A8/W8A16 is explained without listing call-site overrides; TorchAO and other quantizers keep their native forwards; focused tests are removed until Hub fp8 usage is clearer. * Document FP8 mixed default vs none trade-offs without ModelOpt restore boilerplate. Keep serialized restore in the ModelOpt guide and spell out that all-W8A8 is faster but can flicker on multi-step video. * Apply style fixes * Regenerate Cosmos modular auto docstrings for mixed-precision inputs. --------- Co-authored-by: Sayak Paul <spsayakpaul@gmail.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Y
Yiming Zhao committed
759164b7ad116e091e9d3e222211c9aa27d835f6
Parent: 1f7be81
Committed by GitHub <noreply@github.com>
on 9/14/2026, 8:16:31 PM