SIGN IN SIGN UP

[tests] refactor pipeline-level quantization tests (#14435)

* fix(torchao): route root-level tensors around safetensors reconstruction

Models with parameters at the root of the module tree (e.g. Wan's
`scale_shift_table`) crashed torchao's `unflatten_tensor_state_dict`
when loading serialized checkpoints, since flattened tensor names are
assumed to carry a module prefix. Filter such tensors and their
metadata entries out of the reconstruction and merge them back
unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bnb): defer 8-bit weights split from their SCB stats across shards

Sharded serialization can place an 8-bit weight and its `SCB`
statistics in different shard files, in which case the shard-by-shard
loader failed with "Missing quantization component `SCB`". Hold the
incomplete half of the pair back until its counterpart arrives with a
later shard, mirroring the torchao pending mechanism, and disable
parallel shard loading for prequantized 8-bit checkpoints.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [tests] extend model-level quantization tester mixins

Migrate remaining model-level coverage from tests/quantization into the
tester mixins so it runs for every wired model:

- base: buffer-placement assertions in the device-map test, and an
  opt-in sharded-serialization test enabled by setting
  `sharded_serialization_config`
- bnb: serialization across all configs (sharded included), dtype
  assignment and adapter training for 8-bit, device moves preserving
  the memory footprint, corrupted-state-dict loading error, and a fixed
  modules-to-not-convert test (BitsAndBytesConfig only exposes
  llm_int8_skip_modules; the old test passed an unsupported kwarg and
  only survived by being skipped)
- torchao: custom device maps with cpu/disk offload, generalized from
  the Flux-specific test
- gguf: the diffusers-format single-file loading path, wired into the
  Flux model tests

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [tests] fix model-level quantization test wiring

- SD3.5: the quantized testers reused the random-init dummy inputs
  (4 latent channels, fp32) while the tiny Hub checkpoint has
  in_channels=8 and the quantizers load the model in half precision;
  give them matching inputs and relax the 4-bit memory expectation for
  the tiny checkpoint.
- QwenImage / Flux2: the quantized testers had no Hub checkpoint wired
  at all, so every test errored; point them at
  hf-internal-testing/tiny-qwenimage-pipe and tiny-flux2 with matching
  inputs.
- NucleusMoE: no tiny checkpoint exists on the Hub yet; comment the
  testers out like the LTX ones.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [tests] split tests/quantization into pipeline-level and backend-level tiers

Pipeline-level quantization tests (pipeline quality slices, cpu offload,
LoRA loading, compile, PipelineQuantizationConfig) move to
tests/pipelines/testing_utils/quantization.py, marked per backend so the
nightly CI can select them with `pytest -m`. tests/quantization keeps
only backend-level tests that fit neither tier: config validation,
utility warnings, and GGUF CUDA kernel correctness. Tests already
covered by the model-level mixins are dropped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [ci] select nightly quantization tests by marker across test tiers

Each backend job now runs `pytest -m <marker>` over tests/models,
tests/quantization, and tests/pipelines/testing_utils/quantization.py,
giving the model-level mixin tests a nightly home with the backend
dependencies installed. The torchao job additionally installs mslk.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: sayak@huggingface.co <sayak@ip-10-53-90-45.ec2.internal>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
S
Sayak Paul committed
425b113c211c42ec26ff646cbdbaa7d83adb46b5
Parent: a14d630
Committed by GitHub <noreply@github.com> on 8/18/2026, 6:53:32 PM