SIGN IN SIGN UP

[LoRA] add LoRA training for Qwen Image 2.1 (#14808)

* [LoRA] add LoRA training for Qwen Image 2.1

Add DreamBooth LoRA training for Qwen-Image 2.1, text-to-image and
image-to-image, with fast tests and a README section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* default rank and alpha to 16 and document the ratio

`--rank` and `--lora_alpha` are independent arguments, so raising the rank
alone leaves the update scaled by `lora_alpha / rank`. Default both to 16,
which keeps the scale at 1, and add a README section explaining the ratio.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* fix validation when embeddings are precomputed or only the final pass runs

Two fixes from review:

`QwenImage21ValidationPipeline` substituted the cached image-pad mask after
calling the base `encode_prompt`, but that call raises for a missing mask when
it is handed `prompt_embeds` together with a condition image, so the
substitution never ran and image-to-image validation failed. Pass the mask in
before the call instead.

`--final_validation_prompt` was accepted by the final-inference guard, but the
prompt embeddings were only built under `--validation_prompt`, so the final pass
reached a pipeline whose text encoder is `None` with no prompt at all. Build the
embeddings for whichever prompt is set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* fix encoding under --offload without --cache_latents

`vae.encode` sat outside the `offload_models` block, so the context manager had
already moved the VAE back to the CPU by the time it ran, while the prepared
dataloader hands the batch over on the accelerator:

    RuntimeError: Input type (torch.cuda.FloatTensor) and weight type
    (torch.FloatTensor) should be the same

Move the call inside the block and cover that flag combination, which is the
only path that encodes pixels inside the training loop. The image-to-image
trainer already had it right.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* add LoRA tests for the Qwen-Image 2.1 pipeline

Follow the Flux layout and reuse `LoraTesterMixin` and `LoraMemoryTesterMixin`,
so the adapters these trainers produce are covered by the pipeline's own
loading, fusing and memory-offload tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* fix caching a prompt mask that `encode_prompt` returns as None

`encode_prompt` drops the prompt mask when nothing in the batch is padded,
which is the common case for `--caption_column` datasets since captions in a
bucket often tokenize to the same length. The cache is filled per sample, so
slicing that `None` raised during latent caching:

    prompt_embeds_mask_cache[idx] = prompt_embeds_mask[i : i + 1]
    TypeError: 'NoneType' object is not subscriptable

Materialize the mask before filling the cache, and cover the path with a test:
none of the existing ones passed `--caption_column`, which is why this only
showed up on a real dataset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* pass rank and alpha explicitly in the README dog example

The dog example was written for the old 4/4 defaults; now that both default
to 16, spell out 4/4 so the documented command keeps training the size it
was tuned for.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* drop copy-paste leftovers from other trainers

`--output_dir` defaulted to `hidream-dreambooth-lora`, the weight-decay help
mentioned UNet params, and the VAE cast comment talked about the Flux VAE.
All three were inherited verbatim from the trainer this was derived from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* lower the README dog example to lr 1e-4

At 2e-4 the example memorizes the five training images: prompted scenes
collapse to the training backdrop rather than following the prompt. 1e-4
is what the image-to-image example already uses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qnMKe5MVc7B4XXHGPvXuZ

* skip the LoRA scale test for the Qwen-Image 2.1 dummy

Halving the LoRA scale moves this dummy transformer's output by less than the
tolerance the assertion allows, so the two outputs compare equal and the test
fails on CPU. Skip it with the measurement recorded rather than loosening a
tolerance shared with every other pipeline's LoRA tests.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
L
Linoy Tsaban committed
e0118ade2f60234c41bacf40330a7e2f61108849
Parent: 0377f0c
Committed by GitHub <noreply@github.com> on 9/24/2026, 3:03:09 AM