Diffusers 0.41: Qwen-Image 2.1 Lands, LTX-2.5 Gets DFR

Introduction
Hugging Face pushed Diffusers 0.41.0 to PyPI on October 6 at about 14:24 MYT, and it's a big one for anyone who runs open image and video models from Python. The headline: Qwen-Image 2.1 is now an official pipeline, QwenImage21Pipeline, with text-to-image, image editing, multiple reference images, native transparent (RGBA) output, and a ready-made LoRA training script.
The quick facts: Diffusers is Hugging Face's open-source (Apache-2.0) Python library for diffusion models. You can install it today with pip. The release also adds a new detail-boosting video pipeline for LTX-2.5 called DFR, and single-file loading for Krea 2 and MiniMax-H3 checkpoints. The model weights keep their own licenses, so Qwen-Image 2.1 still ships under the Qwen Research License.
Qwen-Image 2.1, the easy way
Until now, running Qwen-Image 2.1 outside ComfyUI meant custom code. Now it's a few lines. Per the docs, Qwen-Image 2.1 reads your prompt and any input images with a Qwen3-VL encoder, then paints the result with a 7B visual generator. Qwen's recommended defaults are baked in: 40 steps and no guidance.
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16).to("cuda")
# Text-to-image
image = pipe("A capybara wearing a wizard hat, oil painting").images[0]
image.save("t2i.png")
# Edit the same image with a new instruction
edited = pipe("Move it to a snowy mountain top", image=image).images[0]
edited.save("edit.png")A few things worth knowing:
- Multiple references: pass a list to
imageand each picture becomes its own block. Order matters, because the model reads them first to last. Think "put the flowers from the first image into the second scene". - Guidance is optional: pass a
negative_promptwithtrue_cfg_scaleabove 1 to switch on classic guidance. It roughly doubles the work per step. - Faster attention: there's an optional flex-attention processor that speeds things up, but only once you compile the model. Uncompiled, it can run out of memory at high resolution.
- Bring your ComfyUI file:
QwenImage21Transformer2DModel.from_single_file()loads the single bf16 checkpoint from the Comfy-Org repo, so you don't need two copies of the weights. - Custom step grids: you can pass your own
sigmaslist per call to experiment with schedules.
Train your own Qwen-Image 2.1 LoRA
The release adds train_dreambooth_lora_qwenimage21.py in the Diffusers examples folder. It follows the familiar DreamBooth LoRA recipe: point it at Qwen/Qwen-Image-2.1, give it a folder of images and an instance prompt, and go. Out of the box it uses LoRA rank 16, 512px resolution, and a 1e-4 learning rate, all of which you can change with flags. If you've been waiting for an official, scriptable path to Qwen-Image 2.1 styles and characters, this is it.
LTX-2.5 DFR: trade time for detail
The second big addition is Diffusion Fidelity Rendering for Lightricks' LTX-2.5, through new LTX2DFRPipeline and LTX2DFRTemporalRefinePipeline classes. The idea is simple once you see it. Video models squeeze many frames into each compressed latent frame. DFR adds a few extra keyframe slots, each spending a full latent frame on one real image frame, so the frames around it are built from genuinely new detail instead of being filled in between.
- Cost: more tokens. The docs give a worked example: at 1024×1536 and 121 frames, stage two grows from 24,576 to 32,256 tokens, about 31% more.
- Memory: if a resolution only just fits today, you may need CPU offload, VAE tiling, or a smaller canvas with DFR switched on.
- Recipe: the suggested starting point is 1088×1920 image-to-video, one 2x temporal refine round, and the 2x spatial detailing IC-LoRA on stage two.
- Requirement: it only works with LTX-2.5 checkpoints, and the pipeline refuses older ones.
Everything else in the box
- Single-file loaders for Krea 2 and MiniMax-H3, plus a MiniMax-H3 VAE decode fix.
- Tensor-parallel loading: each GPU reads only its slice of sharded weights, which cuts loading memory on multi-GPU rigs.
- Cosmos 3: SeaCache support and mixed FP8 denoising.
- Modular Diffusers: Wan 2.2 VACE gets modular blocks.
- Heads-up: ONNX support is now deprecated (use Optimum), and the old Lumina pipeline aliases are gone. Check your scripts before upgrading.
Hugging Face also says future minor releases will be timed around new model integrations, much like its Transformers library.
Run it today
pip install -U "diffusers==0.41.0"Then grab the Qwen-Image 2.1 snippet above, or open the LTX-2.5 section of the Diffusers docs for the full DFR recipe. You'll need an NVIDIA GPU with plenty of memory for both models.
Original Source
- Diffusers v0.41.0 release notes
- Diffusers on PyPI
- Qwen-Image 2.1 model card
- Qwen-Image 2.1 LoRA training script
Conclusion
Diffusers 0.41 turns two of this season's favorite open models into plain, scriptable Python: Qwen-Image 2.1 for images, edits, and transparent cutouts, and LTX-2.5 for sharper video when you can spare the time. If you'd rather skip the setup and play with Krea 2 or MiniMax-H3 in the browser, Gen has you covered.
Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=diffusers-0-41-qwen-image-2-1-ltx-2-5-dfr
—Aurelia ♡
