Introduction

Eight steps. That's all it takes now. On October 9, 2026, Alibaba's Qwen team released Qwen-Image-2.1-Turbo, an accelerated open-weights checkpoint of Qwen-Image-2.1 that does both text-to-image and image editing in just 8 denoising steps. It keeps the same 7B visual generation architecture, ships its own sampling schedule baked into the checkpoint, and is already on Hugging Face, ModelScope, and in ComfyUI-ready form. The same day, Qwen switched on hosted Pro and Turbo APIs on Alibaba Cloud Model Studio.

What shipped

  • Model: Qwen-Image-2.1-Turbo (Alibaba Qwen)
  • Modality: image generation + natural-language image editing, one model
  • Base: Qwen-Image-2.1 (released September 20), same 7B visual generation architecture
  • Steps: 8, with the recommended schedule saved in the checkpoint and loaded automatically
  • Guidance: CFG=1 by default, plus prefix KV caching that reuses text and reference-image context across steps
  • Resolution: the 2.1 presets, from 2048×2048 square up to 2752×1536 at 16:9 (3:2 is 2528×1696)
  • Weights: open on Hugging Face and ModelScope under the Qwen Research License
  • Hosted: Qwen-Image-2.1 Pro and Turbo APIs on Alibaba Cloud Model Studio

Why 8 steps matters

Qwen-Image-2.1 was already pitched as the lean, cost-effective member of the family. Turbo goes further by distilling the sampling path down to eight passes. Because guidance runs at CFG=1, each step does one forward pass instead of two, so the time saved is bigger than the step count alone suggests.

Early local tests on X back that up. One Mac Studio M3 Ultra user reported about 120 seconds per image on 2.1 versus about 25 seconds on Turbo. Another got 2048×2048 frames in roughly 75 seconds on an RTX 3090. Those are community numbers, not official benchmarks, but they point the same way.

The showcase on the model card covers portraits, dance poses, transparent RGBA output, typography-heavy posters, UI layouts, single-image transformations, and multi-reference composition, including one interior scene built from four separate input images.

Run it today

ComfyUI

Comfy-Org has already added Turbo files to its Qwen-Image-2.1 repo:

  • diffusion_models/qwen_image_2.1_turbo_bf16.safetensors
  • diffusion_models/qwen_image_2.1_turbo_int8_convrot.safetensors (lighter INT8 build)
  • loras/qwen_image_2.1_turbo_lora_avg_rank_178_bf16.safetensors (a Turbo LoRA for the base 2.1 model)

ComfyUI has supported Qwen-Image-2.1 natively since day 0, so the existing 2.1 text-to-image and edit templates are the starting point. Swap in the Turbo diffusion model, or keep base 2.1 and load the Turbo LoRA, then drop the sampler to 8 steps with CFG 1. Community GGUF quants (Q8 and Q4) started landing on Hugging Face within hours for lower-VRAM cards.

Diffusers

Turbo needs the latest Diffusers from source:

pip install git+https://github.com/huggingface/diffusers.git
pip install "transformers>=5.17.0" accelerate pillow

import torch
from diffusers import QwenImage21Pipeline

pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1-Turbo",
    dtype=torch.bfloat16,
).to("cuda")

image = pipe(
    prompt="A lantern-lit night market on a rainy hillside street, steam rising from food stalls, warm gold light reflecting on wet stone, cinematic wide shot",
    width=2528,
    height=1696,
    use_kv_cache=True,
    generator=torch.Generator("cpu").manual_seed(42),
).images[0]
image.save("turbo.png")

One gotcha: setting num_inference_steps alone won't override the saved 8-step schedule. If you want to experiment, pass explicit sigmas at call time, and know that Qwen hasn't evaluated other schedules.

For editing, pass an image= input to the same pipeline with a natural-language instruction. The official example turns a pencil sketch of a yacht into a photoreal shot at sea.

Try this prompt

A cozy record shop at dusk seen through a rain-streaked front window, a woman in a mustard cardigan flipping through vinyl crates, neon "OPEN" glow reflected on the glass, warm tungsten interior against cool blue street light, shallow depth of field, 35mm film photograph

The license fine print

The open weights use the Qwen Research License, which allows research and evaluation only. Commercial use needs a separate license from Qwen or the hosted Model Studio APIs. Read it before you ship Turbo inside a product.

Original Source

Conclusion

Qwen-Image-2.1-Turbo turns a solid open image-and-edit model into something you can iterate with at sketch speed: eight steps, 2K output, transparent backgrounds, and multi-reference edits from a single 7B model. If you're already running 2.1 in ComfyUI, the Turbo files are one swap away. Just keep that research-only license in mind before anything leaves your sandbox.

—Aurelia ♡