Introduction

What if a 20-billion-parameter image model could sketch and finish a picture in two quick strokes? That's the promise of PVD (Phase-wise Velocity Distillation), a NeurIPS 2026 paper from The Hong Kong Polytechnic University's VCLab and OPPO Research Institute. The team published open weights on Hugging Face on October 5, with distilled students for FLUX.1-dev, Qwen-Image, and SD3.5 Medium, plus the full training and inference code on GitHub.

The headline: each PVD student generates a 1024×1024 image for roughly the compute of one forward pass of its original teacher, while cutting peak VRAM by about 46–48%. It's image generation, open weights, and you can run it locally today with the official Python scripts.

How two halves beat one

Most speed-up tricks squash a 30-step sampler into one or two steps with the same giant backbone. PVD does something sneakier. It slices the teacher down to a half-depth backbone, then trains two "phase experts" on different stretches of the denoising path:

  • Expert one drafts. It handles the early, noisy phase where global layout and composition get decided.
  • Expert two refines. It picks up the first expert's intermediate state and polishes texture and detail.

Each expert learns the average velocity over its own time window, and a phase-specific adversarial loss sharpens the text-to-image results. At inference the two run back to back, once each. For FLUX.1-dev and Qwen-Image, the two experts are LoRA adapters on one shared distilled backbone (rank 64 for FLUX, rank 32 for Qwen-Image), so you only load the big weights once.

The numbers

From the repo's cost table (teacher → PVD):

  • FLUX.1-dev: 11.90B → 5.98B active parameters, peak VRAM 22.64 → 12.28 GiB, 1.01 forward-pass equivalents.
  • Qwen-Image: 20.43B → 10.40B active parameters, peak VRAM 38.42 → 19.84 GiB, 1.02 forward-pass equivalents.
  • SD3.5 Medium: 2.24B → 1.10B, peak VRAM 4.99 → 2.67 GiB.

Quality holds up surprisingly well. On GenEval, the Qwen-Image student scores 0.8846 against the full multi-step teacher's 0.8720, and it beats the TwinFlow few-step baseline on GenEval, WISE, Qwen-Image-Bench, and aesthetic score. The FLUX student keeps its Qwen-Image-Bench score at 43.14 versus the teacher's 43.52, though it trails the teacher a little on ImageReward (0.6747 vs 0.8139). On ImageNet 256×256, the class-conditional model reaches an FID of 1.48 in a single forward-equivalent, close to the 500-pass LightningDiT teacher's 1.35.

There are also "Unsplash" variants of all three, further trained on high-quality photos for more photoreal output.

Try it today

There's no ComfyUI node yet, so this is a terminal job for now. You'll need a CUDA GPU, the base model's VAE and text encoders, and the PVD weights placed in the code repo root:

git clone https://github.com/PolyU-VCLab/PVD
cd PVD
pip install -r requirements.txt
hf download VCLab-PolyU/PVD --include "weights/pvd_flux/*" --local-dir .

python infer.py --task flux --model-root models/FLUX.1-dev \
  --prompt "A red apple on a wooden table" --output outputs/flux.png

Swap --task to qwenimage, sd35, or an _unsplash variant to pick a different student. The --part1_steps and --part2_steps flags set how many steps each expert takes, and both default to 1.

A quick license note

The Hugging Face card lists Apache-2.0, but the FLUX.1-dev student is still built from FLUX.1-dev, whose own license limits commercial use. Read the base model's terms before shipping anything made with the FLUX variant.

Why creators should care

Halving VRAM is the part that matters most for home setups. A FLUX.1-dev student that peaks around 12 GiB is suddenly in reach of mid-range cards, and a Qwen-Image student under 20 GiB fits on a single 24 GB GPU without heavy quantization. The "draft, then refine" split is also an easy idea to borrow: expect someone to wrap these adapters in a ComfyUI node soon.

Original Source

https://huggingface.co/VCLab-PolyU/PVD

Code and training scripts: PolyU-VCLab/PVD on GitHub

Conclusion

PVD is a neat reminder that speed doesn't always mean fewer, blurrier steps. Sometimes it means giving each half of the job to a specialist. Grab the FLUX or Qwen-Image student, time a render against your usual 28-step workflow, and see how much of the quality survives on your own prompts.

Want to play with Flux without touching a terminal? Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=pvd-phase-wise-distillation-flux-qwen-image

—Aurelia ♡