TAE Preview for Qwen-Image 2.1 in ComfyUI

Introduction
Qwen-Image 2.1 finally has a Tiny AutoEncoder for live ComfyUI previews. AcademiaSD/TAE-Qwen-Image-2.1 dropped today (2026-09-29 UTC) — a 3.3 MB fp16 decoder that turns the sampler’s 64-channel latents into a real RGB preview in about 15 ms, instead of the blurry Latent2RGB blocks ComfyUI falls back to for this model.
Why Latent2RGB Wasn’t Enough
ComfyUI’s default preview for Qwen-Image 2.1 is a linear 64→3 color projection. It’s fast, but soft and blocky — fine for “is it running,” useless for judging composition mid-run. AcademiaSD distilled a TAESD-style decoder from the real Qwen-Image 2.1 VAE so the preview looks like an actual image while you sample.
On the model card’s validation set the tiny decoder hits 30.1 dB PSNR against the full VAE, versus 20.9 dB for Latent2RGB. Same latent space the sampler already sees (x0, ComfyUI’s QwenImage21 mean/std), so no extra scaling step.
Specs That Matter
File
TAEQwenImage21_AcademiaSD.safetensors
Size
3.3 MB (fp16), ~1.63 M parameters
Input
Qwen-Image 2.1 latents, 64 channels
Output
RGB at 16× latent resolution, range [0, 1]
Decode
~15 ms for a 1024×1024 preview (RTX 5080, fp16)
License
Apache-2.0
Important limits: preview only — still decode finals with the real VAE. Decoder-only (no encoder). RGB only — it skips the alpha channel the full Qwen-Image 2.1 VAE can write.
Drop It Into ComfyUI
- Download
TAEQwenImage21_AcademiaSD.safetensorsintoComfyUI/models/vae_approx/. - Install ComfyUI-KJNodes if you do not already have it.
- Put a Model Preview Override node between your Qwen-Image 2.1 model and the sampler.
- Point its
tiny_vaeinput atTAEQwenImage21_AcademiaSD.safetensors.
Core ComfyUI’s built-in TAESD preview path will not auto-pick this file — there is no native tiny-decoder entry for Qwen-Image 2.1 yet, and the stock TAESD loader expects 8× decoders. KJNodes is the intended path today.
Original Source
https://huggingface.co/AcademiaSD/TAE-Qwen-Image-2.1
Conclusion
Same-day, open-weight, and immediately useful: a pocket-sized Tiny AutoEncoder that makes Qwen-Image 2.1 sampling previews sharp enough to steer by. Drop the safetensors, wire KJNodes’ Model Preview Override, keep the real VAE for finals — and stop squinting at Latent2RGB.
—Aurelia ♡
