GRACE Makes Wan2.1-14B Video Up to 15x Faster

Introduction
What if your 14B video model rendered a clip in about a minute and a half instead of a quarter of an hour, with no visible drop in quality? That's the promise of GRACE (Generation-Aware Latent Compression), a new open research release from KAIST AI and Kakao. It retrofits Wan2.1-14B so the diffusion transformer works on 8x fewer latent tokens, and the team reports 11.1x faster generation at 480p and 15.5x faster at 720p-class resolution while matching the original model on VBench. Code is MIT, the GRACE checkpoints are Apache-2.0, and you can run text-to-video and image-to-video locally today from the GitHub repo.
What GRACE Actually Changes
Video diffusion models spend most of their time chewing on latent tokens. Squeeze the video harder in the autoencoder and the transformer has far less to process, but that usually wrecks quality or forces you to retrain the whole model from scratch.
GRACE gets around that with a clever two-stage trick:
- Dual latent: it keeps Wan2.1's original "base" latent frozen, so the transformer stays in a space it already knows, and adds a learned residual latent that carries the detail and motion that heavier compression would throw away.
- Generation-aware alignment: instead of training the autoencoder only to reconstruct pixels, it matches the compressed latent to the original one inside the frozen transformer's own feature space. The paper's punchline is that better reconstruction does not mean better generation.
- Base-ahead denoising: at inference, the base latent is denoised slightly ahead of the residual (a fixed offset of 0.15), so fine detail lands on content that has already settled. The DiT itself only gets a light LoRA adaptation.
The net effect is compression from 8x spatial and 4x temporal to 16x spatial and 8x temporal, taking a 480x832, 81-frame clip from 32.8k tokens down to 4.3k.
The Numbers
All timings are end to end for one 81-frame video on a single A100 80GB at 50 steps, CFG 5, batch 1, bf16:
- 480x832 text-to-video: Wan2.1-14B takes 851.5s, GRACE takes 75.8s
- 480x832 image-to-video: Wan2.1-14B takes 863.2s, GRACE takes 77.7s
- 736x1280 text-to-video: Wan2.1-14B takes 3,361s (about 56 minutes), GRACE takes 215.6s
- 736x1280 image-to-video: Wan2.1-14B takes 3,397s, GRACE takes 218.8s
- VBench-T2V total at 480p: 85.81 for GRACE versus 83.93 for Wan2.1-14B
- VBench-I2V total at 480p: 87.90 for GRACE versus 87.92 for Wan2.1-14B
GRACE is also faster than LTX-Video 0.9.7 at the same 4.3k token count (99.6s for text-to-video at 480p) while scoring higher on VBench-T2V. In a blind user study with 39 raters, text-to-video viewers picked GRACE over the uncompressed Wan2.1-14B for visual quality in 49.4% of votes and preferred the original in 35.3%. Image-to-video was closer, leaning slightly toward the original, so treat this as "roughly matches" rather than "beats" for I2V.
How to Run It Today
GRACE runs through a bundled fork of DiffSynth-Studio, and the authors warn you to use that copy rather than a pip install, because upstream silently ignores two settings the checkpoints need. You'll also need the official Wan2.1 base weights, since the GRACE repo doesn't redistribute them.
git clone https://github.com/cvlab-kaist/GRACE.git
cd GRACE
pip install -r requirements.txt
huggingface-cli download Wan-AI/Wan2.1-T2V-14B --local-dir ./Wan2.1-T2V-14B
huggingface-cli download Wan-AI/Wan2.1-I2V-14B-480P --local-dir ./Wan2.1-I2V-14B-480P
export GRACE_WAN_T2V_DIR=./Wan2.1-T2V-14B
export GRACE_WAN_I2V_DIR=./Wan2.1-I2V-14B-480P
bash scripts/generate_t2v.sh "a shark is swimming in the ocean" outputs/t2v
bash scripts/generate_i2v.sh photo.png "a penguin walking on a beach" outputs/i2vThe GRACE DiT, VAE, and decoder download automatically on first run. Defaults are 480x832, 81 frames, 50 steps, CFG 5. For the larger size, prefix the command with HEIGHT=736 WIDTH=1280.
Things to keep in mind
- It's built on Wan2.1, not the newer Wan releases, so think of it as proof that the trick works rather than a drop-in for your current favorite model.
- There's no ComfyUI node yet, and the Hugging Face demo and training code are still listed as coming soon.
- The quoted speeds are on an A100 80GB. A 14B model still wants serious VRAM, so consumer cards will need the usual offloading tricks.
Why It Matters for Creators
Most speedups we cover come from step distillation, meaning fewer denoising steps. GRACE attacks a different bottleneck, the sheer number of tokens per clip, and its gains grow with resolution. If this recipe gets applied to newer open video models, and ideally stacked with step distillation, high-res local video could go from coffee-break to near-interactive.
Original Source
- Paper: GRACE on arXiv (2610.10524)
- Code: cvlab-kaist/GRACE on GitHub
- Weights: chimaharicox/GRACE on Hugging Face
- Project page and gallery: GRACE project page
Conclusion
GRACE is one of those quietly big ideas: keep the model, shrink what it has to think about, and get a 15x speedup at high resolution without paying for it in quality. If you've got an A100-class box and some Wan2.1 prompts lying around, it's a fun one to try this week. I'll be watching for a ComfyUI port and for someone brave enough to try it on newer models.
—Aurelia ♡
