Introduction

MiniMax H3 makes gorgeous video with sound, but on a home GPU it takes its sweet time. That just got a lot better. The Veda Sparse Attention team has published an official ComfyUI custom node, Veda Sparse Attention (MiniMax H3), on the Comfy Registry. It went live on October 3, and version 0.2.0 followed on October 4.

The idea is simple: most of H3's attention work barely changes the final frame, so Veda skips it. On the team's own RTX 5070 12 GB test, an 8-step 1344×768 clip went from 342 seconds to 130 seconds, about 2.9x faster end to end. The node code is MIT-licensed, the 275 MB predictor is open on Hugging Face, and it runs locally in ComfyUI 0.38.0 or newer today.

What Veda actually does

Veda comes from the ICML 2026 paper "Veda: Scalable Video Diffusion via Distilled Sparse Attention." A small predictor, distilled from the full model, looks at the attention map in tiles and guesses which ones really matter. It keeps roughly the top 10%, and a tile-skipping kernel only fetches those.

Because Veda changes how attention is computed and leaves the weights alone, it's not a LoRA. Your Turbo LoRA, style LoRAs, fine-tuned or quantized H3 checkpoints, and first/last-frame or reference conditioning all keep working.

The numbers

These are the repo's figures for text-to-video-with-audio at 1344×768, 124 frames (5.2 seconds), with the 8-step Turbo LoRA at 90% sparsity on an RTX 5070 12 GB under Windows 11:

  • ComfyUI default attention: 40.7 s per step, 342 s for 8 steps
  • ComfyUI with --use-sage-attention: 24.8 s per step, 231 s for 8 steps
  • Veda sparse INT8: 14.0 s per step, 130 s for 8 steps

On attention alone that's about 7.1x. The team says the gap grows with longer clips, and the Hugging Face predictor card reports up to 6.8x attention speedup and 3.1x end to end on an RTX PRO 6000 Blackwell with 14-second clips.

How to install it

You need ComfyUI 0.38.0 or newer. Search "Veda" in ComfyUI Manager, or use the CLI:

comfy node install veda-sparse-attention

The node doesn't download anything on its own. The easiest route is Workflow → Browse Templates → Veda-on-ComfyUI, then pick the T2VA or R2VA template, and ComfyUI's missing-model dialog will offer the predictor. To grab it by hand:

hf download Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview \
  minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors \
  --local-dir ComfyUI/models/veda

Where it goes in your graph

It's one node with MODEL in and MODEL out. Put it on the MODEL wire after your model and LoRA loaders, last thing before the guider or sampler. Select it and press Ctrl+B to bypass it, and you can render the same seed with full attention for a side-by-side check.

The advanced inputs default to the trained values:

  • generated_sparsity (90%) sets how much of the generated video's attention gets skipped.
  • reference_sparsity (90%) does the same for first/last frames and reference images or videos; set it to 0% to give references full attention.
  • full_attention_layers and full_attention_steps keep chosen DiT blocks or sampling steps at full attention.
  • verbose prints per-phase timing after each run.

One gotcha: don't stack it with ComfyUI's built-in Model Sparse Attention node on H3. That node replaces the attention blocks outright, so Veda would never run. The Veda node warns you if it sees both.

Hardware and fine print

A single Triton kernel targets every NVIDIA GPU from SM80 (RTX 30 / A100) onward on Windows and Linux, and Apple silicon runs through MLX. So far the authors have verified it on an RTX 5070 under Windows 11 and an M3 Pro under macOS 15. RTX 30/40, H100 and B200-class cards have the code path ready but aren't verified yet. Older cards, ROCm and CPU fall back to H3's normal attention, and the node tells you so instead of producing a broken render.

The predictor is labeled a preview. It was trained for 1344×768, 768×1344, 768×768 and 1024×768 at 5, 10 and 14 seconds with the 8-step Turbo LoRA. Other sizes, step counts, and FL2VA or R2VA references still work, but the team suggests comparing those against full attention. A dedicated R2VA fine-tune is promised for a future release. The predictor inherits the MiniMax H3 Community License from its base model.

Original Source

Conclusion

Veda is the rare speedup that asks for nothing in return: no retraining, no new checkpoint, and no giving up your favorite LoRAs. If H3 renders have been testing your patience on a 12 GB card, this is the node to try first, with a quick Ctrl+B comparison to keep yourself honest.

Want to play with MiniMax H3 without touching your GPU at all? Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=veda-sparse-attention-comfyui-minimax-h3

—Aurelia ♡