Introduction

MiniMax H3 just moved out of the data center and onto your desk. Hao AI Lab at UC San Diego (the FastVideo team) published FastH3 on Consumer Hardware on October 6. It brings FastH3 V2, their open 8-step distill of MiniMax H3 that makes video with synchronized audio, to a single RTX 5090, RTX 4090, RTX PRO 6000, DGX Spark or Apple Silicon Mac. They also shipped FastH3 Trim, an experimental pruned version that's 4.2× smaller than base H3 and runs in as little as 8 GB of GPU memory.

This is text-to-video with audio, and the weights are open under the MiniMax H3 Community License. You can run it today in two ways: through the FastVideo Python library, or in ComfyUI with ready-made single-file checkpoints on Hugging Face.

What shipped

The catch with FastH3 V2 used to be size. Its BF16 weights total 137.7 GiB, and a 32 GB RTX 5090 couldn't even hold the diffusion transformer. The new release splits the work by hardware:

  • NVFP4 for Blackwell cards (RTX 50 series, RTX PRO 6000) and DGX Spark
  • FP8 for the RTX 4090 and GPUs with less memory
  • INT6 via MLX for Apple Silicon
  • FastH3 Trim: 42 of H3's 50 transformer blocks, with a rank-16 timestep factorization. It's faster than V2 on every device, at what the team calls minimal cost in quality.

The team says FastH3 V2 already beats base H3 on quality at eight steps. That's their claim, not an independent benchmark. They're also upfront that Trim is experimental and can show less detail in busy scenes, so use V2 when quality matters most.

The speed numbers

All times are end to end on a warm server, from prompt to finished MP4 with audio. Each is the median of two runs on each of two prompts. Format: V2 / Trim.

5-second clip at 832×480:

  • RTX PRO 6000 (96 GB, NVFP4): 13.5 s / 12.0 s
  • RTX 5090 (32 GB, NVFP4): 14.8 s / 13.4 s
  • RTX 4090 (24 GB, FP8): 54.6 s / 43.9 s
  • RTX 4090 capped at 16 GB: 79.9 s / 72.1 s
  • RTX 4090 capped at 12 GB: 91.2 s / 74.1 s
  • RTX 4090 capped at 8 GB: Trim only, 82.0 s
  • DGX Spark (128 GB unified): 141.4 s / 125.8 s
  • 2× DGX Spark: 87.2 s / 78.3 s
  • Mac M4 Max (36 GB, INT6): Trim only, 925.2 s

5-second clip at 1344×768:

  • RTX PRO 6000: 36.5 s / 32.5 s
  • RTX 5090: 38.6 s / 35.4 s
  • RTX 4090: 154.6 s / 132.8 s
  • DGX Spark: Trim only, 340.1 s

One honest footnote: the 16, 12 and 8 GB rows come from capping memory on the same RTX 4090. A real card with that much memory will be slower.

How they shrank it

H3 is really three networks: a Qwen3-VL text encoder, a 50-block diffusion transformer that denoises video and audio together, and two VAEs. Each one got its own diet:

  • Text encoder, 62.1 → 15.3 GiB. H3 only reads hidden state 50 of the 64-layer Qwen3-VL, so they cut the last 14 layers and the language-model head, and stored the rest in NVFP4.
  • Transformer. Attention, MLP and sparse-attention gate weights go to NVFP4 (FP8 where there's no FP4 support). Trim's transformer is 11.1 GiB in NVFP4.
  • VAEs, 10.3 → 6.5 GiB. Video decoding uses the LynnReal lightweight video VAE with Kijai's INT8 weights. It takes 2.3 GiB of GPU memory.
  • No clipping at 4 bits. In block 37, activations reach 368,640, which is 137 times what a unit NVFP4 scale can hold. So the team ran 1,000 prompts through the full sampler and stored one calibrated scale per layer (294 layers in Trim).

That diet is what makes the 5090 work. The first FP4 export left a 20 GB transformer that got shuffled to host memory on every request, so a 480p clip took 26.4 s and 768p didn't fit at all. With everything quantized, the transformer stays on the GPU and the same clip takes 13.4 s.

Run it in ComfyUI

The ComfyUI repacks are already up. Load the FastVideo FastH3: Text to Video template (ComfyUI 0.36.0 or later) and pick a file in its UNETLoader:

  • FastH3 Trim NVFP4 (fastvideo_fasth3_trim_8step_nvfp4.safetensors, 11.9 GB): RTX 50 series, RTX PRO 6000, DGX Spark
  • FastH3 Trim FP8 (19.7 GB): RTX 40 series
  • FastH3 Trim INT8 ConvRot (18.9 GB): any recent NVIDIA GPU, and the template's default format
  • FastH3 V2 NVFP4 (fastvideo_fasth3_8step_v2_pruned_nvfp4.safetensors, 13.6 GB): Blackwell GPUs and DGX Spark

Keep the template's sampler settings: 8 steps, res_multistep, simple, and the MiniMaxH3SigmaShift node at 10 for video and 3 for audio. The text encoder and VAEs sit in their usual folders. For example, a Blackwell setup with Trim looks like this:

ComfyUI/models/diffusion_models/fastvideo_fasth3_trim_8step_nvfp4.safetensors
ComfyUI/models/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
ComfyUI/models/vae/minimax_h3_video_vae_int8_convrot.safetensors
ComfyUI/models/vae/minimax_h3_audio_vae_fp32.safetensors

Run it in Python

Install FastVideo with uv. Use cu130 instead of cu126 on CUDA 13. DGX Spark needs a from-source install, and Macs follow the MLX guide.

UV_TORCH_BACKEND=cu126 uv pip install fastvideo

Here's the team's single-RTX-5090 recipe from the blog:

import os
from huggingface_hub import hf_hub_download
repo = "FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer"
os.environ["FASTVIDEO_H3_PARK_MODULES"] = "vae,audio_vae"  # keep the transformer on the GPU during text encoding
os.environ["FASTVIDEO_H3_ENCODER_LAYERWISE"] = "1"  # stream the text encoder layer by layer
os.environ["FASTVIDEO_H3_ADALN_TABLE"] = hf_hub_download(repo, "transformer/adaln_tables.pt")  # skip 26 GB of AdaLN weights
from fastvideo import VideoGenerator
generator = VideoGenerator.from_config({
    "model_path": repo,
    "engine": {
        "num_gpus": 1,
        "quantization": {"transformer_quant": "NVFP4", "layer_profile": "h3_dit_vsa"},
        "offload": {"text_encoder": True, "pin_cpu_memory": True},
    },
    "pipeline": {"experimental": {"attention_backend": "VIDEO_SPARSE_ATTN_H3", "h3_sequential_load": True}},
})
generator.generate_video(
    prompt="A red fox leaps into deep snow at sunrise and pops back up with snow on its face.",
    height=480, width=832, num_frames=124, guidance_scale=1.0, output_path="out.mp4",
)

For the faster experimental model, swap in FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4 and drop the AdaLN table line. On an RTX 4090, use the -FP8 repos instead. Mac fans should know that the MLX INT6 repos named in the post weren't publicly reachable on Hugging Face when I checked at 17:00 MYT, so keep an eye on the FastH3 collection.

Why it matters

Until this release, FastH3 V2 needed data-center GPUs. Now a 5-second clip with audio takes about 13 seconds on one RTX 5090, and even the team's 12 GB-capped test finished a 480p V2 clip in 91.2 s. Trim is the team's first step toward smaller models, and they say they expect better quality and smaller releases from here. No local GPU? H3 runs in the browser on Gen.

Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=fasth3-v2-trim-consumer-gpus-rtx-mac

Original Source

Conclusion

FastH3 V2 on consumer hardware is the kind of release that changes what a home studio can do. You get eight steps, synced sound, and real numbers on real cards. Trim makes it even smaller if you're willing to trade a little detail. Grab the Comfy checkpoint for your GPU, keep the sampler at 8 steps, and let a fox leap into some snow tonight.

—Aurelia ♡