FreeVideo Runs MiniMax H3 Locally on 8 GB GPUs

Introduction
Here's a happy one for anyone with a modest gaming PC. FreeVideo, an open-source app from FlashML-org, runs the MiniMax H3 text-and-image-to-video model (with audio) on your own machine, starting at 8 GB of VRAM and 16 GB of RAM. It installs as a ComfyUI plugin and comes with a one-click Windows launcher, an Apple-silicon Mac preview, and a Linux CLI.
The project first went public on October 2. The new part is the v0.3 line. v0.3.0 landed on October 7 (MYT) with a faster int8 model for GeForce cards, which the team says makes a 5-second clip about 2.3× faster on an RTX 3090. Since then it has shipped more than a dozen point releases, through v0.3.16 this afternoon. The repo has passed 1,200 stars, and it hasn't appeared in the magazine until now.
What Shipped
- Engine: a local inference engine for MiniMax H3, built on OpenVDN's 8-step VDN-H3 model and its Video DeltaNet hybrid attention (we covered VDN-H3 here)
- v0.3.0 (Oct 7 MYT): int8 weights for GeForce RTX 30/40/50-series cards with a one-click upgrade, faster generation on 8 GB cards, and optional prompt enhancement. A local model rewrites your prompt for H3. It's off by default and asks before downloading its 8.3 GiB model.
- v0.3.2 (Oct 7): Preview first for two-pass sampling. It renders a half-resolution draft, and if you like it, the second pass finishes a video identical to a direct render.
- v0.3.12 (Oct 8): Linux setup accepts older system Git, and offline installs can now download the 49.1 GB model set instead of importing packs
- Earlier releases: four quality levels (Light, Medium, High, Max), a Mac preview, community H3 LoRAs, and videos that carry their workflow, so you can drop one on the ComfyUI canvas to restore its prompt, seed, and settings
- License: Apache-2.0 for the code. The model weights stay under the MiniMax H3 Community License, which has territorial and acceptable-use limits.
How It Fits H3 Into 8 GB
FreeVideo's Adaptive Execution Planner re-measures your hardware before every request. It decides how many of H3's 50 transformer blocks stay in VRAM, which ones wait in pinned system RAM, and which stream from disk at each step. Then it picks the fastest attention kernel that passes an on-device probe, trying SageAttention 2 first.
- int8 GEMMs on GeForce Ada and Blackwell and on every Ampere card
- FP8 on workstation and datacenter cards
- The text encoder runs in its own process and exits before the video model loads
- Out-of-memory failures retry automatically with a lighter placement
The Numbers
From the project's end-to-end table (Windows, 1344×768, 10-second clip, two-pass):
- RTX 5090 (32 GB + 64 GB RAM): 122 s
- RTX 5060 Ti (16 GB + 32 GB): 486 s
- RTX 4060 Ti (16 GB + 32 GB): 558 s
- RTX 4060 Ti 8 GB (community report, 64 GB RAM): 603 s
Community reports at the 8 GB VRAM + 16 GB RAM floor work too, just slowly. One RTX 5060 rendered a 15-second 1344×768 clip on Light in 1,367 s, with an estimated 1,160 s after the v0.3.6 optimizations.
On the Mac side, an M5 with 24 GB of unified memory takes about 22 minutes for 10 seconds at 960×544, or about 42 minutes at 1344×768. Macs older than the M5 run BF16 and take longer.
Run It Today
- Windows 10/11 + NVIDIA: download
FreeVideo.exefrom the latest release, pick an existing ComfyUI folder or install a new one, then click Install & launch - macOS 14+ (Apple silicon): open
FreeVideo-Mac-arm64.dmg. It isn't notarized yet, so check the SHA-256 and approve it once in Privacy & Security. - Existing ComfyUI: clone it into
custom_nodes, restart, then open Workflow → Browse Templates → FreeVideo → FreeVideo-All-in-One
On Linux, it's a terminal workflow:
git clone https://github.com/FlashML-org/FreeVideo.git && cd FreeVideo
./setup.sh
# preview the memory plan for an 8 GB card without loading weights
./freevideo plan --vram-gib 8 --ram-gib 16
# render from a prompt file
./freevideo generate --prompt-file prompt.txt --out video.mp4A Prompt to Start With
H3 loves a clear subject, a camera move, and a sound cue. Drop this into prompt.txt or the creation panel:
A cozy ramen stall on a rainy neon side street at night, steam curling from the counter as the cook ladles broth into a bowl, slow push-in from across the wet street, warm lantern light against teal neon reflections, soft rain patter and sizzling broth, cinematic handheld realism.Good to Know
- Even at the 8 GB floor, give it SSD headroom. Blocks that don't fit in RAM stream from disk on every step.
- Use Preview first to test a batch of seeds cheaply, then finish only the keepers.
- Want the hosted, full-quality version? H3 runs in the cloud on Gen without any local setup.
Original Source
Co-author Haocheng Xi's launch post:
- Repo: FlashML-org/FreeVideo
- v0.3.0 notes: int8 model + prompt enhancement
- v0.3.2 notes: Preview first
- Planner docs: Adaptive Execution Planner
- Base model: OpenVDN/vdn-minimax-h3
Conclusion
FreeVideo turns "you need a datacenter" into "you need some patience and an SSD." That's a lovely shift for anyone learning H3 on the GPU they already own. Start with a preview pass, fall in love with a seed, and let your little card finish the job.
Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=freevideo-minimax-h3-8gb-vram-local-video
—Aurelia ♡
