H3 Video Upsampler: Guided 2x Upscale in ComfyUI

Introduction
Here's a sweet little finishing tool for anyone rendering MiniMax H3 clips at a modest size to spare their VRAM. H3 Video Upsampler is a new open-source ComfyUI custom node from indie developer dntpi (published on Hugging Face as sandpies). It takes a finished H3 clip, or really any clip, and enlarges it with the H3 latent upscaler. Then it runs one light re-draw step in tiles to bring the detail back. The mouths, motion and soundtrack all come from your original clip.
Version 0.1.0 landed on GitHub and Hugging Face on October 6 (13:57 UTC, about 21:57 MYT). It ships with a purpose-trained h3upscale x2 LoRA, so the re-draw can actually look at your clip while it works. It runs locally and it's free.
What Shipped
- The node: H3 Video Upsampler (found under video/upscale in ComfyUI). It returns upscaled frames plus your audio.
- The guide LoRA: h3upscale_v2_x2.safetensors, a rank-64 LoRA trained for 4,000 steps with ai-toolkit on Alissonerdx's h3upscale-4k dataset, using the MiniMax H3 Ref2VA base.
- The trick: with that LoRA loaded, the re-draw step also gets your clip at half the output size as a guide. It restores the clip's own detail instead of inventing new texture.
- The latent upscaler: LBH-123-AI's minimax_h3_latent_upscaler_3d_bf16.safetensors (Apache-2.0), the same network behind the 10Eros upscale workflow we covered in September.
- A ready graph: example_workflows/H3_Video_Upsampler.json. Load it, pick a clip, queue.
- Any size, any length: frames are fitted to 32 px and padded to a valid H3 length, then the padding is trimmed off again.
The Knobs That Matter
- megapixels: output size with the aspect ratio kept. 2 is about 1920×1080.
- denoise: how hard the single re-draw step works. 0.1 is the tested setting, and 0 only enlarges.
- tile_tokens: your VRAM dial. The default is 70000. Try 40000 or 25000 on smaller cards (more, smaller tiles, slower, less memory).
- lora: pick h3upscale_v2_x2 and guided mode switches on automatically, detected from the file header.
- prompt: defaults to the caption the LoRA was trained on. Keep it when you use the LoRA.
- audio_vae: wire it in and the re-draw "hears" your soundtrack (best at 24 fps). Leave it unwired and it hears silence. Either way you get your original audio back.
One tip from the README: feed it the plain H3 diffusion model loader. The LoRA was trained on the bare DiT, and any extra LoRAs stacked under it add their own texture back in.
How Well Does It Work?
The developer says they tested it by eye on H3 renders. Guided at denoise 0.1, the 2 MP result beat the native render in a blind comparison. They're also upfront about its limits: this is an upscaler only. If you try generating video from noise with the LoRA, it flickers every 17 frames, because it was trained on 5-frame clips. Refining an existing clip, which is what the node does, doesn't have that problem.
Early community tests are encouraging. X creator minami_gunma rendered at 0.6 MP and upscaled to 2.2 MP, taking a 640×960 clip to 1280×1920.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/dntpi/ComfyUI-H3-Video-UpsamplerThen restart ComfyUI and place two files:
- minimax_h3_latent_upscaler_3d_bf16.safetensors goes in ComfyUI/models/latent_upscale_models (from LBH-123-AI on Hugging Face)
- h3upscale_v2_x2.safetensors goes in ComfyUI/models/loras (from the sandpies repo's lora folder)
You'll need a ComfyUI build with MiniMax H3 support and the usual H3 models: the DiT, text encoder, and video and audio VAEs.
Licensing Notes
- Node code: MIT
- Latent upscaler weights: Apache-2.0 (LBH-123-AI)
- h3upscale LoRA: marked "other" on Hugging Face. The notices ask you to read the terms of Alissonerdx's h3upscale-4k dataset and the MiniMax H3 base before you redistribute it or use it commercially.
Original Source
- GitHub (node, v0.1.0): https://github.com/dntpi/ComfyUI-H3-Video-Upsampler
- Hugging Face (LoRA and code mirror): https://huggingface.co/sandpies/ComfyUI-H3-Video-Upsampler
- Latent upscaler weights: https://huggingface.co/LBH-123-AI/Minimax_h3_latent_Upscaler
- Parent node pack: https://github.com/dntpi/ComfyUI-Hand-Tie-Clips
Community 0.6 MP to 2.2 MP test:
Conclusion
I love how humble and clever this one is. You don't need a bigger GPU or a pricier render. Make your H3 clip small and fast, keep the take you love, then let a guided re-draw sharpen it while your motion and audio stay put. Start at denoise 0.1, turn tile_tokens down if your card complains, and let me know how your 2 MP renders look.
Want a fresh H3 clip to upscale? Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=h3-video-upsampler-comfyui-guided-2x-lora
—Aurelia ♡
