H3 Speedups on One RTX 5090

One Board, One 5090, Every H3 Trick
This morning’s Video DeltaNet scoop was the research-cluster flex — hybrid attention racing MiniMax H3 past playback on 8×B200. Creators on a single desk GPU just got a different gift: a public bake-off that measures every known MiniMax H3 speed-up on one RTX 5090, same prompt, same seed, with the real clip sitting next to each row.
The live board is plox-1.github.io/h3-speedups (repo plox-1/h3-speedups, generated 2026-09-19). Baseline: 1344×768, 124 frames (~5.2s), 6-step turbo LoRA path, INT8 ConvRot weights, warm runs only. Speed-up is baseline wall time ÷ method wall time; differences under ~5% are noise. Still-frame “pips” are an AI rating of one fixed test shot — the site is honest that motion, lip sync, and audio still need your eyes and ears.
What Actually Wins (and What Just Looks Fast)
Fastest ≠ best. The board’s top wall-clock rows (about 1.7×, ~33s vs ~56s baseline) lean on aggressive step-cutting LoRAs and caches — and they often drop the little things that sell H3, like rain droplets on a face.
The quality-per-speed standout called out on the board is MotionCache, tuned: roughly 1.36× (~41s), still-frame pips maxed, reusing 2 of 6 steps while keeping fine detail. Close companions that also hold still-frame quality:
- Spectrum (single-pass / feature forecasting) — ~1.26×, forecasts 2 of 6 steps with no visible still loss
- Sparse attention (
sol_attn) — ~1.24×, skips most attention blocks with one dense step; near-lossless in stills - SageAttention 2.2 — looks lossless… and sits on the noise floor (~1.00×) for this short clip setup
Meanwhile, a few “famous” names land as dead ends on this exact 5-second desk test:
- VDN-H3 (ComfyUI port) — strong stills, but over 2× slower here (~111s); the board notes short clips are its worst case and suggests retesting at ~15s
- HyperFlow 8-step LoRA — keys wouldn’t load; the timed run was effectively raw 8-step H3 and slower than baseline
- ComfyUI
--fastflags — slightly slower on this INT8 path
How to Use the Board Like a Creator
- Open the live board and sort by Best balance before you chase raw speed-up.
- Expand a row, watch the clip, and drag the embedded ComfyUI workflow into your graph — every sample ships with its graph.
- Treat still-frame pips as a filter, not a verdict. If your work lives in dialogue, rain, hair, or product text, the clip is the boss.
- Stack carefully: sparse attention + a gentle cache may travel better than slamming a 3-step LoRA and praying the weather comes back.
Why This Matters for Gen Creators
H3 has been drowning in folklore — Sol-Attn recipes, Turbo LoRAs, block caches, VDN ports, “just enable Sage.” This board turns folklore into a same-machine, same-seed scorecard. Pair it with this morning’s VDN research read: cluster math and desk math are different sports. When you want the H3 look without babysitting every node, ArtRealmAI Gen keeps the family close.
Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=h3-speedups-rtx-5090-bakeoff
Original Source
https://plox-1.github.io/h3-speedups/
https://github.com/plox-1/h3-speedups
Conclusion
Speed without the raindrops is just a stopwatch brag. Pick the lane that keeps the weather — then go make something wet, golden, and on time.
—Aurelia ♡
