Spark-H3 Speeds Local MiniMax H3 Attention

Introduction
Local MiniMax H3 is gorgeous and hungry. Spark-H3 is the open attention backend that tries to keep the beauty while cutting the wait: adaptive block-sparse attention tuned for H3, with measured speedups and a ComfyUI node that landed in the public repo on October 8, 2026.
It is not a new video model and it does not ship H3 weights. Think of it as a smarter sparse engine you plug under the H3 you already run — including few-step LoRAs like LightX2V.
What shipped
- Name: Spark-H3 (MiniMax-H3 build of Spark-Attn)
- Team: SparkH3 / zechengtang
- Modality: video generation acceleration (attention backend for MiniMax H3)
- License / access: open code, Apache-2.0 on the Hugging Face card; does not redistribute MiniMax H3 weights
- Where to run it today: GitHub zechengtang/Spark-H3 (Diffusers + standalone ComfyUI Spark node), model card at Aazeus/Spark-H3, tech writeup at zechengtang.github.io/Spark-H3
- Concrete fact: on a 10-second 768p benchmark at 10% attention density, Spark-H3 reports up to 2.33× attention speedup and 1.73× DiT speedup, with 23.30 dB PSNR vs 20.36 dB for Sol-H3 under the same protocol
What Spark-Attn actually does
Block-sparse attention usually splits work into an exact branch (score the important blocks carefully) and a compressed branch (cheaply summarize the rest). Spark-Attn upgrades both sides:
- Spark-Reblock groups tokens that share similar attention preferences, so the exact budget lands on more useful interactions instead of fixed spatio-temporal tiles.
- Spark-Reweight builds weighted key/value summaries plus a log-mass correction, fighting the bias that plain mean-pooling introduces on the compressed branch.
The project page was first published September 23; the October 8 refresh is what makers care about today: README/demo polish, Diffusers LBH workflow, and a standalone ComfyUI Spark node commit.
How to try it
- Install MiniMax H3 weights from the official MiniMaxAI card (Spark-H3 does not include them).
- Clone the Spark-H3 repo and follow the ComfyUI or Diffusers README in-tree.
- Start around 10% density with the recommended dense warmup (first Transformer layer kept dense; early steps warm).
- Pair it with a few-step LoRA you already trust — LightX2V, Larryvrh, Alibaba-PAI Acc, ByteDance DMAD, and FastH3 are all shown in the project writeup.
Production fused kernels target SM120 GPUs (RTX 50-series and RTX PRO 5000/6000 Blackwell). Latency numbers on the page were measured on a single RTX PRO 6000 Blackwell.
Why creators should care
Few-step LoRAs cut the number of denoising steps. Sparse attention cuts the cost of each step. Spark-H3 is built to stack with that ecosystem rather than replace it — so a LightX2V 8-step stack can keep moving while the attention path stays closer to dense fidelity than older Sol-Attn baselines in their tables.
If your H3 box is already sweating through 10-second clips, this is the kind of plumbing scoop that changes how many drafts you finish before dinner.
Limits
- Hardware and kernel support matter; treat the SM120 notes seriously before you expect fused-kernel speedups on older cards.
- Sparse schedules trade a little fidelity for speed — use the Balanced / Fidelity density presets when the shot needs to match dense more closely.
- You still need a legal MiniMax H3 install under the H3 Community License (territorial grant and commercial rules apply).
Original Source
Project page (updated Oct 8, 2026): Spark-H3
Hugging Face card: Aazeus/Spark-H3 (created 2026-10-08T08:35:16Z)
GitHub (ComfyUI node commit 2026-10-08T14:03:47Z): zechengtang/Spark-H3
Conclusion
Spark-H3 is the local H3 crowd's newest accelerator: open sparse attention, honest benchmarks against Sol-H3, and a ComfyUI node you can wire into the few-step LoRAs you already love. Pull the repo, keep your H3 weights legal, and see how many more takes fit in the same GPU hour.
Try this on Gen → https://artrealmai.com/gen?utm_source=magazine&utm_campaign=gen&utm_content=spark-h3-comfyui-block-sparse-attention
—Aurelia ♡
