OmniVBench: Omni R2V Where MiniMax H3 Leads

What “omni” R2V actually means
Reference-to-video used to feel like a handful of specialty tricks: keep this character, copy that object, maybe glue two subjects together. Omni R2V is the bigger creative ask — treat references as a toolbox of factors you can mix: who appears, how they move, what style the shot wears, what structure the camera follows, and even how the story continues.
Tencent’s Online Video BU just put a shared floor under that ambition with OmniVBench (the eval) and the Omni-R2V dataset (the training fuel). If you care about character consistency, style transfer, storyboard → video, multi-ref composition, or picking which R2V stack to trust, this drop is worth a bookmark.
What’s new vs older R2V benches
Prior benches (OpenS2V-Eval, VACE-Bench, UniVBench, IntelligentVBench, FashionVideoBench, and friends) mostly lived in the content-reference lane — identity and multi-subject glue — with thinner coverage of motion, style, structure, narrative, and cross-aspect mixes.
OmniVBench stretches the map to 7 task families / 18 fine-grained tasks:
- Content — object, character, scene
- Motion — action + camera motion (from video refs)
- Style — style image → new content
- Structure — greybox, line art, rough storyboard
- Narrative — multi-panel storyboard, story rewrite, preceding-shot continuation
- Multi-content and cross-aspect — several refs, or complementary factors (e.g. content + motion)
That’s closer to how creators actually work: “keep her face, borrow his dance, paint it like this, and follow the boards.”
Factor-grounded scoring (not just “looks consistent”)
The quiet hero is the eval protocol. Instead of one holistic “did the reference stick?” score, OmniVBench ships factor-grounded checklists — 12,172 case-specific items across 813 cases — asking whether the intended factors were:
- Preserved (Reference Fidelity)
- Disentangled and routed to the right target (Instruction Realization)
- Backed by solid video quality
So a model can look “pretty consistent” and still fail for copying the wrong carrier (stealing identity when you only wanted motion), or for missing a required edit. That’s the difference between a vibe check and a creative-tool scorecard.
Omni-R2V: ~340K processed training samples
Training data for omni R2V has been fragmented and expensive. The team releases Omni-R2V at about 340K / 339,570 processed samples spanning the same seven families — image and video references, multi-ref compositions, and processed pairs ready to train on (not “download raw footage and DIY the pairing”).
Built mainly from professional video footage with task-specific pair-construction pipelines (cross-pair matching + inverse construction), it’s the kind of industrial-grade recipe open labs usually don’t get to peek at.
Datasets on Hugging Face:
Who’s leading the board right now
From the paper’s overall OmniVBench table (equal average of the three L1 dimensions):
LaneModelOverall (approx.)
Open
MiniMax H3
~72.41
Closed
Seedance 2.5
~72.68
Also scored: Kling 3.0 Omni, Happy Horse 1.0, Gemini Omni, Seedance 2.0, Vidu-Q2-Pro, Bernini, LoomVideo, OmniWeaving, UniVideo, and more.
Two practical takeaways jump out:
- The open/closed gap on omni R2V is thin at the top — H3 sitting next to Seedance 2.5 is a big deal for creators watching open stacks.
- Content refs are the easy lane. Motion, style, structure, narrative, and multi-ref/cross-aspect still show wider spreads — which matches what we feel in the wild when an identity lock works but the dance, camera, or board-follow falls apart.
What creators should do with this
- Choosing a model: Don’t crown a winner from character-ID demos alone. Peek at OmniVBench’s harder families (motion / structure / narrative / multi-ref) before you lock a pipeline.
- Training or fine-tuning: Omni-R2V gives a ready processed corpus across heterogeneous refs — useful if you’re building LoRAs, adapters, or full R2V systems that need more than “subject still → video.”
- Prompting & workflows: Think in factors. Explicitly separate what to keep, what to borrow, and what to ignore — the bench is basically punishing muddy routing.
Original Source
https://arxiv.org/abs/2609.22069
Also on HF Daily Papers: OmniVBench (2609.22069)
Conclusion
Omni R2V is where video gen is headed — not one reference trick, but a controllable mix of content, motion, style, structure, and story. OmniVBench + Omni-R2V give the community a shared ruler and a training kitchen, and the early scoreboard says MiniMax H3 is already punching in the same weight class as top closed systems. That’s the kind of race night I like to cheer from the pier.
—Aurelia ♡
