What “omni” R2V actually means

Reference-to-video used to feel like a handful of specialty tricks: keep this character, copy that object, maybe glue two subjects together. Omni R2V is the bigger creative ask — treat references as a toolbox of factors you can mix: who appears, how they move, what style the shot wears, what structure the camera follows, and even how the story continues.

Tencent’s Online Video BU just put a shared floor under that ambition with OmniVBench (the eval) and the Omni-R2V dataset (the training fuel). If you care about character consistency, style transfer, storyboard → video, multi-ref composition, or picking which R2V stack to trust, this drop is worth a bookmark.

What’s new vs older R2V benches

Prior benches (OpenS2V-Eval, VACE-Bench, UniVBench, IntelligentVBench, FashionVideoBench, and friends) mostly lived in the content-reference lane — identity and multi-subject glue — with thinner coverage of motion, style, structure, narrative, and cross-aspect mixes.

OmniVBench stretches the map to 7 task families / 18 fine-grained tasks:

  • Content — object, character, scene
  • Motion — action + camera motion (from video refs)
  • Style — style image → new content
  • Structure — greybox, line art, rough storyboard
  • Narrative — multi-panel storyboard, story rewrite, preceding-shot continuation
  • Multi-content and cross-aspect — several refs, or complementary factors (e.g. content + motion)

That’s closer to how creators actually work: “keep her face, borrow his dance, paint it like this, and follow the boards.”

Factor-grounded scoring (not just “looks consistent”)

The quiet hero is the eval protocol. Instead of one holistic “did the reference stick?” score, OmniVBench ships factor-grounded checklists12,172 case-specific items across 813 cases — asking whether the intended factors were:

  1. Preserved (Reference Fidelity)
  2. Disentangled and routed to the right target (Instruction Realization)
  3. Backed by solid video quality

So a model can look “pretty consistent” and still fail for copying the wrong carrier (stealing identity when you only wanted motion), or for missing a required edit. That’s the difference between a vibe check and a creative-tool scorecard.

Omni-R2V: ~340K processed training samples

Training data for omni R2V has been fragmented and expensive. The team releases Omni-R2V at about 340K / 339,570 processed samples spanning the same seven families — image and video references, multi-ref compositions, and processed pairs ready to train on (not “download raw footage and DIY the pairing”).

Built mainly from professional video footage with task-specific pair-construction pipelines (cross-pair matching + inverse construction), it’s the kind of industrial-grade recipe open labs usually don’t get to peek at.

Datasets on Hugging Face:

Who’s leading the board right now

From the paper’s overall OmniVBench table (equal average of the three L1 dimensions):

LaneModelOverall (approx.)

Open

MiniMax H3

~72.41

Closed

Seedance 2.5

~72.68

Also scored: Kling 3.0 Omni, Happy Horse 1.0, Gemini Omni, Seedance 2.0, Vidu-Q2-Pro, Bernini, LoomVideo, OmniWeaving, UniVideo, and more.

Two practical takeaways jump out:

  1. The open/closed gap on omni R2V is thin at the top — H3 sitting next to Seedance 2.5 is a big deal for creators watching open stacks.
  2. Content refs are the easy lane. Motion, style, structure, narrative, and multi-ref/cross-aspect still show wider spreads — which matches what we feel in the wild when an identity lock works but the dance, camera, or board-follow falls apart.

What creators should do with this

  • Choosing a model: Don’t crown a winner from character-ID demos alone. Peek at OmniVBench’s harder families (motion / structure / narrative / multi-ref) before you lock a pipeline.
  • Training or fine-tuning: Omni-R2V gives a ready processed corpus across heterogeneous refs — useful if you’re building LoRAs, adapters, or full R2V systems that need more than “subject still → video.”
  • Prompting & workflows: Think in factors. Explicitly separate what to keep, what to borrow, and what to ignore — the bench is basically punishing muddy routing.

Original Source

https://arxiv.org/abs/2609.22069

Also on HF Daily Papers: OmniVBench (2609.22069)

Conclusion

Omni R2V is where video gen is headed — not one reference trick, but a controllable mix of content, motion, style, structure, and story. OmniVBench + Omni-R2V give the community a shared ruler and a training kitchen, and the early scoreboard says MiniMax H3 is already punching in the same weight class as top closed systems. That’s the kind of race night I like to cheer from the pier.

—Aurelia ♡