Introduction

What if the model that plans your shot is the same one that paints it, animates it, and then tells you what changed? That's the bet behind Rho-1, a research preview Reka published on October 5. It's a 19-billion-parameter "omni" model trained from scratch that understands and generates text, images, and video inside a single network, with no hand-offs to separate image or video specialists.

Quick status check before you get excited: this is a closed research preview, not open weights and not a public app. Reka is inviting builders working on interactive simulation, robotics, and vision-action systems to reach out via contact@reka.ai. So think of this as a preview of where video generation is heading rather than something you can run tonight.

What Rho-1 actually does

Reka's launch demo is a single, unedited five-turn chat. Rho-1 draws a red-and-white lighthouse on a rocky headland, puts a bounding box around the lighthouse, animates a drone flying toward it, turns the clip into a heavy snowstorm while keeping the camera move identical, and then explains what changed between the two videos.

The trick is that every turn reads from and writes to the same memory (KV cache). Because Rho-1 generated the lighthouse itself, the first frame of the video isn't a re-encoded image. It's the original representation still sitting in context, so lighting, geometry, and object identity carry through. There's no bolt-on prompt enhancer either: the same model writes the shot plan and renders it.

Speed numbers worth noticing

  • Base model: generates video at a median 0.79× real time, with a watchable stream starting in roughly 6 seconds.
  • Distilled "Flash" variant: cuts denoising from 99 steps to 8. Reka says it returned a 5.3-second clip in about one second, the fastest of any video model they timed.
  • Prompt-tweak editing: with a fixed seed, swapping a golden retriever for a black labrador in a snowy pine forest re-rendered in about 1.1 seconds.
  • Training budget: 320 H100 GPUs for three months, which Reka points out is a small fraction of frontier video-model compute.

Not just clips: a steerable world

Rho-1 can also stream continuously, clip after clip, while you send new instructions mid-roll. Reka's demos include a robot arm that grips and releases a red ball on command, and a drone rollout that splits into "bank left" and "bank right" futures from the same opening half-second. The gallery spans about 60 single-pass clips across driving, video-game worlds (voxel, low-poly, haunted mansions, torch-lit dungeons), aerial, weather, nature, and interiors.

Under the hood, each transformer block holds two expert streams. One handles understanding, the other denoises images and video, and both share attention. Text is trained with next-token prediction, while images, video, and robot actions are trained with flow matching.

Honest limitations

Reka is refreshingly upfront here:

  • Resolution: native video is currently capped at 672×384.
  • Long-horizon drift: a 30-second stream can keep crisp textures while the room layout quietly stops making sense.
  • Grounding over time: boxes and coordinates work on still images, but not yet across video.
  • Editing stability: targeted edits are still brittle across varied prompts.

Why creators should care

Today's AI video stacks chain an LLM, an image model, and a video model, and each hand-off loses context. Rho-1 is an early, low-resolution proof that one network can hold the whole story: you sketch a scene, animate it, re-weather it, and interrogate it in one conversation. If that scales the way Reka hopes, "edit my shot" could start to feel like talking to a director who never forgets the set.

Original Source

https://reka.ai/news/rho-1-collapsing-the-multimodal-stack

Conclusion

Rho-1 isn't something you can download or prompt yet, but it's a vivid signal: generation, editing, and understanding are collapsing into single, fast, conversational models. Keep an eye on Reka for wider access, and keep dreaming in drafts and snowstorms. ✨

—Aurelia ♡