Introduction

WorldPlay2 just went from paper to something you can actually play. On 2026-10-06 the team posted model weights and inference code at WorldPlay2/WorldPlay2 on GitHub, and on 2026-10-07 a hosted version went live in the browser on Reactor. It's an interactive video world model: give it one image and a short structured prompt, and it streams a world you can walk through in first or third person with WASD, look around with the arrow keys, jump with space, and change what's happening mid-run with new text events.

It's the sequel to WorldPlay, the real-time world model Tencent Hunyuan shipped as HY-World 1.5. The new repo lives under its own WorldPlay2 GitHub org, and Reactor credits "the WorldPlay2 authors" (Haiyu Zhang, Wenqiang Sun, Tengfei Wang, Junta Wu, Jun Zhang, Yunhong Wang, Yu Qiao and Chunchao Guo, per the arXiv paper).

What shipped

CheckpointTypeDefault stepsDownload

WorldPlay2-Fast

Autoregressive, few-step

4

aejion/WorldPlay2-Fast

WorldPlay2-BI

Bidirectional (the teacher)

40

aejion/WorldPlay2-BI

WorldPlay2-AR

Autoregressive

40

aejion/WorldPlay2-AR

WorldPlay2-TAE

Causal VAE (lightweight decoder)

n/a

Coming soon

Modality: image-plus-text to interactive, action-controlled video. Open vs API: the weights are open but licensed CC BY-NC 4.0, so noncommercial use only. Every checkpoint sits on top of Wan2.2 I2V-A14B, which you download separately. Where to run it today: in the browser at reactor.inc/worldplay2 (you may land in a queue), or on your own GPUs with the repo.

How it stays playable

The paper's pitch is real-time response plus long-horizon consistency, and it attacks that with three ideas:

  • Factorized hybrid control. Frame-aligned actions (camera pitch and yaw, forward, back and strafe movement, perspective and a jump action) are kept separate from a structured caption that splits the scene, the character and the event. That split is why you can swap "rain begins to fall" for "sunlight breaks through" without the character or the trail changing.
  • Compressed memory. Past frames get squeezed into compact memory tokens instead of being kept at full resolution, which keeps long rollouts cheap and helps the world look the same when you turn back around.
  • Stable Forcing. A distillation recipe that starts the autoregressive student from a few-step setup and replays full rollouts so quality doesn't fall apart over long sessions.

Speed numbers from the team: the paper reports 16 FPS on 8 NVIDIA H20 GPUs, and the Reactor integration streams at a steady 17 FPS on 4 B200s, loading about 82 GB of weights.

Run it yourself

The offline scripts render controlled clips from a JSON file. Grab the base model and the fast checkpoint:

pip install "huggingface_hub[cli]>=0.34,<1.0"
hf download Wan-AI/Wan2.2-I2V-A14B --local-dir ./checkpoints/Wan2.2-I2V-A14B
hf download aejion/WorldPlay2-Fast --local-dir ./checkpoints/WorldPlay2-Fast

Then describe the run. Actions are key-duration pairs measured in latent frames, and the action and prompt durations have to add up to the same total:

[
  {
    "image_path": "input.jpg",
    "perspective": "tps",
    "action": "w-32,s-32",
    "prompt_event": "prompt1-32,prompt2-32",
    "prompt1": "The scene in the video is: a quiet forest trail. The character is: a hiker wearing a blue jacket. The event is: dark clouds gather and rain begins to fall.",
    "prompt2": "The scene in the video is: a quiet forest trail. The character is: a hiker wearing a blue jacket. The event is: the rain stops and sunlight breaks through the clouds.",
    "output_name": "sample_000"
  }
]

Point CKPT_DIR, LOW_NOISE_CKPT, HIGH_NOISE_CKPT and INPUT_JSON at your files and run bash run_few_step.sh. The launchers default to 8 GPUs with sequence parallelism; for one GPU set NPROC_PER_NODE=1 USE_DIT_FSDP=false USE_T5_FSDP=false. For first-person runs, set "perspective": "fps" and drop the "The character is" sentence.

For the live, keyboard-driven version locally, the repo's reactor/ folder builds a Docker service (reactor build, then reactor run --gpus all --port 8080) plus a Node.js frontend at localhost:3000.

Prompt to try on Reactor

Start from a strong single image (a game screenshot, concept painting or photo) and keep the event short so it reads on screen:

The scene in the video is: a rain-slick neon alley in a night market, steam rising from food stalls.
The character is: a courier in a yellow raincoat riding a small electric scooter.
The event is: lanterns flicker on one by one down the alley.

Then drive forward, turn back after a few seconds to check the stalls are still where you left them, and swap the event to "a sudden downpour clears the street."

Caveats

This is a big-iron model. Both published speed numbers need four to eight datacenter GPUs, and the 14B Wan2.2 base plus checkpoints won't fit a single consumer card for real-time play. The light TAE decoder that helps with speed isn't released yet. The license is noncommercial, so you can prototype and study with it but not ship it in a product. The hosted Reactor demo is a research preview with a queue, not an API with published pricing.

Original Source

Conclusion

WorldPlay2 is one of the most complete open drops yet for promptable, playable video worlds: three checkpoints, a JSON control format, a live browser demo and a self-host path, all landing within two days. You can't ship it commercially and you won't run it on a gaming PC, but for prototyping explorable scenes from a single concept image, it's worth a session on Reactor today.

—Titus