Introduction

On 2026-10-01–02, independent maker jgalego dropped four open DIAMOND-style diffusion world models on Hugging Face: Breakout, Doom / Freedoom E1M1, Ms. Pac-Man, and Space Invaders. Each card ships Apache-2.0 weights, a single play/train script, rollout GIFs, and a one-liner uv + hf download launch — no game engine at play time, only the model.

This is the small-VRAM cousin of DIAMOND / GameNGen: last 4 frames + actions in, next frame out via a tiny U-Net with 3 Euler denoise steps. Makers can steer Atari cabinets or a Freedoom corridor from a laptop GPU today.

What shipped

ModelCreated (UTC)ParamsFrameWeights

Breakout

2026-10-01 ~18:41

~4.2M

64×64

~16 MB

Doom (Freedoom E1M1)

2026-10-01 ~23:14

7.5M

128×96

~30 MB

Ms. Pac-Man

2026-10-02 ~15:27

~4.2M

64×64

~16 MB

Space Invaders

2026-10-02 ~15:27

~4.2M

64×64

~16 MB

Architecture (shared recipe): EDM-preconditioned U-Net; noise level and the last four actions modulate every block; σ from 5.0 with three Euler steps per frame. Training data is random-policy rollouts (200k frames each). Pac-Man and Breakout trained 50k steps on an A10G (56 min); Doom is a shorter 5k-step run (~16.5 min) with EMA weights.

Play Ms. Pac-Man (browser UI, one generated frame per button):

uv run "$(hf download jgalego/ms-pacman-world-model atari.py --quiet)" play --game MsPacman

Play Doom (each press streams four frames ≈ 0.5 s of game time):

uv run "$(hf download jgalego/doom-world-model doom.py --quiet)" play

Python rollout (Pac-Man) after atari.py is on the path:

import torch
from atari import DEVICE, collect, load, rollout, to_float

model = load("jgalego/ms-pacman-world-model")
frames, _, _ = collect(4, seed=0, game="MsPacman")
context = to_float(frames[:4])[None].to(DEVICE)
moves = torch.tensor([[0, 0, 0] + [1] * 16], device=DEVICE)  # UP
video = rollout(model, context, moves)  # (1, 16, 3, 64, 64)

Doom actions: 0 NOOP, 1 FORWARD, 2 BACKWARD, 3 LEFT, 4 RIGHT, 5 ATTACK, 6 USE. Pac-Man: 9-way joystick including diagonals.

Why game makers should care

Clock-hot, fully open, and runnable today — the same bar as the Tetris MaskGIT browser drop, but for classic discrete games and a first-person Doom slice. You get:

  • Weights + loader in one HF repo per game (no “checkpoints coming soon”).
  • A honest DIY path to reproduce DIAMOND-scale next-frame control without a 100M+ stack.
  • Concrete held-out PSNR tables on the cards. Ms. Pac-Man (changed pixels): 21.2 / 17.25 / 15.2 dB at 1 / 5 / 15 frames ahead vs repeat baselines ~11 / 10 / 10 dB.
  • Doom is the ambitious GameNGen-shaped sibling (walk / turn / shoot / USE on Freedoom E1M1) at only 7.5M params.

If you are prototyping action-conditioned world models for arcade loops, discrete planners, or tiny agent imagination, these cards are paste-ready baselines you can A/B in an afternoon.

Limits

jgalego documents the failure modes clearly — read them before you ship a demo. Doom rollouts collapse into haze/walls by ~5 frames (PSNR still beats frame-repeat early, then loses); longer training made that worse. Atari models know early-game random-policy coverage best; 64×64 sprites can smear; errors compound because every generated frame becomes the next context. Four-frame memory is under half a second of history. This is a maker research dump, not a production NetHack-scale world model.

Original Source

Conclusion

Four clock-hot open DIAMOND-style world models — Breakout, Space Invaders, Ms. Pac-Man, and Freedoom Doom — landed on HF Oct 1–2 with weights, play scripts, and one-liner launches. Start with Pac-Man for the cleanest metrics; try Doom if you want the first-person GameNGen sketch at 7.5M params.

—Titus