DIAMOND Atari + Doom World Models on HF

Introduction
On 2026-10-01–02, independent maker jgalego dropped four open DIAMOND-style diffusion world models on Hugging Face: Breakout, Doom / Freedoom E1M1, Ms. Pac-Man, and Space Invaders. Each card ships Apache-2.0 weights, a single play/train script, rollout GIFs, and a one-liner uv + hf download launch — no game engine at play time, only the model.
This is the small-VRAM cousin of DIAMOND / GameNGen: last 4 frames + actions in, next frame out via a tiny U-Net with 3 Euler denoise steps. Makers can steer Atari cabinets or a Freedoom corridor from a laptop GPU today.
What shipped
ModelCreated (UTC)ParamsFrameWeights
Breakout
2026-10-01 ~18:41
~4.2M
64×64
~16 MB
Doom (Freedoom E1M1)
2026-10-01 ~23:14
7.5M
128×96
~30 MB
Ms. Pac-Man
2026-10-02 ~15:27
~4.2M
64×64
~16 MB
Space Invaders
2026-10-02 ~15:27
~4.2M
64×64
~16 MB
Architecture (shared recipe): EDM-preconditioned U-Net; noise level and the last four actions modulate every block; σ from 5.0 with three Euler steps per frame. Training data is random-policy rollouts (200k frames each). Pac-Man and Breakout trained 50k steps on an A10G (56 min); Doom is a shorter 5k-step run (~16.5 min) with EMA weights.
Play Ms. Pac-Man (browser UI, one generated frame per button):
uv run "$(hf download jgalego/ms-pacman-world-model atari.py --quiet)" play --game MsPacmanPlay Doom (each press streams four frames ≈ 0.5 s of game time):
uv run "$(hf download jgalego/doom-world-model doom.py --quiet)" playPython rollout (Pac-Man) after atari.py is on the path:
import torch
from atari import DEVICE, collect, load, rollout, to_float
model = load("jgalego/ms-pacman-world-model")
frames, _, _ = collect(4, seed=0, game="MsPacman")
context = to_float(frames[:4])[None].to(DEVICE)
moves = torch.tensor([[0, 0, 0] + [1] * 16], device=DEVICE) # UP
video = rollout(model, context, moves) # (1, 16, 3, 64, 64)Doom actions: 0 NOOP, 1 FORWARD, 2 BACKWARD, 3 LEFT, 4 RIGHT, 5 ATTACK, 6 USE. Pac-Man: 9-way joystick including diagonals.
Why game makers should care
Clock-hot, fully open, and runnable today — the same bar as the Tetris MaskGIT browser drop, but for classic discrete games and a first-person Doom slice. You get:
- Weights + loader in one HF repo per game (no “checkpoints coming soon”).
- A honest DIY path to reproduce DIAMOND-scale next-frame control without a 100M+ stack.
- Concrete held-out PSNR tables on the cards. Ms. Pac-Man (changed pixels): 21.2 / 17.25 / 15.2 dB at 1 / 5 / 15 frames ahead vs repeat baselines ~11 / 10 / 10 dB.
- Doom is the ambitious GameNGen-shaped sibling (walk / turn / shoot / USE on Freedoom E1M1) at only 7.5M params.
If you are prototyping action-conditioned world models for arcade loops, discrete planners, or tiny agent imagination, these cards are paste-ready baselines you can A/B in an afternoon.
Limits
jgalego documents the failure modes clearly — read them before you ship a demo. Doom rollouts collapse into haze/walls by ~5 frames (PSNR still beats frame-repeat early, then loses); longer training made that worse. Atari models know early-game random-policy coverage best; 64×64 sprites can smear; errors compound because every generated frame becomes the next context. Four-frame memory is under half a second of history. This is a maker research dump, not a production NetHack-scale world model.
Original Source
- Ms. Pac-Man: https://huggingface.co/jgalego/ms-pacman-world-model
- Doom: https://huggingface.co/jgalego/doom-world-model
- Breakout: https://huggingface.co/jgalego/breakout-world-model
- Space Invaders: https://huggingface.co/jgalego/space-invaders-world-model
- DIAMOND paper: https://arxiv.org/abs/2405.12399
- GameNGen paper: https://arxiv.org/abs/2408.14837
Conclusion
Four clock-hot open DIAMOND-style world models — Breakout, Space Invaders, Ms. Pac-Man, and Freedoom Doom — landed on HF Oct 1–2 with weights, play scripts, and one-liner launches. Start with Pac-Man for the cleanest metrics; try Doom if you want the first-person GameNGen sketch at 7.5M params.
—Titus
