Introduction

Building Rome from a Single Image is a new open release from Applied Intuition's AI Research team (with Purdue, UIUC and UC Berkeley co-authors). Give it one photo, indoor or outdoor, and it outputs a complete metric 3D scene mesh: the surfaces you can see, plus plausible geometry for what the camera can't. Inference code landed on GitHub and the fine-tuned checkpoints on Hugging Face on 10 Oct 2026. The paper is arXiv 2610.08790, and the project page has orbitable reconstructions.

  • Modality: single image to 3D scene mesh (geometry, exported as scene.glb)
  • Open vs closed: open code and weights under CC BY-NC 4.0, so non-commercial use only. It sits on top of Microsoft's MIT-licensed TRELLIS.2-4B base weights.
  • Where to run today: locally on Linux with an NVIDIA CUDA GPU. There's no hosted demo yet.
  • One hard number: a street scene reaching 98 m (4 depth bands, 15 chunks) peaks at 15.8 GiB of GPU memory and takes about 6 minutes on one A100 80GB.

What shipped

TRELLIS.2 is built to generate single objects. Building Rome teaches it to generate whole scenes:

  1. MoGe-3 estimates metric depth and the camera from your photo.
  2. DINOv3 image features get lifted onto those observed surfaces.
  3. The scene is tiled with chunks that grow with distance: native 3 m chunks up to 9 m out, then bigger bands further away.
  4. A fine-tuned TRELLIS.2 generates each chunk using latent outpainting, reusing overlap latents across bands, and the chunks are stitched into one mesh.

The release is two small delta files that get added to the public TRELLIS.2 weights:

CheckpointContentsSize

ss_delta.safetensors

Sparse-structure (occupancy) LoRA r32 adapters plus depth-lift and clearance branches, 115M params

438 MiB

slat_delta.safetensors

Shape-model LoRA r32 adapters plus depth-lift branch, 66M params

251 MiB

Training used indoor scenes from NVIDIA's SAGE-10k, Infinigen scenes, and about 4,000 synthetic outdoor scenes.

Run it

git clone https://github.com/Applied-Intuition-Open-Source/BuildRome.git
cd BuildRome
. ./setup.sh --new-env --basic --flash-attn --nvdiffrast --cumesh --flexgemm --o-voxel
hf download AppliedIntuitionResearch/BuildRome --local-dir checkpoints
python inference/image_to_scene.py --image path/to/photo.jpg \
    --ss-ckpt checkpoints/ss_delta.safetensors --slat-ckpt checkpoints/slat_delta.safetensors --out results/

Or load it once and batch images from Python:

from build_rome import ScenePipeline

pipe = ScenePipeline.from_pretrained("AppliedIntuitionResearch/BuildRome")
scene = pipe("photo.jpg")
scene.save("results/photo")  # scene.glb, layout_bands.json, layout_topview.png, spiral.mp4

Some setup notes before you start:

  • MoGe-3 needs PyTorch 2.8+, so the README has you install it in a separate conda env and pass --moge3-python /path/to/envs/moge3/bin/python. You can also feed in a precomputed --depth-pack instead.
  • DINOv3 is gated on Hugging Face, so accept its license and log in before the first run.
  • Out of GPU memory? Lower --max-inflated-voxels (default 100000). Add --no-spiral-video to skip the orbit render.

Why it matters for game builders

The output scene.glb is metric, glTF Y-up, with the camera at the origin looking down -Z, which drops straight into Blender, Godot, Unity or Unreal with real-world scale. That makes it a fast way to block out a level from a location photo or a concept painting: a street, a room, or a plaza with its hidden sides filled in.

Keep the limits in mind. You get geometry only, with no texture. Hidden regions are plausible guesses, not reconstructions. Meshes can have holes and aren't guaranteed watertight. Any depth error from MoGe-3 carries into the scene. And because of CC BY-NC, this is for prototypes and research, not shipped commercial assets.

Conclusion

Building Rome turns TRELLIS.2 from an object generator into a one-photo scene generator, with code and weights you can run today on a 16 GB+ CUDA GPU. Grab the checkpoints, point it at a photo, and open the GLB in your engine to see how much of the level it blocks out for you.

—Titus