Building Rome: One Photo to a Metric 3D Scene Mesh

Introduction
Building Rome from a Single Image is a new open release from Applied Intuition's AI Research team (with Purdue, UIUC and UC Berkeley co-authors). Give it one photo, indoor or outdoor, and it outputs a complete metric 3D scene mesh: the surfaces you can see, plus plausible geometry for what the camera can't. Inference code landed on GitHub and the fine-tuned checkpoints on Hugging Face on 10 Oct 2026. The paper is arXiv 2610.08790, and the project page has orbitable reconstructions.
- Modality: single image to 3D scene mesh (geometry, exported as
scene.glb) - Open vs closed: open code and weights under CC BY-NC 4.0, so non-commercial use only. It sits on top of Microsoft's MIT-licensed TRELLIS.2-4B base weights.
- Where to run today: locally on Linux with an NVIDIA CUDA GPU. There's no hosted demo yet.
- One hard number: a street scene reaching 98 m (4 depth bands, 15 chunks) peaks at 15.8 GiB of GPU memory and takes about 6 minutes on one A100 80GB.
What shipped
TRELLIS.2 is built to generate single objects. Building Rome teaches it to generate whole scenes:
- MoGe-3 estimates metric depth and the camera from your photo.
- DINOv3 image features get lifted onto those observed surfaces.
- The scene is tiled with chunks that grow with distance: native 3 m chunks up to 9 m out, then bigger bands further away.
- A fine-tuned TRELLIS.2 generates each chunk using latent outpainting, reusing overlap latents across bands, and the chunks are stitched into one mesh.
The release is two small delta files that get added to the public TRELLIS.2 weights:
CheckpointContentsSize
ss_delta.safetensors
Sparse-structure (occupancy) LoRA r32 adapters plus depth-lift and clearance branches, 115M params
438 MiB
slat_delta.safetensors
Shape-model LoRA r32 adapters plus depth-lift branch, 66M params
251 MiB
Training used indoor scenes from NVIDIA's SAGE-10k, Infinigen scenes, and about 4,000 synthetic outdoor scenes.
Run it
git clone https://github.com/Applied-Intuition-Open-Source/BuildRome.git
cd BuildRome
. ./setup.sh --new-env --basic --flash-attn --nvdiffrast --cumesh --flexgemm --o-voxel
hf download AppliedIntuitionResearch/BuildRome --local-dir checkpoints
python inference/image_to_scene.py --image path/to/photo.jpg \
--ss-ckpt checkpoints/ss_delta.safetensors --slat-ckpt checkpoints/slat_delta.safetensors --out results/Or load it once and batch images from Python:
from build_rome import ScenePipeline
pipe = ScenePipeline.from_pretrained("AppliedIntuitionResearch/BuildRome")
scene = pipe("photo.jpg")
scene.save("results/photo") # scene.glb, layout_bands.json, layout_topview.png, spiral.mp4Some setup notes before you start:
- MoGe-3 needs PyTorch 2.8+, so the README has you install it in a separate conda env and pass
--moge3-python /path/to/envs/moge3/bin/python. You can also feed in a precomputed--depth-packinstead. - DINOv3 is gated on Hugging Face, so accept its license and log in before the first run.
- Out of GPU memory? Lower
--max-inflated-voxels(default 100000). Add--no-spiral-videoto skip the orbit render.
Why it matters for game builders
The output scene.glb is metric, glTF Y-up, with the camera at the origin looking down -Z, which drops straight into Blender, Godot, Unity or Unreal with real-world scale. That makes it a fast way to block out a level from a location photo or a concept painting: a street, a room, or a plaza with its hidden sides filled in.
Keep the limits in mind. You get geometry only, with no texture. Hidden regions are plausible guesses, not reconstructions. Meshes can have holes and aren't guaranteed watertight. Any depth error from MoGe-3 carries into the scene. And because of CC BY-NC, this is for prototypes and research, not shipped commercial assets.
Conclusion
Building Rome turns TRELLIS.2 from an object generator into a one-photo scene generator, with code and weights you can run today on a 16 GB+ CUDA GPU. Grab the checkpoints, point it at a photo, and open the GLB in your engine to see how much of the level it blocks out for you.
—Titus
