Ordinary RGBD videos
Learn geometry-conditioned RGB generation and joint RGBD denoising. No authored-proxy / real-video pairs are required.
Learning to Generate Worlds
from Proxies Without Seeing Them.
Build only what matters. Generate the rest.
Control the terrain and the subject—without building a finished world. The whitebox is the middle ground.
Proxy definitions are underdetermined: different abstractions can express the same intent. A chosen whitebox fixes key spatial relationships while leaving geometry detail, materials and style open.




Camera-grid control specifies viewpoint motion while leaving scene content open to generation.
Different examples illustrate the three interfaces; this is not a same-scene ablation.
No paired whitebox–real video supervision.
Learn structure and appearance from ordinary RGBD videos.
Learn geometry-conditioned RGB generation and joint RGBD denoising. No authored-proxy / real-video pairs are required.
Specify layout, terrain, subject and camera using a simple editable proxy—not finished art assets.
Use proxy geometry in early steps, then release the explicit depth condition and jointly refine RGBD while retaining camera guidance.

A scene proxy is a designer’s abstraction, not a uniquely recoverable target from video. We therefore learn geometry-conditioned RGB generation and joint RGBD generation from ordinary RGBD observations, without scene-proxy training data. At inference, proxy geometry grounds the early denoising steps; later steps remove that explicit condition while retaining camera guidance, allowing geometry and appearance to be completed together. Camera–geometry scale alignment keeps both controls consistent.
Drag between visual styles. Switch the interaction to change what happens in the world.
Lower the drawbridge, open the gate, and cross into the hall.
Recorded interaction · synchronized 15-second outputsExisting generated videos, not live inference. Select “Scene proxy” on the left to compare structure with appearance.
“Build me a demon castle I can explore and unravel.”
A lightweight proxy becomes a shared creation interface for the player, the AI level designer, and the generative world model.
Design-process replay and recorded generation results.

EDITABLE SCENE PROXYA benchmark for spatial control, visual quality, and interactive world generation.
We introduce ProxyBench, an agent-driven benchmark spanning 25 diverse scenes and 300+ video clips, including interactive scenarios. By systematically varying proxy abstraction, camera trajectory complexity, and motion scale, ProxyBench evaluates how well world models follow spatial structure and scene interactions while generating detailed, visually compelling worlds.

One scene per row. Explore the proxy and camera path, then slide through methods and drag to compare with ours.
Temporary mock-up: each method repeats the same existing video. The 3D proxies and paths are schematic, not recovered geometry or measured camera trajectories. No baseline performance is represented.
| Method | VLM-based evaluation | Camera accuracy | ||||
|---|---|---|---|---|---|---|
| Structure ↑ | Text ↑ | Quality ↑ | Artifact suppression ↑ | Rotation error ↓ | Translation error ↓ | |
| Camera-conditioned methods | ||||||
| Gen3R | — | — | — | — | — | — |
| SCoPE | — | — | — | — | — | — |
| FantasyWorld | — | — | — | — | — | — |
| Ours (T2V, P = 0) | — | — | — | — | — | — |
| Ours (TI2V, P = 0) | — | — | — | — | — | — |
| Depth-conditioned methods | ||||||
| VACE-Depth | — | — | — | — | — | — |
| Cosmos-Depth | — | — | — | — | — | — |
| Ours (T2V, P = 1) | — | — | — | — | — | — |
| Ours (TI2V, P = 1) | — | — | — | — | — | — |
| Proxy-conditioned methods | ||||||
| Coarse2Real | — | — | — | — | — | — |
| Ours (T2V, P = P★) | — | — | — | — | — | — |
| Ours (TI2V, P = P★) | — | — | — | — | — | — |
First-person exploration and third-person gameplay. Select a clip to view the original video.