WHITEBOX → GENERATED WORLDS

Proxy2World

Learning to Generate Worlds
from Proxies Without Seeing Them.

Build only what matters. Generate the rest.

Hongli Xu, Weilong Yan, Anbang Wang, Chunyu Zou, Siyu Hong, Jingwei Huang†
Tencent IEG† Corresponding author

24 proxies. Worlds waiting to unfold.
Watch the introduction · 3:14 · English narration and captions

Why Scene Proxies?

Control the terrain and the subject—without building a finished world. The whitebox is the middle ground.

Less scene specificationMore asset authoring
STRUCTURE, NOT A FINISHED WORLD

A whitebox is a constraint.
Not a unique answer.

Proxy definitions are underdetermined: different abstractions can express the same intent. A chosen whitebox fixes key spatial relationships while leaving geometry detail, materials and style open.

Keep: layout · occupancy · subjectVary: detail · materials · appearance
Shared editable castle whitebox
One chosen whitebox
Same whitebox realized as Gothic dusk
Gothic dusk
Same whitebox realized as Arcane violet
Arcane violet
Same whitebox realized as Crimson sanctum
Crimson sanctum
Camera gridGenerated landscape
0:00

Camera-grid control specifies viewpoint motion while leaving scene content open to generation.

Different examples illustrate the three interfaces; this is not a same-scene ablation.

Learning Without Proxy Pairs

No paired whitebox–real video supervision.
Learn structure and appearance from ordinary RGBD videos.

TRAIN

Ordinary RGBD videos

Learn geometry-conditioned RGB generation and joint RGBD denoising. No authored-proxy / real-video pairs are required.

TRANSFER

Author a whitebox

Specify layout, terrain, subject and camera using a simple editable proxy—not finished art assets.

GENERATE

Ground, then refine

Use proxy geometry in early steps, then release the explicit depth condition and jointly refine RGBD while retaining camera guidance.

Proxy2World: cross-modal joint RGBD learning with a shared DiT, followed by proxy-camera hybrid denoising

A scene proxy is a designer’s abstraction, not a uniquely recoverable target from video. We therefore learn geometry-conditioned RGB generation and joint RGBD generation from ordinary RGBD observations, without scene-proxy training data. At inference, proxy geometry grounds the early denoising steps; later steps remove that explicit condition while retaining camera guidance, allowing geometry and appearance to be completed together. Camera–geometry scale alignment keeps both controls consistent.

From Proxies to Worlds

Drag between visual styles. Switch the interaction to change what happens in the world.

Same scene · same recorded motion
Gothic duskCrimson sanctum
Shared scene proxy

Lower the drawbridge, open the gate, and cross into the hall.

Recorded interaction · synchronized 15-second outputs
0:00

Existing generated videos, not live inference. Select “Scene proxy” on the left to compare structure with appearance.

Create Your Own World

“Build me a demon castle I can explore and unravel.”

A lightweight proxy becomes a shared creation interface for the player, the AI level designer, and the generative world model.

An idea → A level plan → An editable proxy → A visual world

Design-process replay and recorded generation results.

Demon Castle art-direction reference, not gameplay
Art-direction reference
Actual editable Demon Castle whiteboxEDITABLE SCENE PROXY

ProxyBench

A benchmark for spatial control, visual quality, and interactive world generation.

We introduce ProxyBench, an agent-driven benchmark spanning 25 diverse scenes and 300+ video clips, including interactive scenarios. By systematically varying proxy abstraction, camera trajectory complexity, and motion scale, ProxyBench evaluates how well world models follow spatial structure and scene interactions while generating detailed, visually compelling worlds.

ProxyBench design and evaluation overview

Qualitative Comparisons

Layout preview · repeated demo footage

One scene per row. Explore the proxy and camera path, then slide through methods and drag to compare with ours.

Temporary mock-up: each method repeats the same existing video. The 3D proxies and paths are schematic, not recovered geometry or measured camera trajectories. No baseline performance is represented.

Quantitative Results

Results pending in supplied paper
VLM scores: 1–5, higher is better. Camera errors: lower is better. Dashes denote unreported results, not zero.
MethodVLM-based evaluationCamera accuracy
Structure ↑Text ↑Quality ↑Artifact suppression ↑Rotation error ↓Translation error ↓
Camera-conditioned methods
Gen3R——————
SCoPE——————
FantasyWorld——————
Ours (T2V, P = 0)——————
Ours (TI2V, P = 0)——————
Depth-conditioned methods
VACE-Depth——————
Cosmos-Depth——————
Ours (T2V, P = 1)——————
Ours (TI2V, P = 1)——————
Proxy-conditioned methods
Coarse2Real——————
Ours (T2V, P = P★)——————
Ours (TI2V, P = P★)——————

Exploration & Gameplay

First-person exploration and third-person gameplay. Select a clip to view the original video.

FROM WORD TO WORLD
Open full page ↗