Fei-Fei Li’s World Labs Unveils Atlas: The First Multimodal World Model for Pixel-Level Camera Control

9.2 Fei-Fei Li’s World Labs has officially unveiled Atlas, the world’s first multimodal world model capable of pixel-precise camera control for generating images and video frames, while simultaneously reconstructing them in 3D space.

Fei-Fei Li’s World Labs has officially unveiled Atlas, the world’s first multimodal world model capable of pixel-precise camera control for generating images and video frames, while simultaneously reconstructing them in 3D space.

360截图20260902135900

Atlas’s core breakthrough lies in treating camera pose parameters as a native input. Traditional AI video generation relies on text descriptions like “slowly pan left,” leaving outcomes to chance. Atlas accepts precise camera coordinates, allowing users to specify any position and angle for the virtual camera. With just one to six reference photos and a defined camera trajectory, Atlas can generate videos up to one minute long at 1440p resolution with consistent spatial geometry.

World Labs describes Atlas as a multimodal autoregressive diffusion Transformer, pretrained from scratch to natively process text, images, video, camera poses, and 3D depth data—all within a single model for world generation, reconstruction, and simulation.

The model features four core capabilities: Camera-Controlled Generation—produces up to one minute of 1440p video with pixel-level precision; Spatial Reconstruction—reconstructs real scenes from as few as one to twenty-five input images, achieving state-of-the-art results on datasets like DTU and ETH3D (25.3 AbsRel ×10⁻³), outperforming specialized 3D models; Space-Time Simulation—models both space and time from input video, enabling dramatic reframing and Real-to-Sim workflows for robotics—reconstructing large environments from just 24 frames of cell phone footage.

In human evaluations for camera-controlled generation, Atlas achieved preference rates of 75% over MiniMax H3, 81% over Gemini Omni Flash, 86% over Alibaba HappyHorse 1.1, 93% over FLUX 3, and 94% over ByteDance Seedance 2.5.

While other video models offer a "prompt lottery," Atlas hands creators a director's viewfinder. It doesn't just generate flashy visuals—it gives back camera control, letting AI construct worlds from precise geometric instructions. When a foundation model begins to understand space as well as text, the long-discussed concept of a "world model" finally takes tangible form.

https://www.worldlabs.ai/blog/atlas