Articles

World Labs Atlas: Fei-Fei Li's Omni World Model

World Labs unveiled Atlas, an omni world model for spatial intelligence that generates 3D scenes, depth, and Gaussian splats from images or text.

Chisato Chisato · · 6 min read
An abstract swirl of generative color and light against a dark field

World Labs, the startup founded in 2024 by Stanford computer scientist Fei-Fei Li, unveiled Atlas on Monday, September 1, 2026 — a system it calls an “omni world model” for spatial intelligence. Atlas is entering early access with select partners through a request form; the company disclosed no price, no general-availability date, and no details on model size or rate limits. It is the most ambitious release yet from a company that has raised roughly $1.2 billion on the premise that the next frontier of AI is not language but the physical world.

The pitch is straightforward and expansive: rather than stitching together separate systems for image generation, depth estimation, novel-view synthesis, and 3D reconstruction, Atlas does all of it inside one model. World Labs describes it as a multimodal autoregressive diffusion transformer — a single architecture that ingests and generates across text, images, video, and 3D data within a unified spatial framework. The company positions that consolidation as the point: an omni model that reads a handful of ordinary photos and outputs a coherent, navigable scene.

What Atlas actually produces

According to World Labs, Atlas generates pixel-perfect, camera-controlled imagery and video at resolutions up to 1440p, for clips running up to about one minute. More important than the video specs is what sits underneath them. From sparse inputs — in some demonstrations, a few frames of cell-phone footage — the model can produce novel viewpoints, depth maps, point clouds, and 3D Gaussian splats, the volumetric representation that has become the default for photorealistic scene reconstruction.

That output stack is what separates a world model from a video generator. A conventional text-to-video system predicts the next frame; it does not necessarily know where objects sit in space or how a scene looks from an angle the camera never occupied. Atlas is designed to carry an internal, consistent 3D representation, so that a generated environment holds together as a viewer moves through it. World Labs frames this as the difference between a plausible-looking clip and a walkable space — the concept explored in our explainer on what a world model is.

One benchmark claim drew particular attention: World Labs says Atlas, as a single omni model, outperforms specialized single-purpose 3D models on their own tasks. If that holds up under independent testing, it would be an argument for generality over the pipeline of narrow tools that most 3D and simulation workflows rely on today. It is also the claim most in need of scrutiny — the company’s announcement did not publish a full methodology, and at least one early analysis argued Atlas has not yet proven it can genuinely simulate the world as opposed to rendering convincing views of it.

The spatial-intelligence thesis

Atlas is the clearest expression so far of the idea Li has been building toward since leaving her post as director of Stanford’s AI Lab — that intelligence rooted only in text and 2D images is fundamentally incomplete. Her argument, which she calls spatial intelligence, holds that if AI is to understand and act in the physical world, it must reason natively about geometry, perspective, and the persistence of objects across space and time. Language models describe the world; a world model is meant to represent it.

That thesis has attracted an unusually strategic investor base. World Labs’ roughly $1.2 billion in funding includes NVIDIA, AMD, and Autodesk — a chipmaker, a rival chipmaker, and a design-software incumbent, each with an obvious stake in whether spatial AI becomes the substrate for robotics, simulation, and 3D content. The company is not alone in the bet. World models have become a distinct and crowded research front, from Google DeepMind’s Genie line to startups like Runway’s Solaris and Veeda AI, the Sanja Fidler–led venture that raised $90 million on a similar premise. Atlas is World Labs’ move to define the category on its own terms.

Where it fits: robotics, games, and Marble

World Labs is positioning Atlas across three use cases, each at a different stage of readiness.

Robotics is the most concrete. The company describes Atlas as a component for real-to-simulation pipelines — turning captured real environments into synthetic ones for testing and training. Atlas can reconstruct a real space from sparse inputs, including phone footage, and then generate the RGB and depth observations a simulated robot would perceive as it moves through that space. That matters because the bottleneck in modern robot learning is data: physical trials are slow and expensive, and photorealistic simulation is one of the few ways to generate the volume of experience these systems need. It is the same sim-to-real problem that runs through efforts like Gemini Robotics and whole-body humanoid control, approached from the environment side rather than the policy side.

Games and visual effects are the flashier demonstrations. World Labs showed a “bullet time” reframing effect — the frozen-moment, orbiting-camera shot made famous by The Matrix — reconstructed from just a few cell phones and action cameras rather than the specialized rigs the effect traditionally requires. The implication is that camera-controlled, 3D-consistent scenes could be generated and re-shot after the fact, collapsing parts of the VFX pipeline.

Creator tools are the commercial anchor. World Labs already sells Marble, its consumer-facing world-generation product, which became generally available on November 12, 2025 with a free tier and paid plans running up to $95 a month; its World API sells credits at $1 for 1,250 credits (minimum $5), with a single Marble world costing about 1,500 credits, or roughly $1.20. World Labs says Atlas will power future versions of Marble and its other products — meaning the research model announced in early access is intended to become the engine behind a shipping consumer business, not just a lab demonstration.

The unknowns

For all the ambition, the announcement left the practical questions open. World Labs disclosed no pricing for Atlas itself, no general-availability timeline, and no production service-level commitments. The details that determine whether a model is deployable — inference cost, latency, throughput, and how it degrades on inputs unlike the curated demos — were not part of the launch. Early access “with select partners” is a controlled rollout, and the gap between a polished reveal and a dependable API can be wide.

The benchmark question compounds the uncertainty. The claim that one omni model beats specialized 3D models is exactly the kind of result that needs reproduction on independent data before it reshapes how teams build. Generation quality on scenes far from the training distribution — cluttered interiors, unusual lighting, motion — is where world models have historically struggled, and it is precisely what a robotics or VFX customer would stress first. The reliance on Gaussian splats and point clouds also inherits the known limits of those representations for editing and physical accuracy. Atlas is a serious entry; it is not yet a settled one.

What it means

Atlas is the strongest signal yet that world models are hardening into a distinct product category, separate from both large language models and text-to-video systems — and that the race to own it is on.

Who benefits. If the omni-model approach holds, the clearest early winners are robotics teams starved for training environments and 3D content pipelines in games and film, where reconstructing navigable scenes from cheap inputs would remove real cost. World Labs’ investors — NVIDIA and AMD on the compute side, Autodesk on the design-tools side — are positioned to benefit whichever way the category tips, which is part of why the round looks less like a bet on one company than on the substrate itself.

The tension to watch. The launch is heavy on capability and light on the numbers that decide adoption: price, availability, and independently verified benchmarks. A camera-controlled 1440p clip is a compelling demo, but the buyers who matter — robotics labs, studios, simulation platforms — will judge Atlas on cost per usable output and on how it behaves outside the reel. Generality is a strong story only if it survives contact with messy, out-of-distribution scenes.

What to watch next. Three signals over the coming months: whether Atlas graduates from partner-only early access to a priced, documented API with real SLAs; whether outside researchers can reproduce the claim that one omni model beats specialized 3D systems; and how quickly the model actually lands inside Marble and turns spatial intelligence from a thesis into a product people pay for. Li has spent two years arguing that AI has to leave the flat world of text and pixels. Atlas is the test of whether the market agrees.

Takina Takina · · 6 min read

Runway Solaris: The Interface World Model, Explained

Runway unveiled Solaris, an 'Interface World Model' that renders interactive apps frame by frame with no code. How it works, the benchmarks, and the caveats.

#AI #Frontend #World Models