Li Feifei's Atlas Opens the Door to a New World for Embodied AI

09/14 2026 431

Qingping Chuiguo | Author

star All Empty | Editor

Accidental Orange Juice | Visual

Liu Chao | Producer

"The world is everything that is the case."

This famous quote from Ludwig Wittgenstein's Tractatus Logico-Philosophicus, published in 1921, was referenced a century later by Li Feifei, a leading figure in AI, whose company World Labs is enabling AI to understand space. Recently, World Labs released Atlas, a new-generation world model described as the "world's first" multimodal model capable of processing text, images, videos, and 3D data simultaneously.

The demo is truly impressive: feed in photos from two or three angles, and it can reconstruct the entire scene, then move the camera to positions where no one has ever stood.

This undoubtedly opens a new world for the industry:

Traditional 3D reconstruction has long relied on dense photography, requiring multi-angle, high-coverage image collection around a scene to piece together a relatively complete model. This method demands high equipment, time, and labor costs, making it difficult to scale 3D content production and keeping the barrier high.

Atlas is changing this paradigm. Leveraging the powerful prior knowledge and generative capabilities of large models, it achieves high-fidelity 3D reconstruction using only a small number of images captured with an ordinary smartphone.

As a result, the barrier to 3D content creation is significantly lowered. No professional equipment or dense collection is needed, and vast amounts of old photos and videos can be transformed into usable 3D assets, reactivating the dormant potential of data.

But that's not all. Traditional AI generative models primarily serve creative industries like film and gaming, focusing on generating "realistic-looking" content.

Atlas, however, expands AI's generative capabilities into the robotics field through its interactive and deductive "world simulator," enabling Real-to-Sim.

It transforms generated results from static images into dynamic environments that can be entered, operated, and verified.

In robotics, this means robots can undergo extensive trial-and-error and strategy training in AI-generated simulated environments, predicting the consequences of different actions through "imagination," accumulating experience rapidly in the virtual world, and then transferring it to the real world.

This directly addresses the core bottleneck of scarce and costly real-world training data for robots, providing a low-cost, high-efficiency, and scalable training pathway for embodied AI, accelerating the transition of robots from laboratories to real-world applications.

Give Atlas a Photo, and It Can Create a World

Let's first see what Atlas has accomplished.

Over the past two years, camera control in video generation models has relied on random luck.

You provide a sentence or an image, and it imagines the scene, with camera movements largely dependent on random prompts, generating different results each time.

Atlas, however, changes the underlying logic by incorporating the "camera" directly into the model.

Camera pose parameters are native inputs, allowing you to specify angles and trajectories by feeding data directly into coordinates.

World Labs calls this mechanism "spatial context." Each image fed into the model is locked to a specific coordinate in 3D space, with camera posture and depth information brought into the model along with the image.

The model doesn't work with a stack of isolated photos but a world sketch with coordinates.

Thus, you can direct it like sitting in a director's chair.

The most viral segment of the official demo uses three to five ordinary smartphones to capture the same moment, allowing the model to freeze time and move the camera into the splashing milk.

Over two decades ago, the bullet-time shot in The Matrix required hundreds of cameras circling a green screen studio; now, a few smartphones can roughly replicate it.

The 3D reconstruction is equally impressive.

According to World Labs' official blog, the aerial shot of Stanford's Main Quad was reconstructed using only 2 to 25 tourist photos taken casually from the ground.

Traditional 3D reconstruction requires capturing every corner, often involving hundreds of rotations around the site; Atlas needs only a few to dozens of images.

Seeing this, you might be amazed, exclaim (exclaiming) that AI now understands the physical world!

But that's not quite the case.

The official blog and interviews clarify two points: first, Atlas has achieved "spatial context" understanding at the geometric level; second, the team's long-term goal is for the model to understand physical laws. Many confuse the two.

Geometry and physics are fundamentally different.

Geometry is the Entry Ticket; Physics is the Challenge

What's the difference between geometry and physics?

Geometry answers, "Where is the object, what does it look like, and does it retain its shape from a different angle?"; physics answers, "How will it move if pushed, will it topple, how much force will shatter it?"

Atlas's impressive feats mostly fall under geometry, with collision, friction, and deformation—real physics tasks—barely demonstrated in the official demo.

World Labs is aware of this.

In a June article comb, sort out, organize, arrange, streamline (reviewing) world models, they self-deprecatingly called "world model" the most overused term in AI, categorizing products into three tiers: renderers that only produce images, simulators that generate physically plausible states, and planners that output actions.

By this measure, Atlas has taken a significant step from "renderer" toward "simulator" but is still far from truly understanding physics.

Traditionally, evaluating world models focused on visual resemblance, essentially comparing pixel distributions. Models could generate realistic but physically implausible images and still score high.

But the criteria are changing.

WorldArena, an embodied AI benchmark, now singles out "trajectory accuracy" and "physical plausibility" as evaluation metrics.

Notably, in the inaugural WorldArena 2.0 rankings in late August, China's InSpatio-Curious by Yingsu topped 77 models despite scoring nearly 9 points lower in visual quality than the runner-up, thanks to superior trajectory precision, depth accuracy, and spatiotemporal consistency.

Clearly, "visual beauty" is becoming less valuable in new evaluations.

Additionally, WorldBench, released in January by UCLA, Yale, Sony AI, and the U.S. Army Research Lab, tests physical concepts individually.

Their verdict was harsh: none of the tested models passed in generating reliable physical interactions.

Thus, Atlas's performance—without a published paper, model card, or third-party physical evaluation—should be viewed cautiously. Its official demo used precise trajectories while competitors relied on textual descriptions, giving it an inherent advantage.

More accurately, Atlas demonstrates state-of-the-art camera control but hasn't undergone rigorous physical testing. WorldBench's conclusion that all models fail in physical interactions still stands.

Why, then, is everyone rushing toward "physical correctness"?

Atlas's New Pathway Could Become the "Training Ground" for Embodied AI

Because physical correctness is a prerequisite for using simulated data.

In a post-launch interview, Li Feifei stated that data, not chips, is the current bottleneck for robots.

While the internet has abundant text, images, and videos, robots need experiences with physical consequences, like "what happens if I squeeze this at this angle with this force."

But real-world data collection is prohibitively expensive.

The World Labs founding team cited an example: reconstructing a space traditionally required hundreds to thousands of professional shots; for amateurs, wandering around a room for an hour or two was common.

Without cost reduction, robot training data will remain scarce.

Thus, "sim-to-real" has shifted from a lab curiosity to a competitive frontier:

Quickly transfer real environments into virtual worlds, then randomize them—bend wires differently, change cup colors, vary lighting—letting robots make mistakes virtually first.

In July, World Labs quietly acquired SceniX, a robotics simulation specialist, marking Atlas as the first model to embody this strategy.

Zooming out, creating AI "training grounds" is not new.

In the 1930s, Edwin Link built the first flight simulator, training tens of thousands of Allied pilots during WWII before they flew real planes. Back then, even transistors didn't exist; mechanical devices simulated instrument reactions.

In autonomous driving, Waymo disclosed that its systems logged orders of magnitude more test miles in simulation than on real roads.

If planes trained in simulators and cars in virtual roads, now it's AI's turn to train in simulated physical worlds.

Understanding capital flows becomes easier with this historical context.

Billions have poured into world models, betting not on prettier images but on who can most cheaply transfer the real world into models for AI trial-and-error, positioning themselves at the forefront of next-gen AI.

Atlas's approach could slash data collection and simulation costs, enabling startups to train robots without heavy hardware investments, popularizing the "simulate first, validate in real" paradigm.

It's still early for AI to truly understand the physical world.

But Atlas proves that geometric correctness is engineering-solvable and, once achieved, immediately translates to product strength.

The world model race is shifting from "who generates more realistic images" to "who generates more physically plausible states."

From renderer to simulator, this leap is AI's first true touch of the physical world.

Atlas has cracked the door open; physical AI finally has its first foothold.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.