Nvidia unveils Cosmos 3 world model for robots

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Nvidia unveiled Cosmos 3, an open AI world model aimed at helping robots, autonomous vehicles and other physical systems understand and predict real‑world environments.
- Cosmos 3 was trained on 20 trillion tokens of multimodal data, including nearly a billion images, 400 million real and synthetic videos, ambient audio, text and action data from humans and robots.
- Ming‑Yu Liu, VP of Nvidia’s Cosmos Lab, said the model’s action data—such as robot joint angles, gripper positions and trajectories—distinguishes it from regular video generators and is key for autonomous actions.
- Nvidia released two versions now—a “super” model for high‑physics‑accuracy tasks like robot and autonomous‑vehicle training, and a “nano” model that can generate results in fractions of a second—while an “edge” model for local execution is slated for later.
- Nvidia is building a coalition of partners, including Agile Robots, Black Forest Labs and Runway, to support Cosmos 3 and enable customizations for hardware makers.
- Cosmos 3 can simulate rare or dangerous scenarios (e.g., robot collisions, unusual road events) that are costly or unsafe to capture in real life, helping developers train safer physical AI systems.
Why it matters: Hardware makers and robotics developers gain a customizable, open‑source model that can produce high‑fidelity action data and dangerous scenarios, reducing the need for costly real‑world data collection, while Nvidia secures a platform role in the emerging physical‑AI market and positions its ecosystem for future revenue from model licensing and cloud services.


