AI
Nvidia Opens Cosmos 3 and Recasts How Robots Train
Nvidia Cosmos 3 unifies reasoning, world generation, and robot actions in one open model, then pulls the training factory onto its GPUs.
Nvidia released Cosmos 3, an open world foundation model that reasons about scenes, generates physics-aware video, and outputs robot actions from one network. Weights landed on Hugging Face the day after a May 31, 2026 launch at GTC Taipei. Nvidia says the family can cut physical AI training and evaluation cycles from months to days.
That speed claim is the bait. The follow-on effect is a shared physics prior that robot and vehicle teams can fine-tune, then run on Hopper, Blackwell, RTX PRO, and Jetson boxes Nvidia already sells.
Cosmos 3 Puts Reasoning, Worlds, and Action in One Model
Earlier Cosmos tools were a kit. Cosmos Predict made worlds, Cosmos Transfer steered them, Cosmos Reason read a scene, and Cosmos Policy proposed motion. Developers chained those pieces and paid for the glue. Cosmos 3 folds the same jobs into one omnimodel that Nvidia called the first fully open system of its kind.
Under the hood is a Mixture-of-Transformers architecture with two towers that share attention. An autoregressive reasoner reads text, images, and video. A diffusion generator then denoises video, ambient sound, and action tokens. The reasoner can run alone. The generator always uses both towers, so the model looks at object contact and motion before it draws the next frames or a joint trajectory.
Nvidia trained the family on 20 trillion tokens of multimodal data, including nearly a billion images, 400 million real and synthetic videos, audio, and action traces from people and robots. Action is a first-class channel, not a caption bolted onto pretty video. A sample control prompt on the Hugging Face card is blunt: “Put the pot to the left of the purple item.”
The big bang of physical AI is just around the corner thanks to breakthroughs in multimodal reasoning language, vision and world models. The Cosmos 3 family of open, frontier omnimodels gives developers a generational leap in ability to build robots, autonomous vehicles and vision AI that perceive, reason, plan and act in the physical world.
Jensen Huang, founder and CEO of Nvidia, May 31, 2026
Nvidia AI posted the launch demo the same week. The clip is the product pitch in motion: one model that describes a scene, rolls the world forward, and emits the move.
Introducing Cosmos 3: Our latest frontier model for Physical AI
Cosmos 3 is the world’s first fully open omnimodel with native vision reasoning, world and action generation.
Today we’re releasing Super (32B) and Nano (8B) variants. pic.twitter.com/6UfkSA7kzQ
— NVIDIA AI (@NVIDIAAI) June 1, 2026
A Data Problem No Test Fleet Can Finish
Chatbots can scrape the public web. A robot that grasps parts, or a car that meets fog and a darting pedestrian, cannot. Rare cases are costly to stage and unsafe to repeat on public roads. Ming-Yu Liu, vice president of Cosmos Lab, put the bind in one line after the weights went up.
“You can’t collect your way out of an infinite physical world. You have to generate it,” Liu wrote. Cosmos 3 is built as that generator, then as the policy head that trains on what it just imagined.
Depending on the inputs, the same checkpoint can act as a vision-language reasoner, a world simulator, a forward-dynamics model, an inverse-dynamics model, or a robot policy. Cosmos 2.5 still kept perception and generation in separate models and stuck to text, image, and video. Cosmos 3 adds sound and action and runs them in one forward pass.
That collapse is the industrial change. A lab that once hired three model teams can post-train one trunk on its cameras and arms. The catch, which Nvidia does not hide on its product page, is that the trunk is tuned for Nvidia GPUs, with NIM microservices and TAO recipes around it. Open weights travel. The recommended factory does not.
Super, Nano, and Edge on Three Kinds of Silicon
The family is one design at three budgets. Super is the teacher. Nano is the workstation brain. Edge is the on-device slice that arrived later.
COSMOS 3 MODEL LINE
| Variant | Size | Where it runs | Job |
|---|---|---|---|
| Cosmos 3 Super | 64B (32B reasoner + 32B generator) | Hopper and Blackwell datacenter GPUs | Highest-quality synthetic data and post-training |
| Cosmos 3 Nano | 16B (8B + 8B) | RTX PRO 6000 workstations | Fast video and action reasoning |
| Cosmos 3 Edge | 4B (2B Nemotron reasoner) | Jetson, DGX Spark, other edge boxes | On-device control without a cloud hop |
Checkpoints, post-training scripts, and five synthetic datasets (robot scenes, PhysX interactions, human motion, driving, warehouses) sit on Hugging Face under the Linux Foundation OpenMDW-1.1 license. Nvidia’s own Cosmos page says that license is the terms for its world foundation models. Cloud inference is on build.nvidia.com for teams without local H100-class iron.
Liu said the open-weight models took first place on VANTAGE-Bench and Traffic Anomaly Reasoning, on Artificial Analysis text-to-image and image-to-video boards, and on PAI-Bench, Physics-IQ, R-Bench, RoboArena, and RoboLab. Those scores are Nvidia’s leaderboard read, not a third-party audit published beside them.
Edge is the tell for the hardware funnel. Nvidia AI introduced the 4B model on July 20, 2026, as a world model that runs on-device for robots, road scenes, and live video agents. AWS later described Edge producing 32 actions per inference at 15 Hz on Jetson Thor. The free checkpoint still wants a Jetson at the gripper.
Who Already Trains on the Cosmos Stack?
Nvidia did not ship a lonely research demo. It stood up the Cosmos Coalition at launch and listed factories already on the platform.
WHO IS BUILDING ON COSMOS
- Coalition founders: Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI pledged to share models, research, and tests while using Cosmos training tools.
- Robot lines: Agile Robots, Doosan Robotics, LG Electronics, Samsung Electronics, and Skild AI are building robot apps on the stack.
- Driving: Li Auto is using Cosmos for autonomous vehicle work, while Waabi, Wayve, and Foretellix have used Cosmos models to simulate traffic, weather, and pedestrians.
- Vision agents: Centific, Fogsphere, Linker Vision, Milestone Systems, and Yuan are on the platform for industrial and smart-space video.
The car side is where the second-order loop is easiest to see. Nvidia’s September 10, 2026 robotaxi briefing said Omniverse NuRec rebuilds real drives from sensors, then Cosmos writes physically based variants so a few thousand corner cases become millions of weather, lighting, and behavior mixes. Uber plans NVIDIA DRIVE Hyperion robotaxis in 28 cities by 2028 and is building a data factory on Cosmos for rare fleet events. Mercedes-Benz is working with Nvidia and Uber on an S-Class robotaxi that uses DRIVE Hyperion, DRIVE AV L4 software, and Alpamayo models.
A Goldman Sachs projection cited in that briefing puts the robotaxi market at $400 billion by 2035. Whether those fleets mint money is a separate question of where auto AI actually pays off. Cosmos does not settle the P&L. It makes the synthetic miles cheap enough that more programs can stay in the fight, on Nvidia’s three-computer path from DGX training to RTX PRO simulation to DRIVE AGX in the vehicle.
The Gaps World Labs and Genie Still Own
Cosmos 3 is a world model in the industrial sense: rollouts you inspect, label, and pour into a policy. Other labs are selling different worlds, and they raised real money to do it.
THREE WORLD MODELS, THREE PRODUCTS
| Model | Builder | What it emits | How you use it |
|---|---|---|---|
| Cosmos 3 | Nvidia | Video, sound, and robot actions with a physics prior | Offline synthetic data and policy post-training |
| Genie 3 | Google DeepMind | Playable generated environments | interactive worlds at 24 frames per second, about 720p, coherent for a few minutes |
| Atlas | World Labs | Persistent 3D space from a photo | Walkable scenes you keep, rather than a one-shot clip |
World Labs, co-founded by Fei-Fei Li, closed a $1 billion round in February 2026 with a $200 million Autodesk check. Investor lists for that round included Nvidia and AMD. The company launched Atlas on September 1, 2026, as an omni world model trained across text, images, video, and 3D. Yann LeCun’s AMI Labs raised $1.03 billion in March 2026 to build world models around persistent memory and planning, a longer bet than a video generator. Nvidia has been named among backers on that side of the table as well.
So the open Cosmos release does not wipe those companies out. It sets a free baseline for “understand physics, emit video, emit actions.” Atlas can still sell geometry you can load into a game engine or a CAD tool. Genie 3 can still sell a world you drive in real time. AMI can still sell a learning method that is not a giant generative rollout. What they cannot sell, at least not as a scarce idea, is the mere fact of a world foundation model.
Video-first systems such as Cosmos and Genie show you the future in pixels. Latent approaches in LeCun’s line try to skip that paint. Cosmos 3 is Nvidia’s wager on the layered path: generate the world, reason over it, then attach a policy. Labs that only needed a pretty simulator now start from Super. Labs that needed a new theory of intelligence still have to build it.
AWS Parked the Flywheel on a Persistent Cluster
By late summer the launch had become a production recipe. That is the second-order story in plain metal: the model is free, the loop is a cluster you keep up.
FROM GTC TAIPEI TO THE CLUSTER
- May 31, 2026: Nvidia launches Cosmos 3 Super and Nano at GTC Taipei and opens the Cosmos Coalition.
- June 1, 2026: Super and Nano weights, Diffusers pipelines, and synthetic datasets go up on Hugging Face.
- July 20, 2026: Cosmos 3 Edge ships as a 4B on-device world model with a 2B Nemotron reasoner.
- August 27, 2026: AWS lists Edge, Nano, and Super on SageMaker JumpStart.
- September 4, 2026: AWS publishes a walkthrough for a standing Cosmos 3 training flywheel on SageMaker HyperPod.
The HyperPod note, by Nathan Arnold, Daniel Schoonover, and Eric Saleh, treats physical AI as a continuous factory rather than one training job. Synthetic generation, policy post-training, and closed-loop tests share a cluster and a multi-terabyte store. Their example jobs run on p5en.48xlarge nodes with eight H200 GPUs. The metric they care about is goodput, not a single-run FLOPS peak. That is an Physical AI model factory on HyperPod, and the GPUs inside it are still Nvidia’s.
Nvidia’s Cosmos page draws a hard line next to Omniverse. Omniverse builds 3D simulations with RTX rendering. Cosmos generates the video and action data that train the agent. Teams feed Omniverse clips into Cosmos Transfer-style paths to get photoreal synthetic sets, then bring policies back to the sim. Open model, closed loop, same vendor at both ends.
On September 10, 2026, Zheng Tang, a senior deep learning engineer at Nvidia, said a team had adapted Cosmos 3 Nano for generative traffic forecasting and won AI City Challenge Track 5 at ECCV 2026. That is a small lab result on an open checkpoint, which is what the release was for. The weights remain free to download. The factory that makes them useful still meters time on Nvidia silicon.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI3 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO3 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS3 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI3 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS3 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING3 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
