Connect with us

AI

Nvidia Opens Cosmos 3 and Recasts How Robots Train

Nvidia Cosmos 3 unifies reasoning, world generation, and robot actions in one open model, then pulls the training factory onto its GPUs.

Published

on

Nvidia released Cosmos 3, an open world foundation model that reasons about scenes, generates physics-aware video, and outputs robot actions from one network. Weights landed on Hugging Face the day after a May 31, 2026 launch at GTC Taipei. Nvidia says the family can cut physical AI training and evaluation cycles from months to days.

That speed claim is the bait. The follow-on effect is a shared physics prior that robot and vehicle teams can fine-tune, then run on Hopper, Blackwell, RTX PRO, and Jetson boxes Nvidia already sells.

Cosmos 3 Puts Reasoning, Worlds, and Action in One Model

Earlier Cosmos tools were a kit. Cosmos Predict made worlds, Cosmos Transfer steered them, Cosmos Reason read a scene, and Cosmos Policy proposed motion. Developers chained those pieces and paid for the glue. Cosmos 3 folds the same jobs into one omnimodel that Nvidia called the first fully open system of its kind.

Under the hood is a Mixture-of-Transformers architecture with two towers that share attention. An autoregressive reasoner reads text, images, and video. A diffusion generator then denoises video, ambient sound, and action tokens. The reasoner can run alone. The generator always uses both towers, so the model looks at object contact and motion before it draws the next frames or a joint trajectory.

Nvidia trained the family on 20 trillion tokens of multimodal data, including nearly a billion images, 400 million real and synthetic videos, audio, and action traces from people and robots. Action is a first-class channel, not a caption bolted onto pretty video. A sample control prompt on the Hugging Face card is blunt: “Put the pot to the left of the purple item.”

The big bang of physical AI is just around the corner thanks to breakthroughs in multimodal reasoning language, vision and world models. The Cosmos 3 family of open, frontier omnimodels gives developers a generational leap in ability to build robots, autonomous vehicles and vision AI that perceive, reason, plan and act in the physical world.

Jensen Huang, founder and CEO of Nvidia, May 31, 2026

Nvidia AI posted the launch demo the same week. The clip is the product pitch in motion: one model that describes a scene, rolls the world forward, and emits the move.

A Data Problem No Test Fleet Can Finish

Chatbots can scrape the public web. A robot that grasps parts, or a car that meets fog and a darting pedestrian, cannot. Rare cases are costly to stage and unsafe to repeat on public roads. Ming-Yu Liu, vice president of Cosmos Lab, put the bind in one line after the weights went up.

“You can’t collect your way out of an infinite physical world. You have to generate it,” Liu wrote. Cosmos 3 is built as that generator, then as the policy head that trains on what it just imagined.

Depending on the inputs, the same checkpoint can act as a vision-language reasoner, a world simulator, a forward-dynamics model, an inverse-dynamics model, or a robot policy. Cosmos 2.5 still kept perception and generation in separate models and stuck to text, image, and video. Cosmos 3 adds sound and action and runs them in one forward pass.

That collapse is the industrial change. A lab that once hired three model teams can post-train one trunk on its cameras and arms. The catch, which Nvidia does not hide on its product page, is that the trunk is tuned for Nvidia GPUs, with NIM microservices and TAO recipes around it. Open weights travel. The recommended factory does not.

Super, Nano, and Edge on Three Kinds of Silicon

The family is one design at three budgets. Super is the teacher. Nano is the workstation brain. Edge is the on-device slice that arrived later.

COSMOS 3 MODEL LINE

Variant Size Where it runs Job
Cosmos 3 Super 64B (32B reasoner + 32B generator) Hopper and Blackwell datacenter GPUs Highest-quality synthetic data and post-training
Cosmos 3 Nano 16B (8B + 8B) RTX PRO 6000 workstations Fast video and action reasoning
Cosmos 3 Edge 4B (2B Nemotron reasoner) Jetson, DGX Spark, other edge boxes On-device control without a cloud hop

Checkpoints, post-training scripts, and five synthetic datasets (robot scenes, PhysX interactions, human motion, driving, warehouses) sit on Hugging Face under the Linux Foundation OpenMDW-1.1 license. Nvidia’s own Cosmos page says that license is the terms for its world foundation models. Cloud inference is on build.nvidia.com for teams without local H100-class iron.

Liu said the open-weight models took first place on VANTAGE-Bench and Traffic Anomaly Reasoning, on Artificial Analysis text-to-image and image-to-video boards, and on PAI-Bench, Physics-IQ, R-Bench, RoboArena, and RoboLab. Those scores are Nvidia’s leaderboard read, not a third-party audit published beside them.

Edge is the tell for the hardware funnel. Nvidia AI introduced the 4B model on July 20, 2026, as a world model that runs on-device for robots, road scenes, and live video agents. AWS later described Edge producing 32 actions per inference at 15 Hz on Jetson Thor. The free checkpoint still wants a Jetson at the gripper.

Who Already Trains on the Cosmos Stack?

Nvidia did not ship a lonely research demo. It stood up the Cosmos Coalition at launch and listed factories already on the platform.

WHO IS BUILDING ON COSMOS

  • Coalition founders: Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI pledged to share models, research, and tests while using Cosmos training tools.
  • Robot lines: Agile Robots, Doosan Robotics, LG Electronics, Samsung Electronics, and Skild AI are building robot apps on the stack.
  • Driving: Li Auto is using Cosmos for autonomous vehicle work, while Waabi, Wayve, and Foretellix have used Cosmos models to simulate traffic, weather, and pedestrians.
  • Vision agents: Centific, Fogsphere, Linker Vision, Milestone Systems, and Yuan are on the platform for industrial and smart-space video.

The car side is where the second-order loop is easiest to see. Nvidia’s September 10, 2026 robotaxi briefing said Omniverse NuRec rebuilds real drives from sensors, then Cosmos writes physically based variants so a few thousand corner cases become millions of weather, lighting, and behavior mixes. Uber plans NVIDIA DRIVE Hyperion robotaxis in 28 cities by 2028 and is building a data factory on Cosmos for rare fleet events. Mercedes-Benz is working with Nvidia and Uber on an S-Class robotaxi that uses DRIVE Hyperion, DRIVE AV L4 software, and Alpamayo models.

A Goldman Sachs projection cited in that briefing puts the robotaxi market at $400 billion by 2035. Whether those fleets mint money is a separate question of where auto AI actually pays off. Cosmos does not settle the P&L. It makes the synthetic miles cheap enough that more programs can stay in the fight, on Nvidia’s three-computer path from DGX training to RTX PRO simulation to DRIVE AGX in the vehicle.

The Gaps World Labs and Genie Still Own

Cosmos 3 is a world model in the industrial sense: rollouts you inspect, label, and pour into a policy. Other labs are selling different worlds, and they raised real money to do it.

THREE WORLD MODELS, THREE PRODUCTS

Model Builder What it emits How you use it
Cosmos 3 Nvidia Video, sound, and robot actions with a physics prior Offline synthetic data and policy post-training
Genie 3 Google DeepMind Playable generated environments interactive worlds at 24 frames per second, about 720p, coherent for a few minutes
Atlas World Labs Persistent 3D space from a photo Walkable scenes you keep, rather than a one-shot clip

World Labs, co-founded by Fei-Fei Li, closed a $1 billion round in February 2026 with a $200 million Autodesk check. Investor lists for that round included Nvidia and AMD. The company launched Atlas on September 1, 2026, as an omni world model trained across text, images, video, and 3D. Yann LeCun’s AMI Labs raised $1.03 billion in March 2026 to build world models around persistent memory and planning, a longer bet than a video generator. Nvidia has been named among backers on that side of the table as well.

So the open Cosmos release does not wipe those companies out. It sets a free baseline for “understand physics, emit video, emit actions.” Atlas can still sell geometry you can load into a game engine or a CAD tool. Genie 3 can still sell a world you drive in real time. AMI can still sell a learning method that is not a giant generative rollout. What they cannot sell, at least not as a scarce idea, is the mere fact of a world foundation model.

Video-first systems such as Cosmos and Genie show you the future in pixels. Latent approaches in LeCun’s line try to skip that paint. Cosmos 3 is Nvidia’s wager on the layered path: generate the world, reason over it, then attach a policy. Labs that only needed a pretty simulator now start from Super. Labs that needed a new theory of intelligence still have to build it.

AWS Parked the Flywheel on a Persistent Cluster

By late summer the launch had become a production recipe. That is the second-order story in plain metal: the model is free, the loop is a cluster you keep up.

FROM GTC TAIPEI TO THE CLUSTER

  1. May 31, 2026: Nvidia launches Cosmos 3 Super and Nano at GTC Taipei and opens the Cosmos Coalition.
  2. June 1, 2026: Super and Nano weights, Diffusers pipelines, and synthetic datasets go up on Hugging Face.
  3. July 20, 2026: Cosmos 3 Edge ships as a 4B on-device world model with a 2B Nemotron reasoner.
  4. August 27, 2026: AWS lists Edge, Nano, and Super on SageMaker JumpStart.
  5. September 4, 2026: AWS publishes a walkthrough for a standing Cosmos 3 training flywheel on SageMaker HyperPod.

The HyperPod note, by Nathan Arnold, Daniel Schoonover, and Eric Saleh, treats physical AI as a continuous factory rather than one training job. Synthetic generation, policy post-training, and closed-loop tests share a cluster and a multi-terabyte store. Their example jobs run on p5en.48xlarge nodes with eight H200 GPUs. The metric they care about is goodput, not a single-run FLOPS peak. That is an Physical AI model factory on HyperPod, and the GPUs inside it are still Nvidia’s.

Nvidia’s Cosmos page draws a hard line next to Omniverse. Omniverse builds 3D simulations with RTX rendering. Cosmos generates the video and action data that train the agent. Teams feed Omniverse clips into Cosmos Transfer-style paths to get photoreal synthetic sets, then bring policies back to the sim. Open model, closed loop, same vendor at both ends.

On September 10, 2026, Zheng Tang, a senior deep learning engineer at Nvidia, said a team had adapted Cosmos 3 Nano for generative traffic forecasting and won AI City Challenge Track 5 at ECCV 2026. That is a small lab result on an open checkpoint, which is what the release was for. The weights remain free to download. The factory that makes them useful still meters time on Nvidia silicon.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending