Notes on Physical AI &
synthetic data
Notes on Physical AI and synthetic data from the datadoo research team.
Page 1 of 2

Notes on Skild S1 and NVIDIA's September data releases
Skild AI's S1 learns an unseen task from one video, roughly 380 teleoperated demonstrations saved by Skild's count. Notes on that release, NVIDIA's NuRec and robotaxi posts from the same weeks, and the quality-control ratio Skild published alongside.

Cosmos, GR00T, Nemotron: the open Physical AI stack, read from the data side
NVIDIA now ships an open model for each layer of Physical AI: Cosmos 3 generates worlds, GR00T acts in them, Nemotron 3 runs the agents around them. A plain map of the stack as it stands in August 2026, and the data work all three layers still assume you bring.

What SIGGRAPH 2026 means for synthetic training data
NVIDIA's SIGGRAPH 2026 program, read as a synthetic data release: a 4B world model with open weights running on workstation GPUs, and papers that automate scene reconstruction, materials, and motion. Generating data keeps getting cheaper. Knowing whether it transfers is still the expensive part.

Cosmos 3 thinks before it renders
At GTC Taipei, NVIDIA shipped Cosmos 3: a world model that reasons about a scene before it generates one. It folds the generate-and-render core of a synthetic-data pipeline into a single model. The validation half, the part that decides whether the data trains a production model or quietly breaks it, did not move.

NVIDIA opened the stack. The real game just started.
At GTC 2026, NVIDIA released Cosmos and the Physical AI Data Factory Blueprint under their Open Model License. The tooling fight is over. The operational fight — validation traces, sim-to-real proof, regulatory-grade lineage — is starting. Most companies are not ready for it.

Robots Are Shipping. Training Data Is Not.
Humanoids are deploying, $6B flowed into Physical AI in Q1 2026, and the bottleneck has shifted from hardware to training data. Physics-accurate synthetic data is the binding constraint.