Source-linked AI summary

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilatyev, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Vikram Fugro, Prashant Gaikwad, TJ Galda, Katelyn Gao, Yihuai Gao, Wenhang Ge, Sreyan Ghosh, Arushi Goel, Vivek Goel, Akash Gokul, Rama Govindaraju, Jinwei Gu, Miguel Guerrero, Elfie Guo, Aryaman Gupta, Siddharth Gururani, Hugo Hadfield, Song Han, Ankur Handa, Zekun Hao, Mohammad Harrim, Ali Hassani, Nathan Hayes-Roth, Yufan He, Chris Helvig, Cyrus Hogg, Madison Huang, Michael Huang, Sophia Huang, Yufan Huang, Jacob Huffman, DeLesley Hutchins, Suneel Indupuru, Boris Ivanovic, Arihant Jain, Joel Jang, Ryan Ji, Yanan Jian, Dongfu Jiang, Jingyi Jin, Atharva Joshi, Nikhilesh Joshi, Pranjali Joshi, Andy Ju, Jaehun Jung, Weiwei Kang, Scott Kassekert, Jan Kautz, Ashna Khetan, Julia Kiczka, Slawek Kierat, Gwanghyun Kim, Kuno Kim, Sunny Kim, Kezhi Kong, Xin Kong, Zhifeng Kong, Tomasz Kornuta, Egor Krivov, Hui Kuang, Saurav Kumar, Chia-Wen Kuo, George Kurian, Wojciech Kutak, JF Lafleche, Himangshu Lahkar, Omar Laymoun, Jayjun Lee, Sanggil Lee, Gabriele Leone, Boyi Li, Freya Li, Jiajun Li, Jinfeng Li, Ling Li, Pengcheng Li, Shangru Li, Tingle Li, Xiaolong Li, Xuan Li, Zhaoshuo Li, Zhiqi Li, Hao Liang, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Sifei Liu, Zihan Liu, Hai Loc Lu, Xiangyu Lu, Alice Luo, Ruipu Luo, Wenjie Luo, Jiangran Lyu, Martin Ding Ma, Nic Ma, Qianli Ma, Dawid Majchrowski, Louis Marcoux, Miguel Martin, Qing Miao, Ashkan Mirzaei, Shreyas Misra, Kaichun Mo, Durra Mohsin, Hyejin Moon, Pawel Morkisz, Saeid Motiian, Kirill Motkov, Seungjun Nah, Yashraj Narang, Deepak Narayanan, Thabang Ngazimbi, Julian Ouyang, Shubham Pachori, David Page, Yatian Pang, Sehwi Park, Mahesh Patekar, Mostofa Patwary, Marco Pavone, Trung Pham, Wei Ping, Soha Pouya, Shrimai Prabhumoye, Varun Praveen, Delin Qu, Hesam Rabeti, Morteza Ramezanali, Marilyn Reeb, Xuanchi Ren, Kristen Rumley, Wojciech Rymer, Jun Saito, Yeongho Seol, John Shao, Piyush Shekdar, Tianwei Shen, Humphrey Shi, Min Shi, Stella Shi, Kevin Shih, Mohammad Shoeybi, Mateusz Sieniawski, Shuran Song, Alexander Sotelo, Amir Sotoodeh, Sunil Srinivasa, Vignesh Srinivasakumar, Bartosz Stefaniak, Rahul Heinrich Steiger, Shangkun Sun, Jiaxiang Tang, Shitao Tang, Yangyang Tang, Yue Tang, Tolou Tavakkoli, Kayley Ting, Krzysztof Tomala, Wei-Cheng Tseng, Jibin Varghese, Sergei Vasilev, Thomas Volk, Raju Wagwani, Roger Waleffe, Andrew Z. Wang, Boxiang Wang, Haoxiang Wang, Qiao Wang, Shihao Wang, Shijie Wang, Ting-Chun Wang, Yan Wang, Yu Wang, Rohit Watve, David Wehr, Fangyin Wei, Xinshuo Weng, Jay Zhangjie Wu, Kedi Wu, Hongchi Xia, Summer Xiao, Tianjun Xiao, Kevin Xie, Daguang Xu, Jiashu Xu, Mengyao Xu, Ruqing Xu, Xingqian Xu, Yao Xu, Dinghao Yang, Dong Yang, Hans Yang, Xiaodong Yang, Xuning Yang, Yichu Yang, Yurong You, Zhiding Yu, Hao Yuan, Simon Yuen, Xiaohui Zeng, Pengcuo Zeren, Cindy Zha, Haotian Zhang, Jenny Zhang, Jing Zhang, Liangkai Zhang, Paris Zhang, Shun Zhang, Xuanmeng Zhang, Zhizheng Zhang, Ann Zhao, Yilin Zhao, Yuliya Zhautouskaya, Charles Zhou, Fengzhe Zhou, Shilin Zhu, Yuke Zhu, Dima Zhylko, Artur Zolkowski

arXiv:2606.02800v4cs.CVcs.AIcs.LGcs.MMcs.RO

TL;DR

Physical AI training requires models that can both understand observations and generate plausible futures, while prior work largely treated these capabilities separately. Cosmos 3 unifies language, image, video, audio, and action understanding and generation in one architecture, achieving state-of-the-art performance across most evaluated capabilities.

  • Problem

    Physical AI needs coupled understanding and generation for safe, scalable learning, but prior work largely treated these capabilities in separate models.

  • Method

    Cosmos 3 jointly models language, image, video, audio, and action across understanding and generation modes within one omnimodal framework.

  • Results

    Cosmos 3 establishes a new state-of-the-art across most evaluated capabilities, remaining highly competitive with or outperforming specialized models.

  • Takeaways & Limitations

    Cosmos 3 provides a unified foundation for developing Physical AI agents across multimodal understanding and generation.

  • Takeaways & Limitations

    Public PAIBench-G Image-to-Video leaderboard results were not reproducible, so the authors used an internal judge and independent leaderboard submissions.

Abstract

from arXiv · show

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial Analysis, and the best policy model by RoboArena at the time the technical report was written. To accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated synthetic datasets, and evaluation benchmark available under the Linux Foundation's OpenMDW-1.1 License at https://github.com/nvidia/cosmos and https://huggingface.co/collections/nvidia/cosmos3. The project website is available at https://research.nvidia.com/labs/cosmos-lab/cosmos3.

1. Introduction

Cosmos 3 addresses the costly and risky training of Physical AI by unifying understanding and generation across language, image, video, audio, and action in one omnimodal world-model framework. It serves as a general-purpose backbone and establishes a new state-of-the-art across most evaluated capabilities.

  • Motivation: Training Physical AI directly in the real world is slow, expensive, and potentially dangerous, motivating safe and scalable learning in simulated worlds.Simulated training must support both understanding and generation, which are fundamentally coupled capabilities for Physical AI agents.
  • Motivation: Separating understanding from generation is fundamentally limiting because each requires reasoning about world evolution, actions, structured representations, and agent behaviors.The paper argues that a unified scalable framework is essential for Physical AI.
  • Cosmos 3: Cosmos 3 jointly models language, image, video, audio, and action for understanding and generation, unifying diverse Physical AI model classes in one framework.Its input-output configurations support modes including vision-language understanding and reasoning, image and audio-visual generation, policy or world-action modeling, and dynamics modeling.
  • Adaptation: Cosmos 3 can be post-trained on target data for distinct applications without architectural modifications, including synthetic data generation and robot policy.The paper presents post-training for better synthetic data generation and better robot policy.
  • Results: Cosmos 3 establishes a new state-of-the-art across most capabilities while remaining highly competitive with or outperforming specialized models.The results overview reports consistent gains over specialized open-source baselines across capabilities.

2. Model Architecture

Cosmos 3 uses a unified mixture-of-transformers architecture that processes and generates language, vision, audio, and action through modality-specific encoders and shared multimodal representations. Its action interface maps heterogeneous embodiment controls into a common latent space while preserving domain-specific structure.

  • Unified multimodal backbone: Cosmos 3 treats language, vision, audio, and action as core modalities, using dedicated action tokens to connect physical control signals with language reasoning and video world modeling.Modality-specific encoders project inputs into the transformer representation space, while action tokens directly support physically grounded interaction.
  • Unified multimodal backbone: Learnable modality-specific embeddings are added to every non-language modality so shared transformer parameters and positional embeddings distinguish modalities.Language, vision, audio, and action inputs are first embedded into a unified representation space.
  • Modality encoders: Visual processing uses separate ViT and video-VAE encoders for understanding and generation, respectively.The ViT uses 16 × 16 patches, followed by token merging and projection into the transformer latent space.
  • Action representation: Action modeling supports autonomous vehicles, camera motion, robots, and egocentric human motion by mapping native controls into a unified action interface.Representations can include ego poses, effector poses, and grasp states, with domain-aware projections preserving embodiment-specific structure while sharing the MoT backbone.
  • Token and task organization: The token sequence places an autoregressive reasoning subsequence before a diffusion generation subsequence, routing them to separate parameter sets while coupling them through joint attention.Autoregressive tokens handle language and ViT-encoded visual inputs; diffusion tokens include VAE-encoded vision, audio, and action and are iteratively denoised.

24 FPS

The section highlights language, audio, and action as packed modalities.

  • Language, audio, and action are presented together as packed modalities.

30 FPS

Cosmos 3 spaces temporal positions according to real-world time rather than token count. FPS modulation aligns temporal coordinates across video, audio, and action data with different sampling rates.

  • Temporal position spacing reflects real-world time rather than token count across modalities.
  • FPS modulation accounts for differing physical time intervals caused by modality- and source-specific sampling rates.For example, a temporal-index increment for 24-FPS video corresponds to an interval 2.5 times longer than at 60 FPS.
  • Temporal resolution is characterized with TPS, computed from frame rate and temporal compression for video, hop size for audio, and sampling frequency for action.The video temporal compression factor is 4, while audio TPS uses 48 kHz and a 1920 hop size, giving approximately 25 TPS.
  • The base temporal resolution is set to TPSbase = 24/4 = 6, reflecting the prevalence of video data and 24-FPS video.

3. Data

Cosmos 3 trains complementary Reasoner and Generator pathways with distinct multimodal data curricula. The data strategy combines broad pre-training, Physical AI specialization, quality filtering, structured captions, and action-conditioned supervision.

  • Training objectives: The Reasoner learns world understanding from paired vision-language data, while the Generator learns synthesis, simulation, and action from complementary multimodal data.Both pathways share transformer and token representations but use different training-data types.
  • Training curricula: Both pathways use evolving multi-stage curricula, with broad pre-training followed by specialized Physical AI training or progressively added modalities.Reasoner specialization covers robotics, autonomous driving, and spatial intelligence; Generator training adds actions, control-conditioned transfer, and targeted synthetic data.
  • Reasoner data: 24.2M samples comprise the Reasoner curriculum: 22.0M for pre-training and 2.2M for supervised fine-tuning.Pre-training emphasizes image–text and text-only data, whereas supervised fine-tuning shifts toward Physical AI specialization and video–text samples.
  • Data curation: 4.23% of data is removed as multimodal near-duplicate supervision, while judge thresholds of 2 for pre-training and 5 for SFT balance coverage against annotation quality.Threshold 2 preserves reasoning, grounding, OCR, captioning, and VQA coverage; threshold 5 retains only the highest-confidence SFT examples.
  • Pre-training composition: 9.44M OCR samples constitute 42.9% of the approximately 22M-sample final pre-training mixture, followed by 3.62M 2D-grounding samples at 16.5%.Other major components include visual QA at 2.48M samples (11.3%) and image reasoning at 1.66M (7.5%).
  • Action data: Paired text-video-action data exposes the Generator to controllable interventions that connect observed world states across time.The data supports learning how different robot commands, camera trajectories, vehicle routes, or human hand motions alter future evolution.

4. Training

Cosmos 3 training combines staged multimodal pre-training and curated Physical AI fine-tuning for the Reasoner, followed by Generator initialization and a progressive curriculum spanning images, video, audio, and actions. Mid-training and post-training extend the unified model into specialized text-to-image, image-to-video, and robot-policy systems with leading reported benchmark and leaderboard results.

  • Generator Training: Reasoner weights initialize the Generator, which then follows a progressive curriculum from image, video, and audio pre-training to action and transfer data and curated Physical AI post-training.The shared transformer-block architecture transfers semantic and world knowledge into generation of pixels, audio, and actions.
  • Reasoner Training: The Reasoner is trained through large-scale image–text and video–text pre-training followed by supervised fine-tuning on curated Physical AI tasks.Fine-tuning targets robotics, autonomous driving, and smart infrastructure while preserving broad capabilities acquired during pre-training.
  • Generator Training: The Generator uses rectified flow matching across modalities, predicting constant velocity from noisy latents with masked mean-squared error.The noisy latent is formed by straight-line interpolation between clean targets and Gaussian noise, with conditioning tokens supplied as context.
  • Multimodal Integration: Mid-training expands image, video, and audio generation into a unified Physical AI model that consumes and synthesizes action and control signals while retaining existing visual capabilities.Action, audio, control, and video tokens share the same temporal modeling framework under the clean-prefix/noisy-target formulation.
  • Post-Training: Post-training produced specialized checkpoints that ranked top-1 among open-weight models on Artificial Analysis Text-to-Image and Image-to-Video leaderboards, while Cosmos3-Nano-Policy-DROID ranked first on RoboLab, RoboArena, and MolmoSpaces.Cosmos3-Super-Text2Image also achieved 91.36 on UniGenBench’s full benchmark.

5. Infrastructure

Cosmos 3’s infrastructure unifies scalable multimodal data curation, distributed training mechanisms, and optimized attention, checkpointing, and inference. These systems improve efficiency while supporting inspection, retrieval, deduplication, and flexible model execution.

  • Data infrastructure: SILA unifies storage, metadata management, distributed processing, semantic retrieval, visualization, and debugging for large-scale multimodal data curation.It integrates assets, curation signals, embeddings, vector indexes, and execution state within a Lance-backed substrate.
  • Data infrastructure: 10× throughput increase and startup latency reduced from 30–60 minutes to roughly 5 minutes demonstrate SILA’s efficiency gains over the previous infrastructure.Staged execution, checkpointing, fragment-level coordination, and improved cluster utilization contributed to the improvement.
  • Training infrastructure: The data loader combines token-budgeted packing, joint stream loading, rank-synchronous selection, and look-ahead buffering to reduce padding without permanently reordering samples.Look-ahead diverts oversized samples temporarily and restores buffered samples in their original arrival order.
  • Training infrastructure: 22% improvement in end-to-end training throughput results from two variable-length attention kernel launches that jointly support causal and bidirectional masking.The custom two-way flat attention mechanism eliminates padding overhead in fixed-length implementations.
  • Training infrastructure: 13% improvement in end-to-end training throughput for Cosmos3-Nano at a per-batch token budget of 74,000 tokens is achieved with SAC and no numerical-result change.Asynchronous checkpointing further reduces end-to-end training time by 4% for Cosmos3-Nano and 9% for Cosmos3-Super at a 30-minute interval.
  • Inference infrastructure: CUDA Graphs speed up T2I generation by 30% to 60%, while T2V request batching improves 256p throughput by 8% to 55%.Batching benefits diminish at 480p and disappear at 720p because the context window admits only B=1.

6. Results

Cosmos 3 achieves strong results across general, image, video, audio, and embodied Physical AI evaluations. Its post-trained variants set leading open-source or state-of-the-art results on several public and internal benchmarks, while audio quality remains a relative limitation.

  • General benchmarks: Cosmos 3 shows stronger general capabilities than Cosmos-Reason2 and outperforms open- and closed-source models in robotics, smart infrastructure, and driving domains.Its gains are attributed to 20% additional pre-training data that increases data diversity, although it still trails Gemini 3.1 Pro.
  • Text-to-image generation: Cosmos3-Super-Text2Image ranked #1 among open-weight models and #4 among all models on the Artificial Analysis Text-to-Image leaderboard.The evaluation used crowdsourced public voting in a real-world submission with an agentic harness.
  • Video generation: Cosmos3-Super achieves the highest overall PAIBench-G scores on both T2V and I2V, while Cosmos3-Nano matches the second-best RBench result.Both variants significantly outperform their Cosmos-Predict2.5 predecessors; internal and public evaluations rank Cosmos3-Super and Cosmos3-Nano first and second overall.
  • Video generation: Cosmos3-Super is the best open-source model on HUE T2V at 89.3 and HUE I2V at 89.6, trailing Veo-3.1 by 0.1 points on I2V.On HWB, it achieves 71.9, outperforming Veo-3.1 at 67.8 and Wan2.2-A14B at 60.7; Cosmos3-Nano scores 66.9.
  • Image-to-video generation: Cosmos3-Super-Image2Video ranked first among all open-weight models and #22 globally on the Artificial Analysis Image-to-Video leaderboard.The model was evaluated through crowdsourced public voting without audio and was described as on par with proprietary offerings such as Veo 3.1.
  • Audio-visual generation: Cosmos 3 is strongest on semantic and alignment audio metrics, with Cosmos3-Nano obtaining the best SAV, SA, and AVAlign scores and Cosmos3-Super the best visual-support score.Closed-source systems retain higher overall AVQ through stronger perceptual audio quality; Cosmos 3’s remaining headroom is concentrated in low-level audio fidelity.
  • Robot manipulation: Cosmos3-Nano-Policy-DROID achieves new state-of-the-art results across RoboLab, RoboArena, and MolmoSpaces, demonstrating effective policy post-training for omnimodal world models.The policy is reported to tolerate failures, retry when necessary, and remain robust to human interventions.

7. Related Work

Related work spans latent and generative world models, multimodal reasoning, video generation, action modeling, and embodied intelligence. Cosmos 3 extends omnimodal modeling toward Physical AI by unifying understanding and generation while addressing physical-world dynamics, action-conditioned generation, inverse dynamics, and embodied control.

  • World Models: World models divide into predictive latent models for compact planning and control and generative models that expose future observations as multimodal simulations.Generative interfaces make errors in geometry, contact, timing, and sound directly observable.
  • Multimodal Understanding: Physical intelligence requires maintained, actionable scene estimates that track space, time, affordances, object state, and task progress rather than isolated recognition or captioning.Cosmos-Reason1 addresses this setting through physical common sense and embodied chain-of-thought reasoning.
  • Video Generation: Video-generation research increasingly targets long-horizon realism, prompt adherence, controllability, and high-fidelity dynamics, but perceptual plausibility alone is insufficient for Physical AI world simulation.World simulation requires rollouts to preserve physical consistency, including timing and object identity.
  • Action Modeling: Action modeling covers forward dynamics, inverse dynamics, and policy mapping across heterogeneous action spaces including robotics, vehicles, cameras, human bodies, and egocentric motion.These settings differ in units, rates, and causal scopes.
  • Omnimodal Foundation Models: Existing omnimodal systems unify understanding and generation, but emphasize text-image or general media modeling more than physical dynamics, action conditioning, inverse dynamics, and embodied control; Cosmos 3 extends this direction to Physical AI.The paper describes Cosmos 3 as pairing an autoregressive reasoner tower for multimodal understanding with further capabilities.

8. Conclusion

Cosmos 3 is a family of omnimodal world models for Physical AI that unifies multimodal understanding and generation across language, image, video, audio, and action. Its single architecture reduces the need to compose separate models for vision-language understanding, video generation, world modeling, and action.

  • Cosmos 3 presents a family of omnimodal world models designed for Physical AI.
  • The unified architecture supports multimodal understanding and generation across language, image, video, audio, and action.
  • Cosmos 3 reduces the need to compose separate vision-language, video generation, world, and action models.
  • The formulation uses modality-specific encoders, structured token arrangements, and a Mixture-of-Transformers backbone.

A. Caption Details … B. Default Prompts and Prompt Upsampling Templates for Generator

The appendix describes structured image and video captioning pipelines, including spatial and temporal annotation strategies, schemas, and implementation choices. It also specifies modality-specific prompt-upsampling templates that convert user inputs or conditioning media into generator-ready JSON prompts.

  • A. Caption Details: In-house captioners emit predefined JSON semantic fields rather than free-form captions, improving control, detail recall, and annotation consistency.The schemas cover visual attributes such as subjects, composition, background, lighting, aesthetics, style, viewpoint, and cinematography.
  • A.1. Captioner Models: The captioning pipeline scans four image quadrants plus the center independently to improve coverage of multiple subjects and complex spatial layouts.Region descriptions are merged into the final image caption.
  • A.1. Captioner Models: LoRA fine-tuning of Qwen3-VL-8B provided the best balance between benchmark metrics and inference efficiency for final image and video captioners.Video inputs use 8 FPS, generation temperature 0.7, and Qwen3-VL pixel bounds of 131, 072 to 25, 165, 824.
  • A.2. Image Schema: The image schema represents subjects, scene context, lighting, aesthetics, cinematography, style, visible text, spatial regions, captions, and image metadata in structured fields.Subject records can include appearance, placement, pose, human attributes, and count-sensitive anatomy such as numbers of subjects, arms, hands, fingers, and legs.
  • A.3. Video Schema: The video schema extends image annotations with actions, state changes, temporal segments, transitions, camera motion, temporal captions, audio descriptions, duration, and FPS.The number of temporal entries varies with video complexity, including scene changes, interactions, and camera movements.
  • B. Default Prompts and Prompt Upsampling Templates for Generator: Cosmos 3 Reasoner prompt upsampling converts user prompts into structured JSON for text-to-video, text-to-image, image-to-video, and video-to-video generation.Templates enforce modality-specific fields, fixed or copied sampling metadata, temporal consistency, and continuity with attached conditioning media when applicable.
  • B. Default Prompts and Prompt Upsampling Templates for Generator: The post-trained Cosmos3-Super-Text2Image upsampler densely populates every template field with plausible, scene-consistent details while allowing empty values only for truly inapplicable fields.It requires all top-level keys and forbids adding keys beyond the template.
  • B. Default Prompts and Prompt Upsampling Templates for Generator: The Cosmos3-Super-Image2Video template grounds motion descriptions in the starting frame and requires chronological, physically accurate, causally ordered, persistent, spatially explicit actions.It also specifies unambiguous body-side references, singular pronouns for single subjects, restrained camera motion, and cinematography-aware wording.

C. Synthetic Dataset for Generator Training … C.5. SDG-Warehouse

The generator mid-training uses five synthetic data-generation datasets spanning physical interactions, robotics, autonomous driving, digital humans, and warehouse safety events. Together, they provide diverse videos with simulator-derived supervision and targeted coverage of phenomena that are difficult to obtain from real footage.

  • C. Synthetic Dataset for Generator Training: Five SDG datasets are used for generator mid-training, with their scale, modalities, and dataset cards summarized in Table 22.The datasets cover SDG-PhyxSim, SDG-RobotSim, SDG-DriveSim, SDG-SynHuman, and SDG-Warehouse.
  • C.1. SDG-PhyxSim: SDG-PhyxSim provides large-scale synthetic videos of physically simulated multi-object interactions with synchronized views, segmentation, metric depth, and per-object physics annotations.It includes four fixed camera viewpoints and simulator-read dynamics such as velocity, rotation, and center-of-mass displacement.
  • C. Synthetic Dataset for Generator Training: Across these datasets, simulation infrastructure enables procedural variation and synchronized multimodal supervision for training physical-world generators.The pipelines use Isaac Sim and related systems to vary scenes, agents, sensors, environments, and rendering conditions while producing structured annotations.
  • C.1. SDG-PhyxSim: 76,489 independent simulation runs in SDG-PhyxSim yield approximately 57 M RGB frames and 1,529,752 rendered MP4 files across varied rigid-body interaction scenes.Most clips are 5 s and 150 frames, while ball_mixer clips are 8 s and 240 frames; all are rendered at 1920×1080 and 30 fps.
  • C.2. SDG-RobotSim: SDG-RobotSim is a fully synthetic robotics corpus targeting embodiment persistence, contact understanding, long-horizon video modeling, and action-conditioned reasoning across motion, manipulation, and collision.Its public v1.0 release contains 386,270 RGB MP4 clips, with simulator metadata and generator-specific annotations such as robot state, contacts, and task success.
  • C.3. SDG-DriveSim: SDG-DriveSim targets rare autonomous-driving interactions under varied environments, including emergency vehicles, obstacle nudging, cut-ins, degraded weather, and non-standard pedestrian crossings.Its 264,000 clips total approximately 1,467 hours and are distributed across seven scenario families.
  • C.4. SDG-SynHuman: SDG-SynHuman supplies temporally consistent digital-human videos with dense metric-depth and deterministic camera supervision across diverse people, environments, lighting, animations, and trajectories.The final release contains 236,937 clips totaling 5,841 hours, with paired depth frames and camera parameters at the same temporal resolution.
  • C.5. SDG-Warehouse: SDG-Warehouse renders four indoor industrial-safety scenarios—near-misses, fires, shelf collisions, and box pickup—with multiple synchronized camera viewpoints and dense per-frame annotations.The release contains approximately 123K clips and approximately 412 hours at 1920×1080 and 30 fps, including depth, segmentation, edges, and 2D/3D boxes.

C.6. Distribution of SDG Datasets

The analysis tests whether synthetic content complements the pre-training corpus by locating SDG embeddings relative to the pre-training distribution and quantifying their manifold distances. Together, Fig. 36 and Tab. 25 indicate that SDG occupies long-tail regions.

  • C.6. Distribution of SDG Datasets: The study evaluates synthetic-content complementarity by comparing SDG embeddings with clusters from the pre-training distribution.Fig. 36 visualizes the joint geometry of pre-training and SDG embeddings in the Cosmos-Embed1 video embedding space.
  • C.6. Distribution of SDG Datasets: Tab. 25 quantifies each SDG source’s distance from the pre-training manifold using an in-distribution self-reference as the baseline.
  • C.6. Distribution of SDG Datasets: The combined geometry and distance analyses show that SDG occupies long-tail regions of the pre-training distribution.

C.7. Ablation Study: Impact of SDG Datasets

The ablation shows that individual SDG sources produce complementary domain-specific gains but also introduce trade-offs, especially a persistent decline in Human performance. Combining all sources yields broad, balanced improvements and the best Quality score.

  • Evaluation setup: Variants are fine-tuned on individual or joint SDG sources and evaluated with PAIBench-G T2V across Overall, Quality, and six domain-specific scores.The domains are Common Sense, AV, Robot, Industry, Human, and Physics.
  • Sim-to-real gap and domain trade-offs: Human declines in every individual SDG variant, ranging from −0.38 for SDG-SynHuman to −0.69 for SDG-PhyxSim.Even the human-centric SDG-SynHuman source does not recover the Human score, indicating a pronounced sim-to-real gap for human-related content.
  • Sim-to-real gap and domain trade-offs: SDG-RobotSim hurts AV by −1.03 and Industry by −1.22, showing that domain-specialized sources can degrade orthogonal categories.These trade-offs accompany the source-specific gains and reflect differing simulator content distributions.
  • Combined SDG-All: SDG-All improves eight of nine metrics, retains a Human dip of −0.47, and achieves the best Quality score of 72.56.The joint dataset produces consistent gains across the remaining domain categories, suggesting that diversity suppresses individual source biases.

D. Cosmos3-Edge LLM Model Training · E. Additional Ablation Study

Cosmos3-Edge is a dense 2B model trained through staged pre-training, long-context extension, and supervised fine-tuning, while its evaluation shows strengths in mathematical and scientific reasoning. Additional ablations examine architectural, data, and training choices underlying Cosmos 3’s unified omnimodal capabilities for Physical AI.

  • D. Cosmos3-Edge LLM Model Training: Cosmos3-Edge uses a dense 2B backbone trained from scratch with base pre-training, long-context extension, and supervised fine-tuning stages.Training uses BF16 precision throughout.
  • D. Cosmos3-Edge LLM Model Training: 15T tokens train the 2B backbone during base pre-training with an 8,192-token sequence length and general-to-higher-quality data mixtures.The data mixture is hot-swapped by resuming continued pre-training from the general-pre-training checkpoint.
  • D. Cosmos3-Edge LLM Model Training: 26M SFT samples are packed into roughly 2.6M sequences of up to 128K tokens for a single supervised fine-tuning stage.The data spans mathematics, coding, science, general chat, instruction following, tool use, and code-agent tasks.
  • D. Cosmos3-Edge LLM Model Training: Evaluation covers reasoning, science, instruction following, long context, and general capabilities against the same-size Qwen3.5-2B baseline.Benchmarks include HMMT25 Feb, GPQA, MMLU-Pro, AA-LCR, IFBench, and Scale AI Multi-Challenge.
  • D. Cosmos3-Edge LLM Model Training: Cosmos3-Edge substantially outperforms Qwen3.5-2B on HMMT25 Feb and GPQA, matches it comparably on IFBench and AA-LCR, and trails it on MMLU-Pro.The reported strengths are in math reasoning and science, while the weakness is in general-domain evaluation.
  • E. Additional Ablation Study: The ablation studies investigate Reasoner–Generator interactions and the effects of architectural, data, and training decisions within the Mixture-of-Transformers framework.They also analyze world and action representations across domains to clarify design principles for a unified omnimodal world model.

E.1. How the Reasoner Benefits the Generator · E.2. Choice of FPS Control

The Cosmos3-Nano Reasoner improves the Generator’s Physical AI domain performance over Qwen3-VL-8B, while preserving comparable quality. For FPS control, combining MRoPE modulation with text conditioning yields the strongest composite score, with gains concentrated in motion fidelity rather than perceptual quality.

  • E.1. How the Reasoner Benefits the Generator: The ablation trains Cosmos3-Nano models with either Qwen3-VL-8B or Cosmos3-Nano Reasoner as the understanding tower, while training the Generator from scratch.Training uses 256p and 480p resolutions, 0–200-frame clips, a 25K sequence length, and joint image–video training.
  • E.1. How the Reasoner Benefits the Generator: On T2V, replacing Qwen3-VL-8B with the Reasoner raises the overall Domain score from 73.7 to 75.7, including a +4.8 gain for Robot.Other reported gains include Physics (+0.5, 88.7 →89.2) and AV (+2.3, 52.6 →54.9).
  • E.1. How the Reasoner Benefits the Generator: On I2V, the Reasoner raises the Domain score from 80.0 to 80.8, driven by gains in Common sense, Industry, Human, and Physics.The reported gains are +0.7, +0.8, +1.0, and +0.2, respectively.
  • E.1. How the Reasoner Benefits the Generator: The results suggest that the Reasoner provides better Physical AI embeddings for Generator learning, while the two models achieve comparable quality scores.The comparison concerns the understanding-tower initialization; the Generator tower is trained from scratch in both variants.
  • E.2. Choice of FPS Control: FPS conditioning compares Text Control, MRoPE FPS Modulation, and their combination using videos from 10, 15, 24, and 30 FPS bands, generated at 480p for 5 s across three seeds.The evaluation set contains approximately 100 videos per FPS band with ±2 FPS tolerance.
  • E.2. Choice of FPS Control: The evaluation scores Video Quality and Dynamic Degree, then combines motion fidelity with perceptual quality into a Composite Score averaged across FPS bands.Composite Score is computed as E_FPS-band[(VQ) · MF], with DD measured on a 0–1 scale.
  • E.2. Choice of FPS Control: +1.30 over Base is the best composite gain, achieved by combining Text Control and MRoPE FPS Modulation; MRoPE alone gains +1.12 versus +0.77 for Text Control.VQ remains within a 0.2-point window across settings, indicating that improvements primarily target motion fidelity and temporal behavior.

E.3. Audio Data in Pre-Training · E.4. Synergy Between Action Modes · E.5. Video-Action Consistency

Cosmos 3’s ablations show that joint video-audio pre-training benefits video metrics, joint action-mode training improves inverse-dynamics accuracy and policy coverage with a modest FD tradeoff, and predicted videos closely match simulator rollouts. These findings support cross-modal and cross-action consistency in Cosmos 3 models.

  • E.3. Audio Data in Pre-Training: Two Cosmos3-Nano variants were continued-pretrained for 20k iterations on 128 GPUs using either video-only or joint video-audio data at 256p and 480p.The variants started from the same pre-trained checkpoint and shared the video data.
  • E.3. Audio Data in Pre-Training: Continued pre-training without audio reduced both PAIBench T2V and I2V scores, indicating that joint video-audio training does not degrade video generation quality.The results suggest a modest benefit from audio-inclusive pre-training even on video-centric metrics.
  • E.4. Synergy Between Action Modes: Cosmos 3 evaluates shared structure across forward dynamics, inverse dynamics, and joint video-action prediction using single-mode and joint checkpoints on PushT with Cosmos3-Edge.Three single-mode checkpoints were trained for 2K steps each, while one joint FD/ID/policy checkpoint was trained for 6K steps.
  • E.4. Synergy Between Action Modes: 72% relative reduction: ID MSE decreased from 1.11×10−3 to 3.09×10−4, while policy coverage increased from 74.1% to 77.3% with the joint checkpoint.Lower is better for ID MSE, whereas higher is better for policy coverage.
  • E.4. Synergy Between Action Modes: FD PSNR decreased from 27.13 to 26.22, indicating a modest tradeoff in forward-dynamics reconstruction fidelity under joint action-mode training.Despite this tradeoff, the joint checkpoint achieved the best policy coverage and ID accuracy under the same per-mode optimization budget.
  • E.5. Video-Action Consistency: Cosmos3-Nano-Policy-DROID’s predicted video and action streams were evaluated for consistency on held-out RoboLab by comparing predicted videos with simulator rollouts from identical initial states.For each predicted action chunk, the same chunk was executed in the RoboLab simulator and PSNR was computed between the predicted video and resulting rollout.
  • E.5. Video-Action Consistency: The predicted video closely matches the simulator rollout for both the wrist and left cameras.Figure 37 compares simulator-enacted video with video predicted jointly with the action chunk.

F. Cosmos-HumanEval Benchmark (Cosmos-HUE) … Other Contributions

Cosmos-HumanEval defines a human-referenced binary benchmark for video generation, with reliability-controlled annotation and strong Cosmos3 results across T2V and I2V. The appendix also documents contributors across model, data, training, infrastructure, benchmark, and other contribution groups.

  • F. Cosmos-HumanEval Benchmark (Cosmos-HUE): Cosmos-HUE evaluates video generation using VLM-generated atomic binary questions across Semantic Alignment, Physical Laws, Geometric Reasoning, and Visual Integrity.The benchmark supplies a human-reference signal and formalizes its annotation and scoring procedures.
  • F. Cosmos-HumanEval Benchmark (Cosmos-HUE): Unclear responses are scored as No, preventing models from being rewarded when prompted events are absent, occluded, or too difficult to assess.Each Layer 3 question accepts Yes, No, or Unclear under the stated response schema.
  • F. Cosmos-HumanEval Benchmark (Cosmos-HUE): Each video receives up to 16 questions, with two independent ratings and third-reviewer adjudication for disagreements.The workflow records one canonical answer per video-question pair and limits any single annotator’s influence.
  • F. Cosmos-HumanEval Benchmark (Cosmos-HUE): Up to 10,000 binary observations per checkpoint produce confidence-interval widths around ±0.6 points for top-tier scores, while real-video ground truth scores 93.6 on T2V and 94.4 on I2V prompts.The evaluation pool uses 100 PAIBench-G prompts, five random-seed generations per prompt, and up to 20 questions per video.
  • F. Cosmos-HumanEval Benchmark (Cosmos-HUE): 89.3 makes Cosmos3-Super the best open-source T2V generator, while Veo-3.1 leads overall at 91.3 and Seedance-1.5-Pro follows at 90.0.Cosmos3-Super is best open source on 9 of 12 axes and leads all generators on AV at 87.7 and Physics at 91.5.
  • F. Cosmos-HumanEval Benchmark (Cosmos-HUE): 89.6 places Cosmos3-Super 0.1 points behind Veo-3.1’s 89.7 overall on I2V, while Cosmos3-Super leads Visual Integrity at 94.2, Robotics at 91.1, and Miscellaneous at 94.8.Cosmos3-Super ties Veo-3.1 on Semantic Alignment at 90.3; Cosmos3-Nano wins AV at 87.6 and Wan2.2-A14B wins Physics at 91.9.
  • G. Contributors and Acknowledgments: Contributors are listed alphabetically by last name across supervision, model architecture, reasoner and generator data, training recipe, infrastructure, and results-and-benchmarks groups.The listed groups cover reasoner pre-training and post-training data, image/video/audio/action/transfer generator data, and benchmark contributions.
  • Other Contributions: The acknowledgments additionally identify contributors for image and video data, audio and action generation, transfer, training, infrastructure, results and benchmarks, and other contributions.These groups include deduplication and filtering, infrastructure, results, and the Other Contributions roster.
Loading 2606.02800v4…