World Simulation with Video Foundation Models for Physical AI

We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2World, Image2World, and Video2World generation in a single model and leverages [Cosmos-Reason1], a Physical AI vision-language model, to provide richer text grounding and finer control of world simulation. Trained on 200M curated video clips and refined with reinforcement learning-based post-training, [Cosmos-Predict2.5] achieves substantial improvements over [Cosmos-Predict1] in video quality and instruction alignment, with models released at 2B and 14B scales. These capabilities enable more reliable synthetic data generation, policy evaluation, and closed-loop simulation for robotics and autonomous systems. We further extend the family with [Cosmos-Transfer2.5], a control-net style framework for Sim2Real and Real2Real world translation. Despite being 3.5$\times$ smaller than [Cosmos-Transfer1], it delivers higher fidelity and robust long-horizon video generation. Together, these advances establish [Cosmos-Predict2.5] and [Cosmos-Transfer2.5] as versatile tools for scaling embodied intelligence. To accelerate research and deployment in Physical AI, we release source code, pretrained checkpoints, and curated benchmarks under the NVIDIA Open Model License at https://github.com/nvidia-cosmos/cosmos-predict2.5 and https://github.com/nvidia-cosmos/cosmos-transfer2.5. We hope these open resources lower the barrier to adoption and foster innovation in building the next generation of embodied intelligence.

LTX-Video: RealtimeVideo Latent DiffusionLTX-Video: Realtime Video Latent DiffusionGenie Envisioner: AUnified World Foundatio…Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationPAI-Bench: AComprehensive Benchmark…PAI-Bench: A Comprehensive Benchmark For Physical AICosmos-Transfer1:Conditional World…Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal ControlDreamGen: UnlockingGeneralization in Robot…DreamGen: Unlocking Generalization in Robot Learning through Neural TrajectoriesVideoPhy-2: AChallenging…VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationCosmos World FoundationModel Platform for…Cosmos World Foundation Model Platform for Physical AICosmos-Drive-Dreams:Scalable Synthetic…Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation ModelsV-JEPA 2:Self-Supervised Video…V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningWan: Open and AdvancedLarge-Scale Video…Wan: Open and Advanced Large-Scale Video Generative ModelsEWMBench: EvaluatingScene, Motion, and…EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World ModelsIRASim: A Fine-GrainedWorld Model for Robot…IRASim: A Fine-Grained World Model for Robot ManipulationGigaWorld-0: WorldModels as Data Engine t…GigaWorld-0: World Models as Data Engine to Empower Embodied AIPAI-Bench: AComprehensive Benchmark…PAI-Bench: A Comprehensive Benchmark For Physical AIDriveLaW:UnifyingPlanning and Video…DriveLaW:Unifying Planning and Video Generation in a Latent Driving Worldmimic-video:Video-Action Models for…mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAsAdvancing Open-sourceWorld ModelsAdvancing Open-source World ModelsRethinking VideoGeneration Model for th…Rethinking Video Generation Model for the Embodied WorldCosmos 3: OmnimodalWorld Models for…Cosmos 3: Omnimodal World Models for Physical AISelf-Refining VideoSamplingSelf-Refining Video SamplingWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot PoliciesDreamDojo: A GeneralistRobot World Model from…DreamDojo: A Generalist Robot World Model from Large-Scale Human VideosRISE: Self-ImprovingRobot Policy with…RISE: Self-Improving Robot Policy with Compositional World ModelDiT4DiT: JointlyModeling Video Dynamics…DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot ControlWorld Simulation withVideo Foundation Models…World Simulation with Video Foundation Models for Physical AI過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。