Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Recent video generation models demonstrate remarkable ability to capture complex physical interactions and scene evolution over time. To leverage their spatiotemporal priors, robotics works have adapted video models for policy learning but introduce complexity by requiring multiple stages of post-training and new architectural components for action generation. In this work, we introduce Cosmos Policy, a simple approach for adapting a large pretrained video model (Cosmos-Predict2) into an effective robot policy through a single stage of post-training on the robot demonstration data collected on the target platform, with no architectural modifications. Cosmos Policy learns to directly generate robot actions encoded as latent frames within the video model's latent diffusion process, harnessing the model's pretrained priors and core learning algorithm to capture complex action distributions. Additionally, Cosmos Policy generates future state images and values (expected cumulative rewards), which are similarly encoded as latent frames, enabling test-time planning of action trajectories with higher likelihood of success. In our evaluations, Cosmos Policy achieves state-of-the-art performance on the LIBERO and RoboCasa simulation benchmarks (98.5% and 67.1% average success rates, respectively) and the highest average score in challenging real-world bimanual manipulation tasks, outperforming strong diffusion policies trained from scratch, video model-based policies, and state-of-the-art vision-language-action models fine-tuned on the same robot demonstrations. Furthermore, given policy rollout data, Cosmos Policy can learn from experience to refine its world model and value function and leverage model-based planning to achieve even higher success rates in challenging tasks. We release code, models, and training data at https://research.nvidia.com/labs/dir/cosmos-policy/

Genie Envisioner: AUnified World Foundatio…Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationVideo Generators areRobot PoliciesVideo Generators are Robot PoliciesVidar: Embodied VideoDiffusion Model for…Vidar: Embodied Video Diffusion Model for Generalist Bimanual ManipulationUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsFLARE: Robot Learningwith Implicit World…FLARE: Robot Learning with Implicit World ModelingVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsDreamGen: UnlockingGeneralization in Robot…DreamGen: Unlocking Generalization in Robot Learning through Neural TrajectoriesDual-Stream Diffusionfor World-Model…Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelUnified Video ActionModelUnified Video Action ModelFlowVLA: Thinking inMotion with a Visual…FlowVLA: Thinking in Motion with a Visual Chain of ThoughtCogVLA:Cognition-Aligned…CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & SparsificationGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsDual-Stream Diffusionfor World-Model…Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?DiT4DiT: JointlyModeling Video Dynamics…DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot ControlWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot PoliciesCausal World Modelingfor Robot ControlCausal World Modeling for Robot ControlStarVLA: A Lego-likeCodebase for…StarVLA: A Lego-like Codebase for Vision-Language-Action Model DevelopingBeing-H0.7: A LatentWorld-Action Model from…Being-H0.7: A Latent World-Action Model from Egocentric VideosVAG: Dual-StreamVideo-Action Generation…VAG: Dual-Stream Video-Action Generation for Embodied Data SynthesisWALL-WM: Carving WorldAction Modeling at the…WALL-WM: Carving World Action Modeling at the Event JointsGigaWorld-Policy: AnEfficient…GigaWorld-Policy: An Efficient Action-Centered World-Action ModelInteractive WorldSimulator for Robot…Interactive World Simulator for Robot Policy Training and EvaluationGigaBrain-0.5M*: a VLAThat Learns From World…GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement LearningCosmos Policy:Fine-Tuning Video Model…Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.