Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Imitation learning has emerged as a promising approach towards building generalist robots. However, scaling imitation learning for large robot foundation models remains challenging due to its reliance on high-quality expert demonstrations. Meanwhile, large amounts of video data depicting a wide range of environments and diverse behaviors are readily available. This data provides a rich source of information about real-world dynamics and agent-environment interactions. Leveraging this data directly for imitation learning, however, has proven difficult due to the lack of action annotation. In this work, we present Unified World Models (UWM), a framework that allows for leveraging both video and action data for policy learning. Specifically, a UWM integrates an action diffusion process and a video diffusion process within a unified transformer architecture, where independent diffusion timesteps govern each modality. By controlling each diffusion timestep, UWM can flexibly represent a policy, a forward dynamics, an inverse dynamics, and a video generator. Through simulated and real-world experiments, we show that: (1) UWM enables effective pretraining on large-scale multitask robot datasets with both dynamics and action predictions, resulting in more generalizable and robust policies than imitation learning, (2) UWM naturally facilitates learning from action-free video data through independent control of modality-specific diffusion timesteps, further improving the performance of finetuned policies. Our results suggest that UWM offers a promising step toward harnessing large, heterogeneous datasets for scalable robot learning, and provides a simple unification between the often disparate paradigms of imitation learning and world modeling. Videos and code are available at https://weirdlabuw.github.io/uwm/.

ImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseBehavior Transformers:Cloning k modes with on…Behavior Transformers: Cloning k modes with one stoneVideo Diffusion ModelsVideo Diffusion ModelsOpen X-Embodiment:Robotic Learning…Open X-Embodiment: Robotic Learning Datasets and RT-X ModelsStable Video Diffusion:Scaling Latent Video…Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsDROID: A Large-ScaleIn-The-Wild Robot…DROID: A Large-Scale In-The-Wild Robot Manipulation Datasetπ0: AVision-Language-Action…π0: A Vision-Language-Action Flow Model for General Robot ControlOpenVLA: An Open-SourceVision-Language-Action…OpenVLA: An Open-Source Vision-Language-Action ModelUnleashing Large-ScaleVideo Generative…Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsLatent ActionPretraining from VideosLatent Action Pretraining from VideosCosmos World FoundationModel Platform for…Cosmos World Foundation Model Platform for Physical AIMotus: A Unified LatentAction World ModelMotus: A Unified Latent Action World Modelmimic-video:Video-Action Models for…mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAsF1: AVision-Language-Action…F1: A Vision-Language-Action Model Bridging Understanding and Generation to ActionsCtrl-World: AControllable Generative…Ctrl-World: A Controllable Generative World Model for Robot ManipulationV-JEPA 2:Self-Supervised Video…V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningDreamVLA: AVision-Language-Action…DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?Cosmos Policy:Fine-Tuning Video Model…Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot PoliciesBeing-H0.7: A LatentWorld-Action Model from…Being-H0.7: A Latent World-Action Model from Egocentric VideosPALM: Progress-AwarePolicy Learning via…PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationOA-WAM:Object-Addressable Worl…OA-WAM: Object-Addressable World Action Model for Robust Robot ManipulationUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.