World Action Models are Zero-shot Policies

State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained video diffusion backbone. Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions, using video as a dense representation of how the world evolves. By jointly modeling video and action, DreamZero learns diverse skills effectively from heterogeneous robot data without relying on repetitive demonstrations. This results in over 2x improvement in generalization to new tasks and environments compared to state-of-the-art VLAs in real robot experiments. Crucially, through model and system optimizations, we enable a 14B autoregressive video diffusion model to perform real-time closed-loop control at 7Hz. Finally, we demonstrate two forms of cross-embodiment transfer: video-only demonstrations from other robots or humans yield a relative improvement of over 42% on unseen task performance with just 10-20 minutes of data. More surprisingly, DreamZero enables few-shot embodiment adaptation, transferring to a new embodiment with only 30 minutes of play data while retaining zero-shot generalization.

mimic-video:Video-Action Models for…mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAsVideo Generators areRobot PoliciesVideo Generators are Robot PoliciesLarge Video PlannerEnables Generalizable…Large Video Planner Enables Generalizable Robot ControlUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsFLARE: Robot Learningwith Implicit World…FLARE: Robot Learning with Implicit World ModelingGenie Envisioner: AUnified World Foundatio…Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsV-JEPA 2:Self-Supervised Video…V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsDual-Stream Diffusionfor World-Model…Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelCosmos Policy:Fine-Tuning Video Model…Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningCausal World Modelingfor Robot ControlCausal World Modeling for Robot ControlFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?GigaWorld-Policy: AnEfficient…GigaWorld-Policy: An Efficient Action-Centered World-Action ModelStarVLA: A Lego-likeCodebase for…StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developingπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesInteractive WorldSimulator for Robot…Interactive World Simulator for Robot Policy Training and EvaluationWALL-WM: Carving WorldAction Modeling at the…WALL-WM: Carving World Action Modeling at the Event JointsAction Images:End-to-End Policy…Action Images: End-to-End Policy Learning via Multiview Video GenerationOA-WAM:Object-Addressable Worl…OA-WAM: Object-Addressable World Action Model for Robust Robot ManipulationFrom Imagined Futures toExecutable Actions…From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationDo World Action ModelsGeneralize Better than…Do World Action Models Generalize Better than VLAs? A Robustness StudyVAG: Dual-StreamVideo-Action Generation…VAG: Dual-Stream Video-Action Generation for Embodied Data SynthesisWorld Model for RobotLearning: A…World Model for Robot Learning: A Comprehensive SurveyWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot Policies過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。