Video Generators are Robot Policies

Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is constrained by the size of human demonstration data. In this paper, we use video generation as a proxy for robot policy learning to address both limitations simultaneously. We propose Video Policy, a modular framework that combines video and action generation that can be trained end-to-end. Our results demonstrate that learning to generate videos of robot behavior allows for the extraction of policies with minimal demonstration data, significantly improving robustness and sample efficiency. Our method shows strong generalization to unseen objects, backgrounds, and tasks, both in simulation and the real world. We further highlight that task success is closely tied to the generated video, with action-free video data providing critical benefits for generalizing to novel tasks. By leveraging large-scale video generative models, we achieve superior performance compared to traditional behavior cloning, paving the way for more scalable and data-efficient robot policy learning.

Video (language)modeling: a baseline fo…Video (language) modeling: a baseline for generative models of natural videosStochastic AdversarialVideo PredictionStochastic Adversarial Video PredictionCCNet: Extracting HighQuality Monolingual…CCNet: Extracting High Quality Monolingual Datasets from Web Crawl DataMasked VisualPre-training for Motor…Masked Visual Pre-training for Motor ControlStable Video Diffusion:Scaling Latent Video…Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsZero-Shot RoboticManipulation with…Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion ModelsI2VGen-XL: High-QualityImage-to-Video Synthesi…I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsGR-2: A GenerativeVideo-Language-Action…GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationVideo Language PlanningVideo Language PlanningVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsA Careful Examination ofLarge Behavior Models…A Careful Examination of Large Behavior Models for Multitask Dexterous ManipulationGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid Robotsmimic-video:Video-Action Models for…mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAsNovaFlow: Zero-ShotManipulation via…NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated VideosDual-Stream Diffusionfor World-Model…Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?Cosmos Policy:Fine-Tuning Video Model…Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningInteractive WorldSimulator for Robot…Interactive World Simulator for Robot Policy Training and EvaluationWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot Policiesπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesWorld Model for RobotLearning: A…World Model for Robot Learning: A Comprehensive SurveyFrom Imagined Futures toExecutable Actions…From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationDiT4DiT: JointlyModeling Video Dynamics…DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot ControlRethinking VideoGeneration Model for th…Rethinking Video Generation Model for the Embodied WorldVideo Generators areRobot PoliciesVideo Generators are Robot Policies過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。