Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information, often neglecting the dynamic aspects vital for embodied tasks. Recently, video diffusion models (VDMs) demonstrate the ability to predict future frames and showcase a strong understanding of physical world. We hypothesize that VDMs inherently produce visual representations that encompass both current static information and predicted future dynamics, thereby providing valuable guidance for robot action learning. Based on this hypothesis, we propose the Video Prediction Policy (VPP), which learns implicit inverse dynamics model conditioned on predicted future representations inside VDMs. To predict more precise future, we fine-tune pre-trained video foundation model on robot datasets along with internet human manipulation data. In experiments, VPP achieves a 18.6\% relative improvement on the Calvin ABC-D generalization benchmark compared to the previous state-of-the-art, and demonstrates a 31.6\% increase in success rates for complex real-world dexterous manipulation tasks. Project page at https://video-prediction-policy.github.io

Zero-Shot RoboticManipulation with…Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion ModelsRT-2:Vision-Language-Action…RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlVIP: Towards UniversalVisual Reward and…VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingGen2Act: Human VideoGeneration in Novel…Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot ManipulationVision-LanguageFoundation Models as…Vision-Language Foundation Models as Effective Robot ImitatorsUnleashing Large-ScaleVideo Generative…Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationIGOR: Image-GOalRepresentations are the…IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AIMultimodal DiffusionTransformer: Learning…Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal GoalsOcto: An Open-SourceGeneralist Robot PolicyOcto: An Open-Source Generalist Robot PolicyRoboUniView:Visual-Language Model…RoboUniView: Visual-Language Model with Unified View Representation for Robotic ManipulaitonLatent ActionPretraining from VideosLatent Action Pretraining from VideosImprovingVision-Language-Action…Improving Vision-Language-Action Model with Online Reinforcement LearningUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsCtrl-World: AControllable Generative…Ctrl-World: A Controllable Generative World Model for Robot ManipulationVideo Generators areRobot PoliciesVideo Generators are Robot PoliciesUnified Video ActionModelUnified Video Action ModelGR-3 Technical ReportGR-3 Technical Reportvilla-X: EnhancingLatent Action Modeling…villa-X: Enhancing Latent Action Modeling in Vision-Language-Action ModelsF1: AVision-Language-Action…F1: A Vision-Language-Action Model Bridging Understanding and Generation to ActionsUniCoD: Enhancing RobotPolicy via Unified…UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation LearningTriVLA: ATriple-System-Based…TriVLA: A Triple-System-Based Unified Vision-Language-Action Model for General Robot ControlFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?PALM: Progress-AwarePolicy Learning via…PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationVeo-Act: How Far CanFrontier Video Models…Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation?Video Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。