Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation can significantly benefit from large-scale video generative pre-training. We introduce GR-1, a straightforward GPT-style model designed for multi-task language-conditioned visual robot manipulation. GR-1 takes as inputs a language instruction, a sequence of observation images, and a sequence of robot states. It predicts robot actions as well as future images in an end-to-end manner. Thanks to a flexible design, GR-1 can be seamlessly finetuned on robot data after pre-trained on a large-scale video dataset. We perform extensive experiments on the challenging CALVIN benchmark and a real robot. On CALVIN benchmark, our method outperforms state-of-the-art baseline methods and improves the success rate from 88.9% to 94.9%. In the setting of zero-shot unseen scene generalization, GR-1 improves the success rate from 53.3% to 85.4%. In real robot experiments, GR-1 also outperforms baseline methods and shows strong potentials in generalization to unseen scenes and objects. We provide inaugural evidence that a unified GPT-style transformer, augmented with large-scale video generative pre-training, exhibits remarkable generalization to multi-task visual robot manipulation. Project page: https://GR1-Manipulation.github.io

Fixing Weight DecayRegularization in AdamFixing Weight Decay Regularization in AdamVIMA: General RobotManipulation with…VIMA: General Robot Manipulation with Multimodal PromptsDo As I Can, Not As ISay: Grounding Language…Do As I Can, Not As I Say: Grounding Language in Robotic AffordancesInner Monologue:Embodied Reasoning…Inner Monologue: Embodied Reasoning through Planning with Language ModelsA Generalist AgentA Generalist AgentLearning UniversalPolicies via Text-Guide…Learning Universal Policies via Text-Guided Video GenerationRT-2:Vision-Language-Action…RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlPaLM-E: An EmbodiedMultimodal Language…PaLM-E: An Embodied Multimodal Language ModelVIP: Towards UniversalVisual Reward and…VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingRoboCat: ASelf-Improving…RoboCat: A Self-Improving Foundation Agent for Robotic ManipulationRT-1: RoboticsTransformer for…RT-1: Robotics Transformer for Real-World Control at ScaleInteractive Language:Talking to Robots in…Interactive Language: Talking to Robots in Real TimeHiRT: Enhancing RoboticControl with…HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers3D Diffuser Actor:Policy Diffusion with 3…3D Diffuser Actor: Policy Diffusion with 3D Scene RepresentationsDiffusion TransformerPolicyDiffusion Transformer PolicyPrediction with Action:Visual Policy Learning…Prediction with Action: Visual Policy Learning via Joint Denoising ProcessVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsUP-VLA: A UnifiedUnderstanding and…UP-VLA: A Unified Understanding and Prediction Model for Embodied AgentFast-in-Slow: ADual-System Foundation…Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow ReasoningUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsWorldVLA: TowardsAutoregressive Action…WorldVLA: Towards Autoregressive Action World ModelFrom Multimodal LLMs toGeneralist Embodied…From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsCoT-VLA: VisualChain-of-Thought…CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsPALM: Progress-AwarePolicy Learning via…PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationUnleashing Large-ScaleVideo Generative…Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。