Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

If generalist robots are to operate in truly unstructured environments, they need to be able to recognize and reason about novel objects and scenarios. Such objects and scenarios might not be present in the robot's own training data. We propose SuSIE, a method that leverages an image-editing diffusion model to act as a high-level planner by proposing intermediate subgoals that a low-level controller can accomplish. Specifically, we finetune InstructPix2Pix on video data, consisting of both human videos and robot rollouts, such that it outputs hypothetical future "subgoal" observations given the robot's current observation and a language command. We also use the robot data to train a low-level goal-conditioned policy to act as the aforementioned low-level controller. We find that the high-level subgoal predictions can utilize Internet-scale pretraining and visual understanding to guide the low-level goal-conditioned policy, achieving significantly better generalization and precision than conventional language-conditioned policies. We achieve state-of-the-art results on the CALVIN benchmark, and also demonstrate robust generalization on real-world manipulation tasks, beating strong baselines that have access to privileged information or that utilize orders of magnitude more compute and training data. The project website can be found at http://rail-berkeley.github.io/susie .

Fixing Weight DecayRegularization in AdamFixing Weight Decay Regularization in AdamCACTI: A Framework forScalable Multi-Task…CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation LearningGenAug: Retargetingbehaviors to unseen…GenAug: Retargeting behaviors to unseen situations via Generative AugmentationLearning UniversalPolicies via Text-Guide…Learning Universal Policies via Text-Guided Video GenerationRT-2:Vision-Language-Action…RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlPaLM-E: An EmbodiedMultimodal Language…PaLM-E: An Embodied Multimodal Language ModelOpen-World ObjectManipulation using…Open-World Object Manipulation using Pre-Trained Vision-Language ModelsRT-1: RoboticsTransformer for…RT-1: Robotics Transformer for Real-World Control at ScaleCompositional FoundationModels for Hierarchical…Compositional Foundation Models for Hierarchical PlanningLIV: Language-ImageRepresentations and…LIV: Language-Image Representations and Rewards for Robotic ControlScaling Robot Learningwith Semantically…Scaling Robot Learning with Semantically Imagined ExperienceDiffusion Policy:Visuomotor Policy…Diffusion Policy: Visuomotor Policy Learning via Action Diffusion3D Diffuser Actor:Policy Diffusion with 3…3D Diffuser Actor: Policy Diffusion with 3D Scene RepresentationsAny-point TrajectoryModeling for Policy…Any-point Trajectory Modeling for Policy LearningOcto: An Open-SourceGeneralist Robot PolicyOcto: An Open-Source Generalist Robot PolicyPrediction with Action:Visual Policy Learning…Prediction with Action: Visual Policy Learning via Joint Denoising ProcessTinyVLA: Towards Fast,Data-Efficient…TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic ManipulationDiffusion TransformerPolicyDiffusion Transformer PolicyVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsCoT-VLA: VisualChain-of-Thought…CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsUP-VLA: A UnifiedUnderstanding and…UP-VLA: A Unified Understanding and Prediction Model for Embodied AgentUnifiedVision-Language-Action…Unified Vision-Language-Action ModelGenerative ArtificialIntelligence in Robotic…Generative Artificial Intelligence in Robotic Manipulation: A Surveyπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesZero-Shot RoboticManipulation with…Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.