Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

How can robot manipulation policies generalize to novel tasks involving unseen object types and new motions? In this paper, we provide a solution in terms of predicting motion information from web data through human video generation and conditioning a robot policy on the generated video. Instead of attempting to scale robot data collection which is expensive, we show how we can leverage video generation models trained on easily available web data, for enabling generalization. Our approach Gen2Act casts language-conditioned manipulation as zero-shot human video generation followed by execution with a single policy conditioned on the generated video. To train the policy, we use an order of magnitude less robot interaction data compared to what the video prediction model was trained on. Gen2Act doesn't require fine-tuning the video model at all and we directly use a pre-trained model for generating human videos. Our results on diverse real-world scenarios show how Gen2Act enables manipulating unseen object types and performing novel motions for tasks not present in the robot data. Videos are at https://homangab.github.io/gen2act/

CACTI: A Framework forScalable Multi-Task…CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation LearningZero-Shot RoboticManipulation with…Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion ModelsGenAug: Retargetingbehaviors to unseen…GenAug: Retargeting behaviors to unseen situations via Generative AugmentationVIP: Towards UniversalVisual Reward and…VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingMimicPlay: Long-HorizonImitation Learning by…MimicPlay: Long-Horizon Imitation Learning by Watching Human PlayWhere are we in thesearch for an Artificia…Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Track2Act: PredictingPoint Tracks from…Track2Act: Predicting Point Tracks from Internet Videos enables Diverse Zero-shot Robot ManipulationDreamitate: Real-WorldVisuomotor Policy…Dreamitate: Real-World Visuomotor Policy Learning via Video GenerationAny-point TrajectoryModeling for Policy…Any-point Trajectory Modeling for Policy LearningUnleashing Large-ScaleVideo Generative…Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationGeneral Flow asFoundation Affordance…General Flow as Foundation Affordance for Scalable Robot LearningSemanticallyControllable…Semantically Controllable Augmentations for Generalizable Robot LearningScalableVision-Language-Action…Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity VideosPhantom: Training RobotsWithout Robots Using…Phantom: Training Robots Without Robots Using Only Human VideosPoint Policy: UnifyingObservations and Action…Point Policy: Unifying Observations and Actions with Key Points for Robot ManipulationVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsDreamGen: UnlockingGeneralization in Robot…DreamGen: Unlocking Generalization in Robot Learning through Neural TrajectoriesRobotic Manipulation byImitating Generated…Robotic Manipulation by Imitating Generated Videos Without Physical DemonstrationsMotionTrans: Human VRData Enable Motion-Leve…MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation PoliciesGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsDream2Flow: BridgingVideo Generation and…Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object FlowDreamVLA: AVision-Language-Action…DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeGazeVLA: Learning HumanIntention for Robotic…GazeVLA: Learning Human Intention for Robotic ManipulationPALM: Progress-AwarePolicy Learning via…PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationGen2Act: Human VideoGeneration in Novel…Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。