EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we explore training Vision-Language-Action (VLA) models using egocentric human videos. The benefit of using human videos is not only for their scale but more importantly for the richness of scenes and tasks. With a VLA trained on human video that predicts human wrist and hand actions, we can perform Inverse Kinematics and retargeting to convert the human actions to robot actions. We fine-tune the model using a few robot manipulation demonstrations to obtain the robot policy, namely EgoVLA. We propose a simulation benchmark called Ego Humanoid Manipulation Benchmark, where we design diverse bimanual manipulation tasks with demonstrations. We fine-tune and evaluate EgoVLA with Ego Humanoid Manipulation Benchmark and show significant improvements over baselines and ablate the importance of human data. Videos can be found on our website: https://rchalyang.github.io/EgoVLA

On the Continuity ofRotation Representation…On the Continuity of Rotation Representations in Neural NetworksR3M: A Universal VisualRepresentation for Robo…R3M: A Universal Visual Representation for Robot ManipulationORBIT: A UnifiedSimulation Framework fo…ORBIT: A Unified Simulation Framework for Interactive Robot Learning EnvironmentsIntroducing HOT3D: AnEgocentric Dataset for…Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object TrackingLearning Manipulation byPredicting InteractionLearning Manipulation by Predicting InteractionEvaluating Real-WorldRobot Manipulation…Evaluating Real-World Robot Manipulation Policies in SimulationOpen-TeleVision:Teleoperation with…Open-TeleVision: Teleoperation with Immersive Active Visual FeedbackTACO: BenchmarkingGeneralizable Bimanual…TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingHumanoid Policy HumanPolicyHumanoid Policy Human PolicyLatent ActionPretraining from VideosLatent Action Pretraining from Videosπ0.5: aVision-Language-Action…π0.5: a Vision-Language-Action Model with Open-World GeneralizationSpatiotemporalPredictive Pre-training…Spatiotemporal Predictive Pre-training for Robotic Motor ControlScalableVision-Language-Action…Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity VideosEmergence of Human toRobot Transfer in…Emergence of Human to Robot Transfer in Vision-Language-Action ModelsIn-N-On: ScalingEgocentric Manipulation…In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task DataMasquerade: Learningfrom In-the-wild Human…Masquerade: Learning from In-the-wild Human Videos using Data-EditingMotionTrans: Human VRData Enable Motion-Leve…MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation PoliciesCoMo: LearningContinuous Latent Motio…CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot LearningEgoScale: ScalingDexterous Manipulation…EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human DataRobot Learning fromHuman Videos: A SurveyRobot Learning from Human Videos: A SurveyGazeVLA: Learning HumanIntention for Robotic…GazeVLA: Learning Human Intention for Robotic ManipulationUniDex: A RobotFoundation Suite for…UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human VideosEgoHumanoid: UnlockingIn-the-Wild…EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric DemonstrationActiveGlasses: LearningManipulation with Activ…ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human DemonstrationEgoVLA: LearningVision-Language-Action…EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。