Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel scenarios and tasks, primarily due to their reliance on synthetic data with significant sim-to-real gaps or teleoperated demonstrations lacking scale and diversity. To address this data bottleneck, we propose leveraging human hands as a foundation manipulator, capitalizing on the rich dexterity and scalability present in web data. Our approach centers on physical instruction tuning, a novel training paradigm that combines large-scale VLA pretraining from human videos, physical space alignment for 3D reasoning, and post-training adaptation for robotic tasks. Additionally, we introduce a part-level motion tokenization method which achieves millimeter-level reconstruction accuracy to model precise hand trajectories for action learning. To support our proposed paradigm, we further develop a comprehensive data curation pipeline that integrates heterogeneous sources -- including motion capture, VR, and RGB-only videos -- into a large-scale dataset with millions of motion-based instructional instances. We empirically show the excellence of Being-H0 in hand motion generation and instruction following, and it also scales well with model and data sizes. Importantly, we observe the expected gains of Being-H0 in real-world robotic manipulation as physical instruction tuning is applied. More details are available at https://beingbeyond.github.io/Being-H0.

OpenVLA: An Open-SourceVision-Language-Action…OpenVLA: An Open-Source Vision-Language-Action ModelQwen2-VL: EnhancingVision-Language Model's…Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any ResolutionEgoDex: LearningDexterous Manipulation…EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric VideoHumanoid Policy HumanPolicyHumanoid Policy Human PolicyGraspVLA: a GraspingFoundation Model…GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action DataHuman2LocoMan: LearningVersatile Quadrupedal…Human2LocoMan: Learning Versatile Quadrupedal Manipulation with Human PretrainingBeing-0: A HumanoidRobotic Agent with…Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular SkillsEgoMimic: ScalingImitation Learning via…EgoMimic: Scaling Imitation Learning via Egocentric VideoPhantom: Training RobotsWithout Robots Using…Phantom: Training Robots Without Robots Using Only Human VideosAgiBot World Colosseo: ALarge-scale Manipulatio…AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsRDT-1B: a DiffusionFoundation Model for…RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationDexGraspVLA: AVision-Language-Action…DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingSpatial-Aware VLAPretraining through…Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human VideosScalableVision-Language-Action…Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity VideosMotionTrans: Human VRData Enable Motion-Leve…MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation PoliciesIn-N-On: ScalingEgocentric Manipulation…In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task DataGR-Dexter TechnicalReportGR-Dexter Technical ReportDexMan: LearningBimanual Dexterous…DexMan: Learning Bimanual Dexterous Manipulation from Human and Generated VideosBeing-H0.5: ScalingHuman-Centric Robot…Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment GeneralizationA Systematic Study ofData Modalities and…A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot ManipulationBeing-H0.7: A LatentWorld-Action Model from…Being-H0.7: A Latent World-Action Model from Egocentric VideosVLA-Adapter: AnEffective Paradigm for…VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action ModelRDT2: Exploring theScaling Limit of UMI…RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment GeneralizationJoint-Aligned LatentAction: Towards Scalabl…Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the WildBeing-H0:Vision-Language-Action…Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。