GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

We present GR-2, a state-of-the-art generalist robot agent for versatile and generalizable robot manipulation. GR-2 is first pre-trained on a vast number of Internet videos to capture the dynamics of the world. This large-scale pre-training, involving 38 million video clips and over 50 billion tokens, equips GR-2 with the ability to generalize across a wide range of robotic tasks and environments during subsequent policy learning. Following this, GR-2 is fine-tuned for both video generation and action prediction using robot trajectories. It exhibits impressive multi-task learning capabilities, achieving an average success rate of 97.7% across more than 100 tasks. Moreover, GR-2 demonstrates exceptional generalization to new, previously unseen scenarios, including novel backgrounds, environments, objects, and tasks. Notably, GR-2 scales effectively with model size, underscoring its potential for continued growth and application. Project page: \url{https://gr2-manipulation.github.io}.

Masked VisualPre-training for Motor…Masked Visual Pre-training for Motor ControlR3M: A Universal VisualRepresentation for Robo…R3M: A Universal Visual Representation for Robot ManipulationZero-Shot RoboticManipulation with…Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion ModelsRT-2:Vision-Language-Action…RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlUnleashing Large-ScaleVideo Generative…Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationAny-point TrajectoryModeling for Policy…Any-point Trajectory Modeling for Policy LearningVision-LanguageFoundation Models as…Vision-Language Foundation Models as Effective Robot ImitatorsOcto: An Open-SourceGeneralist Robot PolicyOcto: An Open-Source Generalist Robot PolicyOpenVLA: An Open-SourceVision-Language-Action…OpenVLA: An Open-Source Vision-Language-Action ModelRoboAgent:Generalization and…RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking3D Diffuser Actor:Policy Diffusion with 3…3D Diffuser Actor: Policy Diffusion with 3D Scene RepresentationsLearning InteractiveReal-World SimulatorsLearning Interactive Real-World SimulatorsUnifiedVision-Language-Action…Unified Vision-Language-Action ModelWorldVLA: TowardsAutoregressive Action…WorldVLA: Towards Autoregressive Action World ModelSpatialVLA: ExploringSpatial Representations…SpatialVLA: Exploring Spatial Representations for Visual-Language-Action ModelCoT-VLA: VisualChain-of-Thought…CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsRoboTron-Mani:All-in-One Multimodal…RoboTron-Mani: All-in-One Multimodal Large Model for Robotic ManipulationEnerVerse: EnvisioningEmbodied Future Space…EnerVerse: Envisioning Embodied Future Space for Robotics ManipulationVideo Generators areRobot PoliciesVideo Generators are Robot PoliciesReal-Time Execution ofAction Chunking Flow…Real-Time Execution of Action Chunking Flow PoliciesScalableVision-Language-Action…Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity VideosA Step Toward WorldModels: A Survey on…A Step Toward World Models: A Survey on Robotic Manipulationπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot PoliciesGR-2: A GenerativeVideo-Language-Action…GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。