GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.

RT-2:Vision-Language-Action…RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlPaLM-E: An EmbodiedMultimodal Language…PaLM-E: An Embodied Multimodal Language ModelRT-1: RoboticsTransformer for…RT-1: Robotics Transformer for Real-World Control at ScaleGR-2: A GenerativeVideo-Language-Action…GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulationπ0: AVision-Language-Action…π0: A Vision-Language-Action Flow Model for General Robot Control3D-VLA: A 3DVision-Language-Action…3D-VLA: A 3D Vision-Language-Action Generative World ModelOpenVLA: An Open-SourceVision-Language-Action…OpenVLA: An Open-Source Vision-Language-Action ModelUnleashing Large-ScaleVideo Generative…Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationGen2Act: Human VideoGeneration in Novel…Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot ManipulationVision-LanguageFoundation Models as…Vision-Language Foundation Models as Effective Robot ImitatorsTinyVLA: Towards Fast,Data-Efficient…TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic ManipulationAgiBot World Colosseo: ALarge-scale Manipulatio…AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsEmergence of Human toRobot Transfer in…Emergence of Human to Robot Transfer in Vision-Language-Action ModelsFast-in-Slow: ADual-System Foundation…Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow ReasoningUnifiedVision-Language-Action…Unified Vision-Language-Action ModelThinkAct:Vision-Language-Action…ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningπRL: Online RLFine-tuning for…πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action ModelsORV: 4DOccupancy-centric Robot…ORV: 4D Occupancy-centric Robot Video GenerationHiF-VLA: Hindsight,Insight and Foresight…HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action ModelsCO-RFT: EfficientFine-Tuning of…CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learningℰ0: EnhancingGeneralization and…ℰ0: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete DiffusionWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot Policiesπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesLearning Physics fromPretrained Video Models…Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic ManipulationGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.