A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

Robot manipulation has seen tremendous progress in recent years, with imitation learning policies enabling successful performance of dexterous and hard-to-model tasks. Concurrently, scaling data and model size has led to the development of capable language and vision foundation models, motivating large-scale efforts to create general-purpose robot foundation models. While these models have garnered significant enthusiasm and investment, meaningful evaluation of real-world performance remains a challenge, limiting both the pace of development and inhibiting a nuanced understanding of current capabilities. In this paper, we rigorously evaluate multitask robot manipulation policies, referred to as Large Behavior Models (LBMs), by extending the Diffusion Policy paradigm across a corpus of simulated and real-world robot data. We propose and validate an evaluation pipeline to rigorously analyze the capabilities of these models with statistical confidence. We compare against single-task baselines through blind, randomized trials in a controlled setting, using both simulation and real-world experiments. We find that multi-task pretraining makes the policies more successful and robust, and enables teaching complex new tasks more quickly, using a fraction of the data when compared to single-task baselines. Moreover, performance predictably increases as pretraining scale and diversity grows. Project page: https://toyotaresearchinstitute.github.io/lbm1/

MimicGen: A DataGeneration System for…MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrationsπ0: AVision-Language-Action…π0: A Vision-Language-Action Flow Model for General Robot ControlOpenVLA: An Open-SourceVision-Language-Action…OpenVLA: An Open-Source Vision-Language-Action ModelRobot Learning as anEmpirical Science: Best…Robot Learning as an Empirical Science: Best Practices for Policy EvaluationEvaluating Real-WorldRobot Manipulation…Evaluating Real-World Robot Manipulation Policies in SimulationManiSkill3: GPUParallelized Robotics…ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AIAutoEval: AutonomousEvaluation of Generalis…AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real Worldπ0.5: aVision-Language-Action…π0.5: a Vision-Language-Action Model with Open-World GeneralizationGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsEmpirical Analysis ofSim-and-Real Cotraining…Empirical Analysis of Sim-and-Real Cotraining Of Diffusion Policies For Planar Pushing from PixelsRDT-1B: a DiffusionFoundation Model for…RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSim-and-RealCo-Training: A Simple…Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic ManipulationGR-3 Technical ReportGR-3 Technical ReportLarge Video PlannerEnables Generalizable…Large Video Planner Enables Generalizable Robot ControlReal-to-Sim Robot PolicyEvaluation with Gaussia…Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body InteractionsGR-RL: Going Dexterousand Precise for…GR-RL: Going Dexterous and Precise for Long-Horizon Robotic ManipulationReliable and ScalableRobot Policy Evaluation…Reliable and Scalable Robot Policy Evaluation with Imperfect SimulatorsWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot Policiesπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesInteractive WorldSimulator for Robot…Interactive World Simulator for Robot Policy Training and EvaluationA Systematic Study ofData Modalities and…A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot ManipulationA Mechanistic Analysisof Sim-and-Real…A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot PoliciesCaP-X: A Framework forBenchmarking and…CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot ManipulationMEM: Multi-ScaleEmbodied Memory for…MEM: Multi-Scale Embodied Memory for Vision Language Action ModelsA Careful Examination ofLarge Behavior Models…A Careful Examination of Large Behavior Models for Multitask Dexterous ManipulationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.