RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Simulation-based data synthesis has emerged as a powerful paradigm for advancing real-world robotic manipulation. Yet existing datasets remain insufficient for robust bimanual manipulation due to (1) the lack of scalable task generation methods and (2) oversimplified simulation environments. We present RoboTwin 2.0, a scalable framework for automated, large-scale generation of diverse and realistic data, together with unified evaluation protocols for dual-arm manipulation. At its core is RoboTwin-OD, an object library of 731 instances across 147 categories with semantic and manipulation-relevant annotations. Building on this, we design an expert data synthesis pipeline that leverages multimodal language models (MLLMs) and simulation-in-the-loop refinement to automatically generate task-level execution code. To improve sim-to-real transfer, RoboTwin 2.0 applies structured domain randomization along five axes: clutter, lighting, background, tabletop height, and language, enhancing data diversity and policy robustness. The framework is instantiated across 50 dual-arm tasks and five robot embodiments. Empirically, it yields a 10.9% gain in code generation success rate. For downstream policy learning, a VLA model trained with synthetic data plus only 10 real demonstrations achieves a 367% relative improvement over the 10-demo baseline, while zero-shot models trained solely on synthetic data obtain a 228% gain. These results highlight the effectiveness of RoboTwin 2.0 in strengthening sim-to-real transfer and robustness to environmental variations. We release the data generator, benchmark, dataset, and code to support scalable research in robust bimanual manipulation. Project Page: https://robotwin-platform.github.io/, Code: https://github.com/robotwin-Platform/robotwin/.

RoboMIND: Benchmark onMulti-embodiment…RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot ManipulationCogACT: A FoundationalVision-Language-Action…CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic ManipulationOcto: An Open-SourceGeneralist Robot PolicyOcto: An Open-Source Generalist Robot Policyπ0: AVision-Language-Action…π0: A Vision-Language-Action Flow Model for General Robot ControlMobile ALOHA: LearningBimanual Mobile…Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body TeleoperationRDT-1B: a DiffusionFoundation Model for…RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationGraspVLA: a GraspingFoundation Model…GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action DataBenchmarkingGeneralizable Bimanual…Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS WorkshopFine-TuningVision-Language-Action…Fine-Tuning Vision-Language-Action Models: Optimizing Speed and SuccessAgiBot World Colosseo: ALarge-scale Manipulatio…AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsRoboVerse: Towards aUnified Platform…RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot LearningDexVLA: Vision-LanguageModel with Plug-In…DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot ControlX-VLA: Soft-PromptedTransformer as Scalable…X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelBenchmarkingGeneralizable Bimanual…Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS WorkshopSimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningAnyPos: AutomatedTask-Agnostic Actions…AnyPos: Automated Task-Agnostic Actions for Bimanual ManipulationRLinf-VLA: A Unified andEfficient Framework for…RLinf-VLA: A Unified and Efficient Framework for VLA+RL TrainingMixture of Horizons inAction ChunkingMixture of Horizons in Action ChunkingDexbotic: Open-SourceVision-Language-Action…Dexbotic: Open-Source Vision-Language-Action ToolboxRobo-Dopamine: GeneralProcess Reward Modeling…Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic ManipulationGigaWorld-Policy: AnEfficient…GigaWorld-Policy: An Efficient Action-Centered World-Action ModelCausal World Modelingfor Robot ControlCausal World Modeling for Robot ControlUniVTAC: A UnifiedSimulation Platform for…UniVTAC: A Unified Simulation Platform for Visuo-Tactile Manipulation Data Generation, Learning, and BenchmarkingBiManiBench: AHierarchical Benchmark…BiManiBench: A Hierarchical Benchmark for Evaluating Bimanual Coordination of Multimodal Large Language ModelsRoboTwin 2.0: A ScalableData Generator and…RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.