DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning. While vision-language-action (VLA) models have shown promise for generalizable robot skills, realizing their full potential requires addressing limitations in action representation and efficient training. Current VLA models often focus on scaling the vision-language model (VLM) component, while the action space representation remains a critical bottleneck. This paper introduces DexVLA, a novel framework designed to enhance the efficiency and generalization capabilities of VLAs for complex, long-horizon tasks across diverse robot embodiments. DexVLA features a novel diffusion-based action expert, scaled to one billion parameters, designed for cross-embodiment learning. A novel embodiment curriculum learning strategy facilitates efficient training: (1) pre-training the diffusion expert that is separable from the VLA on cross-embodiment data, (2) aligning the VLA model to specific embodiments, and (3) post-training for rapid adaptation to new tasks. We conduct comprehensive experiments across multiple embodiments, including single-arm, bimanual, and dexterous hand, demonstrating DexVLA's adaptability to challenging tasks without task-specific adaptation, its ability to learn dexterous skills on novel embodiments with limited data, and its capacity to complete complex, long-horizon tasks using only direct language prompting, such as laundry folding. In all settings, our method demonstrates superior performance compared to state-of-the-art models like Octo, OpenVLA, and Diffusion Policy.

Diffusion-VLA: ScalingRobot Foundation Models…Diffusion-VLA: Scaling Robot Foundation Models via Unified Diffusion and AutoregressionTinyVLA: Towards Fast,Data-Efficient…TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation3D-VLA: A 3DVision-Language-Action…3D-VLA: A 3D Vision-Language-Action Generative World ModelConsistency Policy:Accelerated Visuomotor…Consistency Policy: Accelerated Visuomotor Policies via Consistency DistillationFine-TuningVision-Language-Action…Fine-Tuning Vision-Language-Action Models: Optimizing Speed and SuccessHybridVLA: CollaborativeDiffusion and…HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelFAST: Efficient ActionTokenization for…FAST: Efficient Action Tokenization for Vision-Language-Action Modelsπ0.5: aVision-Language-Action…π0.5: a Vision-Language-Action Model with Open-World GeneralizationCoT-VLA: VisualChain-of-Thought…CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsData Scaling Laws inImitation Learning for…Data Scaling Laws in Imitation Learning for Robotic ManipulationVLAS:Vision-Language-Action…VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot ManipulationChatVLA-2:Vision-Language-Action…ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained KnowledgePointVLA: Injecting the3D World into…PointVLA: Injecting the 3D World into Vision-Language-Action ModelsDita: Scaling DiffusionTransformer for…Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action PolicyObjectVLA: End-to-EndOpen-World Object…ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstrationπ0.5: aVision-Language-Action…π0.5: a Vision-Language-Action Model with Open-World GeneralizationdVLA: DiffusionVision-Language-Action…dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-ThoughtFast-in-Slow: ADual-System Foundation…Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow ReasoningForceVLA: Enhancing VLAModels with a…ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationOpenHelix: A ShortSurvey, Empirical…OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic ManipulationLoHoVLA: A UnifiedVision-Language-Action…LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied TasksMemoryVLA:Perceptual-Cognitive…MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic ManipulationEfficientVLA:Training-Free…EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsDexVLA: Vision-LanguageModel with Plug-In…DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。