X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/

π0: AVision-Language-Action…π0: A Vision-Language-Action Flow Model for General Robot ControlMemoryVLA:Perceptual-Cognitive…MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic ManipulationUnifiedVision-Language-Action…Unified Vision-Language-Action ModelKnowledge InsulatingVision-Language-Action…Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize BetterGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsThinkAct:Vision-Language-Action…ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningDiscrete Diffusion VLA:Bringing Discrete…Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action PoliciesFine-TuningVision-Language-Action…Fine-Tuning Vision-Language-Action Models: Optimizing Speed and SuccessFAST: Efficient ActionTokenization for…FAST: Efficient Action Tokenization for Vision-Language-Action ModelsRoboTwin 2.0: A ScalableData Generator and…RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationSpatialVLA: ExploringSpatial Representations…SpatialVLA: Exploring Spatial Representations for Visual-Language-Action ModelTraceVLA: Visual TracePrompting Enhances…TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic PoliciesRynnVLA-002: A UnifiedVision-Language-Action…RynnVLA-002: A Unified Vision-Language-Action and World ModelMixture of Horizons inAction ChunkingMixture of Horizons in Action ChunkingUnifying Perception andAction: A…Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationStarVLA: A Lego-likeCodebase for…StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developingπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesCausal World Modelingfor Robot ControlCausal World Modeling for Robot ControlBeing-H0.5: ScalingHuman-Centric Robot…Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment GeneralizationLAP: Language-ActionPre-Training Enables…LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment TransferHoloBrain-0 TechnicalReportHoloBrain-0 Technical ReportInternVLA-A1: UnifyingUnderstanding…InternVLA-A1: Unifying Understanding, Generation and Action for Robotic ManipulationOA-WAM:Object-Addressable Worl…OA-WAM: Object-Addressable World Action Model for Robust Robot ManipulationStarVLA-α: ReducingComplexity in…StarVLA-α: Reducing Complexity in Vision-Language-Action SystemsX-VLA: Soft-PromptedTransformer as Scalable…X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.