LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Visual-Language-Action (VLA) models report impressive success rates on robotic manipulation benchmarks, yet these results may mask fundamental weaknesses in robustness. We perform a systematic vulnerability analysis by introducing controlled perturbations across seven dimensions: objects layout, camera viewpoints, robot initial states, language instructions, light conditions, background textures and sensor noise. We comprehensively analyzed multiple state-of-the-art models and revealed consistent brittleness beneath apparent competence. Our analysis exposes critical weaknesses: models exhibit extreme sensitivity to perturbation factors, including camera viewpoints and robot initial states, with performance dropping from 95% to below 30% under modest perturbations. Surprisingly, models are largely insensitive to language variations, with further experiments revealing that models tend to ignore language instructions completely. Our findings challenge the assumption that high benchmark scores equate to true competency and highlight the need for evaluation practices that assess reliability under realistic variation.

Evaluating Real-WorldRobot Manipulation…Evaluating Real-World Robot Manipulation Policies in SimulationCogACT: A FoundationalVision-Language-Action…CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic ManipulationInteractivePost-Training for…Interactive Post-Training for Vision-Language-Action ModelsNORA: A SmallOpen-Sourced Generalist…NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied TasksWorldVLA: TowardsAutoregressive Action…WorldVLA: Towards Autoregressive Action World ModelUniVLA: Learning to ActAnywhere with…UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsVLABench: A Large-ScaleBenchmark for…VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsFrom Intention toExecution: Probing the…From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action ModelsWhat Can RL Bring to VLAGeneralization? An…What Can RL Bring to VLA Generalization? An Empirical StudyExploring the Limits ofVision-Language-Action…Exploring the Limits of Vision-Language-Action Manipulations in Cross-task GeneralizationVLA-RL: TowardsMasterful and General…VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningVLA-Arena: AnOpen-Source Framework…VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action ModelsSRPO: Self-ReferentialPolicy Optimization for…SRPO: Self-Referential Policy Optimization for Vision-Language-Action ModelsAVA-VLA: ImprovingVision-Language-Action…AVA-VLA: Improving Vision-Language-Action models with Active Visual AttentionDeepThinkVLA: EnhancingReasoning Capability of…DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action ModelsπRL: Online RLFine-tuning for…πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action ModelsABot-M0: VLA FoundationModel for Robotic…ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold LearningLIBERO-X: RobustnessLitmus for…LIBERO-X: Robustness Litmus for Vision-Language-Action ModelsVLA-JEPA: EnhancingVision-Language-Action…VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelStable Language Guidancefor…Stable Language Guidance for Vision-Language-Action ModelsOA-WAM:Object-Addressable Worl…OA-WAM: Object-Addressable World Action Model for Robust Robot ManipulationVLANeXt: Recipes forBuilding Strong VLA…VLANeXt: Recipes for Building Strong VLA ModelsACoT-VLA: ActionChain-of-Thought for…ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLIBERO-Plus: In-depthRobustness Analysis of…LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。