π*0.6: a VLA That Learns From Experience

We study how vision-language-action (VLA) models can improve through real-world deployments via reinforcement learning (RL). We present a general-purpose method, RL with Experience and Corrections via Advantage-conditioned Policies (RECAP), that provides for RL training of VLAs via advantage conditioning. Our method incorporates heterogeneous data into the self-improvement process, including demonstrations, data from on-policy collection, and expert teleoperated interventions provided during autonomous execution. RECAP starts by pre-training a generalist VLA with offline RL, which we call $π^{*}_{0.6}$, that can then be specialized to attain high performance on downstream tasks through on-robot data collection. We show that the $π^{*}_{0.6}$ model trained with the full RECAP method can fold laundry in real homes, reliably assemble boxes, and make espresso drinks using a professional espresso machine. On some of the hardest tasks, RECAP more than doubles task throughput and roughly halves the task failure rate.

GRAPE: GeneralizingRobot Policy via…GRAPE: Generalizing Robot Policy via Preference AlignmentPolicy Agnostic RL:Offline RL and Online R…Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and BackboneπRL: Online RLFine-tuning for…πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action ModelsConRFT: A ReinforcedFine-tuning Method for…ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency PolicyVLA-RL: TowardsMasterful and General…VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningWhat Can RL Bring to VLAGeneralization? An…What Can RL Bring to VLA Generalization? An Empirical StudySimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningInteractivePost-Training for…Interactive Post-Training for Vision-Language-Action ModelsAVision-Language-Action-…A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement LearningCO-RFT: EfficientFine-Tuning of…CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement LearningImprovingVision-Language-Action…Improving Vision-Language-Action Model with Online Reinforcement LearningRLDG: Robotic GeneralistPolicy Distillation via…RLDG: Robotic Generalist Policy Distillation via Reinforcement LearningRLinf-VLA: A Unified andEfficient Framework for…RLinf-VLA: A Unified and Efficient Framework for VLA+RL TrainingRobust Finetuning ofVision-Language-Action…Robust Finetuning of Vision-Language-Action Robot Policies via Parameter MergingEmbodied RobotManipulation in the Era…Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning PerspectivesRISE: Self-ImprovingRobot Policy with…RISE: Self-Improving Robot Policy with Compositional World ModelIG-RFT: AnInteraction-Guided RL…IG-RFT: An Interaction-Guided RL Framework for VLA Models in Long-Horizon Robotic ManipulationVLAW: IterativeCo-Improvement of…VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Modelπ0.7: a SteerableGeneralist Robotic…π0.7: a Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesLearning whileDeploying: Fleet-Scale…Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot PoliciesWorld-VLA-Loop:Closed-Loop Learning of…World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA PolicyLaST-R1: ReinforcingAction via Adaptive…LaST-R1: Reinforcing Action via Adaptive Physical Latent Reasoning for VLA ModelsUnified Noise Steeringfor Efficient…Unified Noise Steering for Efficient Human-Guided VLA AdaptationTOPReward: TokenProbabilities as Hidden…TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Roboticsπ*0.6: a VLA That LearnsFrom Experienceπ*0.6: a VLA That Learns From Experience過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。