F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Executing language-conditioned tasks in dynamic visual environments remains a central challenge in embodied AI. Existing Vision-Language-Action (VLA) models predominantly adopt reactive state-to-action mappings, often leading to short-sighted behaviors and poor robustness in dynamic scenes. In this paper, we introduce F1, a pretrained VLA framework which integrates the visual foresight generation into decision-making pipeline. F1 adopts a Mixture-of-Transformer architecture with dedicated modules for perception, foresight generation, and control, thereby bridging understanding, generation, and actions. At its core, F1 employs a next-scale prediction mechanism to synthesize goal-conditioned visual foresight as explicit planning targets. By forecasting plausible future visual states, F1 reformulates action generation as a foresight-guided inverse dynamics problem, enabling actions that implicitly achieve visual goals. To endow F1 with robust and generalizable capabilities, we propose a three-stage training recipe on an extensive dataset comprising over 330k trajectories across 136 diverse tasks. This training scheme enhances modular reasoning and equips the model with transferable visual foresight, which is critical for complex and dynamic environments. Extensive evaluations on real-world tasks and simulation benchmarks demonstrate F1 consistently outperforms existing approaches, achieving substantial gains in both task success rate and generalization ability.

Towards Generalist RobotPolicies: What Matters…Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Modelsπ0: AVision-Language-Action…π0: A Vision-Language-Action Flow Model for General Robot ControlEmbodiedOneVision:Interleaved…EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot ControlInstructVLA:Vision-Language-Action…InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to ManipulationDreamVLA: AVision-Language-Action…DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeFlowVLA: Thinking inMotion with a Visual…FlowVLA: Thinking in Motion with a Visual Chain of ThoughtGR-3 Technical ReportGR-3 Technical ReportVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsHume: IntroducingSystem-2 Thinking in…Hume: Introducing System-2 Thinking in Visual-Language-Action ModelGraspVLA: a GraspingFoundation Model…GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action DataWorldVLA: TowardsAutoregressive Action…WorldVLA: Towards Autoregressive Action World ModelUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsInternVLA-M1: ASpatially Guided…InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot PolicyMotus: A Unified LatentAction World ModelMotus: A Unified Latent Action World ModelUnified Diffusion VLA:Vision-Language-Action…Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion ProcessManualVLA: A Unified VLAModel for…ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic ManipulationAsyncVLA: AsynchronousFlow Matching for…AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action ModelsInternVLA-A1: UnifyingUnderstanding…InternVLA-A1: Unifying Understanding, Generation and Action for Robotic ManipulationACoT-VLA: ActionChain-of-Thought for…ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsOA-WAM:Object-Addressable Worl…OA-WAM: Object-Addressable World Action Model for Robust Robot ManipulationAffordanceVLA: AVision-Language-Action…AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware UnderstandingRoboInter: A HolisticIntermediate…RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationFutureVLA: JointVisuomotor Prediction…FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action ModelVLANeXt: Recipes forBuilding Strong VLA…VLANeXt: Recipes for Building Strong VLA ModelsF1: AVision-Language-Action…F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。