MolmoAct: Action Reasoning Models that can Reason in Space

Reasoning is central to purposeful action, yet most robotic foundation models map perception and instructions directly to control, which limits adaptability, generalization, and semantic grounding. We introduce Action Reasoning Models (ARMs), a class of robotic foundation models that integrate perception, planning, and control through a structured three-stage pipeline. Our model, MolmoAct, encodes observations and instructions into depth-aware perception tokens, generates mid-level spatial plans as editable trajectory traces, and predicts precise low-level actions, enabling explainable and steerable behavior. MolmoAct-7B-D achieves strong performance across simulation and real-world settings: 70.5% zero-shot accuracy on SimplerEnv Visual Matching tasks, surpassing closed-source Pi-0 and GR00T N1.5; 86.6% average success on LIBERO, including an additional 6.3% gain over ThinkAct on long-horizon tasks; and in real-world fine-tuning, an additional 10% (single-arm) and an additional 22.7% (bimanual) task progression over Pi-0-FAST. It also outperforms baselines by an additional 23.3% on out-of-distribution generalization and achieves top human-preference scores for open-ended instruction following and trajectory steering. Furthermore, we release, for the first time, the MolmoAct Dataset -- a mid-training robot dataset comprising over 10,000 high quality robot trajectories across diverse scenarios and tasks. Training with this dataset yields an average 5.5% improvement in general performance over the base model. We release all model weights, training code, our collected dataset, and our action reasoning dataset, establishing MolmoAct as both a state-of-the-art robotics foundation model and an open blueprint for building ARMs that transform perception into purposeful action through structured reasoning. Blogpost: https://allenai.org/blog/molmoact

Robotic Control viaEmbodied…Robotic Control via Embodied Chain-of-Thought ReasoningRoboPoint: AVision-Language Model…RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for RoboticsCogACT: A FoundationalVision-Language-Action…CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic ManipulationThinkAct:Vision-Language-Action…ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planningπ0.5: aVision-Language-Action…π0.5: a Vision-Language-Action Model with Open-World GeneralizationHAMSTER: HierarchicalAction Models For…HAMSTER: Hierarchical Action Models For Open-World Robot ManipulationTraceVLA: Visual TracePrompting Enhances…TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic PoliciesNORA: A SmallOpen-Sourced Generalist…NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied TasksWorldVLA: TowardsAutoregressive Action…WorldVLA: Towards Autoregressive Action World ModelAHA: AVision-Language-Model…AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic ManipulationSpatialVLA: ExploringSpatial Representations…SpatialVLA: Exploring Spatial Representations for Visual-Language-Action ModelData Scaling Laws inImitation Learning for…Data Scaling Laws in Imitation Learning for Robotic ManipulationFlowVLA: Thinking inMotion with a Visual…FlowVLA: Thinking in Motion with a Visual Chain of ThoughtEmbodiedOneVision:Interleaved…EmbodiedOneVision: Interleaved Vision-Text-Action Pretraining for General Robot ControlGemini Robotics 1.5:Pushing the Frontier of…Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion TransferInternVLA-M1: ASpatially Guided…InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot PolicyVlaser:Vision-Language-Action…Vlaser: Vision-Language-Action Model with Synergistic Embodied ReasoningUnified Diffusion VLA:Vision-Language-Action…Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion ProcessRobobench: AComprehensive Evaluatio…Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied BrainMolmoAct2: ActionReasoning Models for…MolmoAct2: Action Reasoning Models for Real-world DeploymentBeing-H0.5: ScalingHuman-Centric Robot…Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment GeneralizationST4VLA: Spatially GuidedTraining for…ST4VLA: Spatially Guided Training for Vision-Language-Action ModelsFast-ThinkAct: EfficientVision-Language-Action…Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent PlanningFrom Noise to Intent:Anchoring Generative VL…From Noise to Intent: Anchoring Generative VLA Policies with Residual BridgesMolmoAct: ActionReasoning Models that…MolmoAct: Action Reasoning Models that can Reason in Space過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。