Motus: A Unified Latent Action World Model

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this paper, we propose Motus, a unified latent action world model that leverages existing general pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformer (MoT) architecture to integrate three experts (i.e., understanding, video generation, and action) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (i.e., world models, vision-language-action models, inverse dynamics models, video generation models, and video-action joint prediction models). Motus further leverages the optical flow to learn latent actions and adopts a recipe with three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level "delta action" and enabling large-scale action pretraining. Experiments show that Motus achieves superior performance against state-of-the-art methods in both simulation (a +15% improvement over X-VLA and a +45% improvement over Pi0.5) and real-world scenarios(improved by +11~48%), demonstrating unified modeling of all functionalities and priors significantly benefits downstream robotic tasks.

F1: AVision-Language-Action…F1: A Vision-Language-Action Model Bridging Understanding and Generation to ActionsVidar: Embodied VideoDiffusion Model for…Vidar: Embodied Video Diffusion Model for Generalist Bimanual ManipulationUnified World Models:Coupling Video and…Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsRoboTwin 2.0: A ScalableData Generator and…RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationX-VLA: Soft-PromptedTransformer as Scalable…X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelUniVLA: Learning to ActAnywhere with…UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsEgoDex: LearningDexterous Manipulation…EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric VideoFlowVLA: Thinking inMotion with a Visual…FlowVLA: Thinking in Motion with a Visual Chain of ThoughtAgiBot World Colosseo: ALarge-scale Manipulatio…AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsCoMo: LearningContinuous Latent Motio…CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot LearningShow-o2: Improved NativeUnified Multimodal…Show-o2: Improved Native Unified Multimodal ModelsQwen2.5-VL TechnicalReportQwen2.5-VL Technical ReportFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?Causal World Modelingfor Robot ControlCausal World Modeling for Robot ControlGigaWorld-Policy: AnEfficient…GigaWorld-Policy: An Efficient Action-Centered World-Action ModelInternVLA-A1: UnifyingUnderstanding…InternVLA-A1: Unifying Understanding, Generation and Action for Robotic ManipulationVLA-JEPA: EnhancingVision-Language-Action…VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelFrom Imagined Futures toExecutable Actions…From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationDiT4DiT: JointlyModeling Video Dynamics…DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot ControlBeing-H0.7: A LatentWorld-Action Model from…Being-H0.7: A Latent World-Action Model from Egocentric VideosLDA-1B: Scaling LatentDynamics Action Model…LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data IngestionWorld Model for RobotLearning: A…World Model for Robot Learning: A Comprehensive SurveyDo World Action ModelsGeneralize Better than…Do World Action Models Generalize Better than VLAs? A Robustness StudyABot-M0: VLA FoundationModel for Robotic…ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold LearningMotus: A Unified LatentAction World ModelMotus: A Unified Latent Action World ModelEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.