Causal World Modeling for Robot Control

This work highlights that video world modeling, alongside vision-language pre-training, establishes a fresh and independent foundation for robot learning. Intuitively, video world models provide the ability to imagine the near future by understanding the causality between actions and visual dynamics. Inspired by this, we introduce LingBot-VA, an autoregressive diffusion framework that learns frame prediction and policy execution simultaneously. Our model features three carefully crafted designs: (1) a shared latent space, integrating vision and action tokens, driven by a Mixture-of-Transformers (MoT) architecture, (2) a closed-loop rollout mechanism, allowing for ongoing acquisition of environmental feedback with ground-truth observations, (3) an asynchronous inference pipeline, parallelizing action prediction and motor execution to support efficient control. We evaluate our model on both simulation benchmarks and real-world scenarios, where it shows significant promise in long-horizon manipulation, data efficiency in post-training, and strong generalizability to novel configurations. The code and model are made publicly available to facilitate the community.

Motus: A Unified LatentAction World ModelMotus: A Unified Latent Action World Modelmimic-video:Video-Action Models for…mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAsVidar: Embodied VideoDiffusion Model for…Vidar: Embodied Video Diffusion Model for Generalist Bimanual ManipulationX-VLA: Soft-PromptedTransformer as Scalable…X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelMixture of Horizons inAction ChunkingMixture of Horizons in Action ChunkingGR-3 Technical ReportGR-3 Technical ReportUnifiedVision-Language-Action…Unified Vision-Language-Action ModelRoboTwin 2.0: A ScalableData Generator and…RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationDiscrete Diffusion VLA:Bringing Discrete…Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action PoliciesMemoryVLA:Perceptual-Cognitive…MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic ManipulationTraining-Time ActionConditioning for…Training-Time Action Conditioning for Efficient Real-Time ChunkingCosmos Policy:Fine-Tuning Video Model…Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningFast-WAM: Do WorldAction Models Need…Fast-WAM: Do World Action Models Need Test-time Future Imagination?StarVLA: A Lego-likeCodebase for…StarVLA: A Lego-like Codebase for Vision-Language-Action Model DevelopingDiT4DiT: JointlyModeling Video Dynamics…DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot ControlWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot PoliciesHoloBrain-0 TechnicalReportHoloBrain-0 Technical ReportFrom Imagined Futures toExecutable Actions…From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationBeing-H0.7: A LatentWorld-Action Model from…Being-H0.7: A Latent World-Action Model from Egocentric VideosOA-WAM:Object-Addressable Worl…OA-WAM: Object-Addressable World Action Model for Robust Robot ManipulationWorld Model for RobotLearning: A…World Model for Robot Learning: A Comprehensive SurveyVAG: Dual-StreamVideo-Action Generation…VAG: Dual-Stream Video-Action Generation for Embodied Data SynthesisWorld-Language-ActionModel for Unified World…World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action SynthesisWALL-WM: Carving WorldAction Modeling at the…WALL-WM: Carving World Action Modeling at the Event JointsCausal World Modelingfor Robot ControlCausal World Modeling for Robot ControlEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.