VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their ability to generalize to novel tasks, objects, and settings remains limited. In this work, we present VideoVLA, a simple approach that explores the potential of transforming large video generation models into robotic VLA manipulators. Given a language instruction and an image, VideoVLA predicts an action sequence as well as the future visual outcomes. Built on a multi-modal Diffusion Transformer, VideoVLA jointly models video, language, and action modalities, using pre-trained video generative models for joint visual and action forecasting. Our experiments show that high-quality imagined futures correlate with reliable action predictions and task success, highlighting the importance of visual imagination in manipulation. VideoVLA demonstrates strong generalization, including imitating other embodiments' skills and handling novel objects. This dual-prediction strategy - forecasting both actions and their visual consequences - explores a paradigm shift in robot learning and unlocks generalization capabilities in manipulation systems.

LLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsCogACT: A FoundationalVision-Language-Action…CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic ManipulationOpenVLA: An Open-SourceVision-Language-Action…OpenVLA: An Open-Source Vision-Language-Action ModelHunyuanVideo: ASystematic Framework Fo…HunyuanVideo: A Systematic Framework For Large Video Generative ModelsPhi-3 Technical Report:A Highly Capable…Phi-3 Technical Report: A Highly Capable Language Model Locally on Your PhoneGR-2: A GenerativeVideo-Language-Action…GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationDINOv2: Learning RobustVisual Features without…DINOv2: Learning Robust Visual Features without SupervisionVideo Prediction Policy:A Generalist Robot…Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsSpatialVLA: ExploringSpatial Representations…SpatialVLA: Exploring Spatial Representations for Visual-Language-Action ModelRDT-1B: a DiffusionFoundation Model for…RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationUnified Video ActionModelUnified Video Action ModelCogVideoX: Text-to-VideoDiffusion Models with A…CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerBeing-H0.7: A LatentWorld-Action Model from…Being-H0.7: A Latent World-Action Model from Egocentric VideosDriveVA: Video ActionModels are Zero-Shot…DriveVA: Video Action Models are Zero-Shot DriversDiT4DiT: JointlyModeling Video Dynamics…DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot ControlVP-VLA: Visual Promptingas an Interface for…VP-VLA: Visual Prompting as an Interface for Vision-Language-Action ModelsNoiseGate: LearningPer-Latent Timestep…NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action ModelsTwinBrainVLA: Unleashingthe Potential of…TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-TransformersCausal World Modelingfor Robot ControlCausal World Modeling for Robot ControlAfford-VLA:Action-Aligned Visual…Afford-VLA: Action-Aligned Visual Planning via Internalized AffordanceSelf-Supervised FlowMatching for Scalable…Self-Supervised Flow Matching for Scalable Multi-Modal SynthesisNavDreamer: Video Modelsas Zero-Shot 3D…NavDreamer: Video Models as Zero-Shot 3D NavigatorsGigaWorld-Policy: AnEfficient…GigaWorld-Policy: An Efficient Action-Centered World-Action ModelDream.exe: Can VideoGeneration Models Dream…Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?VideoVLA: VideoGenerators Can Be…VideoVLA: Video Generators Can Be Generalizable Robot Manipulators過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。