V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data (robot trajectories), to develop models capable of understanding, predicting, and planning in the physical world. We first pre-train an action-free joint-embedding-predictive architecture, V-JEPA 2, on a video and image dataset comprising over 1 million hours of internet video. V-JEPA 2 achieves strong performance on motion understanding (77.3 top-1 accuracy on Something-Something v2) and state-of-the-art performance on human action anticipation (39.7 recall-at-5 on Epic-Kitchens-100) surpassing previous task-specific models. Additionally, after aligning V-JEPA 2 with a large language model, we demonstrate state-of-the-art performance on multiple video question-answering tasks at the 8 billion parameter scale (e.g., 84.0 on PerceptionTest, 76.9 on TempCompass). Finally, we show how self-supervised learning can be applied to robotic planning tasks by post-training a latent action-conditioned world model, V-JEPA 2-AC, using less than 62 hours of unlabeled robot videos from the Droid dataset. We deploy V-JEPA 2-AC zero-shot on Franka arms in two different labs and enable picking and placing of objects using planning with image goals. Notably, this is achieved without collecting any data from the robots in these environments, and without any task-specific training or reward. This work demonstrates how self-supervised learning from web-scale data and a small amount of robot interaction data can yield a world model capable of planning in the physical world.

GAIA-1: A GenerativeWorld Model for…GAIA-1: A Generative World Model for Autonomous DrivingRevisiting FeaturePrediction for Learning…Revisiting Feature Prediction for Learning Visual Representations from VideoTD-MPC2: Scalable,Robust World Models for…TD-MPC2: Scalable, Robust World Models for Continuous ControlUnleashing Large-ScaleVideo Generative…Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationDINO-WM: World Models onPre-trained Visual…DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningPerception Encoder: Thebest visual embeddings…Perception Encoder: The best visual embeddings are not at the output of the networkGR00T N1: An OpenFoundation Model for…GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsPerceptionLM:Open-Access Data and…PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingLearning fromReward-Free Offline…Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsTarsier2: AdvancingLarge Vision-Language…Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video UnderstandingGAIA-2: A ControllableMulti-View Generative…GAIA-2: A Controllable Multi-View Generative World Model for Autonomous DrivingLLaVA-Video: VideoInstruction Tuning With…LLaVA-Video: Video Instruction Tuning With Synthetic DataPlanning with Reasoningusing Vision Language…Planning with Reasoning using Vision Language World ModelWorld Models CanLeverage Human Videos…World Models Can Leverage Human Videos for Dexterous ManipulationGigaBrain-0: A WorldModel-Powered…GigaBrain-0: A World Model-Powered Vision-Language-Action ModelOrbis: OvercomingChallenges of…Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World ModelsWorld Action Models areZero-shot PoliciesWorld Action Models are Zero-shot PoliciesDreamDojo: A GeneralistRobot World Model from…DreamDojo: A Generalist Robot World Model from Large-Scale Human VideosWorld Action Models: TheNext Frontier in…World Action Models: The Next Frontier in Embodied AIWorld Models: AComprehensive Survey of…World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and ApplicationsVisuo-Tactile WorldModelsVisuo-Tactile World ModelsDo World Action ModelsGeneralize Better than…Do World Action Models Generalize Better than VLAs? A Robustness StudyAgentic World Modeling:Foundations…Agentic World Modeling: Foundations, Capabilities, Laws, and BeyondBeyond LanguageModeling: An Exploratio…Beyond Language Modeling: An Exploration of Multimodal PretrainingV-JEPA 2:Self-Supervised Video…V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.