Revisiting Feature Prediction for Learning Visual Representations from Video

This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K.

Bootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersAn Empirical Study ofTraining Self-Supervise…An Empirical Study of Training Self-Supervised Vision TransformersAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleMasked Siamese Networksfor Label-Efficient…Masked Siamese Networks for Label-Efficient LearningSimMIM: A SimpleFramework for Masked…SimMIM: A Simple Framework for Masked Image ModelingBEiT: BERT Pre-Trainingof Image TransformersBEiT: BERT Pre-Training of Image TransformersCoCa: ContrastiveCaptioners are…CoCa: Contrastive Captioners are Image-Text Foundation ModelsMasked Autoencoders AreScalable Vision LearnersMasked Autoencoders Are Scalable Vision LearnersContext Autoencoder forSelf-Supervised…Context Autoencoder for Self-Supervised Representation LearningDINOv2: Learning RobustVisual Features without…DINOv2: Learning Robust Visual Features without SupervisionModeling CaptionDiversity in Contrastiv…Modeling Caption Diversity in Contrastive Vision-Language PretrainingDynaMo: In-DomainDynamics Pretraining fo…DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor ControlLarge Concept Models:Language Modeling in a…Large Concept Models: Language Modeling in a Sentence Representation SpaceOnlineVPO: Align VideoDiffusion Model with…OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference OptimizationLearning from StreamingVideo with Orthogonal…Learning from Streaming Video with Orthogonal GradientsUniVLA: Learning to ActAnywhere with…UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsCritiques of WorldModelsCritiques of World ModelsDisMo: DisentangledMotion Representations…DisMo: Disentangled Motion Representations for Open-World Motion TransferDiffusion-Based ActionRecognition Generalizes…Diffusion-Based Action Recognition Generalizes to Untrained DomainsOneVision-Encoder:Codec-Aligned Sparsity…OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal IntelligenceSapiens2Sapiens2Agentic World Modeling:Foundations…Agentic World Modeling: Foundations, Capabilities, Laws, and BeyondRevisiting FeaturePrediction for Learning…Revisiting Feature Prediction for Learning Visual Representations from VideoEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.