Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative approach for self-supervised learning from images. The idea behind I-JEPA is simple: from a single context block, predict the representations of various target blocks in the same image. A core design choice to guide I-JEPA towards producing semantic representations is the masking strategy; specifically, it is crucial to (a) sample target blocks with sufficiently large scale (semantic), and to (b) use a sufficiently informative (spatially distributed) context block. Empirically, when combined with Vision Transformers, we find I-JEPA to be highly scalable. For instance, we train a ViT-Huge/14 on ImageNet using 16 A100 GPUs in under 72 hours to achieve strong downstream performance across a wide range of tasks, from linear classification to object counting and depth prediction.

Bootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningMomentum Contrast forUnsupervised Visual…Momentum Contrast for Unsupervised Visual Representation LearningA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersAn Empirical Study ofTraining Self-Supervise…An Empirical Study of Training Self-Supervised Vision TransformersAre Large-scale DatasetsNecessary for…Are Large-scale Datasets Necessary for Self-Supervised Pre-training?Exploring Simple SiameseRepresentation LearningExploring Simple Siamese Representation LearningMasked FeaturePrediction for…Masked Feature Prediction for Self-Supervised Visual Pre-TrainingBEiT: BERT Pre-Trainingof Image TransformersBEiT: BERT Pre-Training of Image TransformersSimMIM: A SimpleFramework for Masked…SimMIM: A Simple Framework for Masked Image ModelingThe Hidden UniformCluster Prior in…The Hidden Uniform Cluster Prior in Self-Supervised LearningContext Autoencoder forSelf-Supervised…Context Autoencoder for Self-Supervised Representation LearningThe effectiveness of MAEpre-pretraining for…The effectiveness of MAE pre-pretraining for billion-scale pretrainingHMSN: HyperbolicSelf-Supervised Learnin…HMSN: Hyperbolic Self-Supervised Learning by Clustering with Ideal PrototypesMC-JEPA: AJoint-Embedding…MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content FeaturesTemporal DINO: ASelf-supervised Video…Temporal DINO: A Self-supervised Video Strategy to Enhance Action PredictionScalable Pre-training ofLarge Autoregressive…Scalable Pre-training of Large Autoregressive Image ModelsSD-DiT: Unleashing thePower of Self-Supervise…SD-DiT: Unleashing the Power of Self-Supervised Discrimination in Diffusion Transformer*You Don't NeedDomain-Specific Data…You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised LearningEEG2Rep: EnhancingSelf-supervised EEG…EEG2Rep: Enhancing Self-supervised EEG Representation Through Informative Masked InputsRethinkingGeneralizability and…Rethinking Generalizability and Discriminability of Self-Supervised Learning from Evolutionary Game Theory PerspectiveThe Dynamic Duo ofCollaborative Masking…The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder LearningDiffusion-Based ActionRecognition Generalizes…Diffusion-Based Action Recognition Generalizes to Untrained DomainsSelf-Distillation ofHidden Layers for…Self-Distillation of Hidden Layers for Self-Supervised Representation LearningSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。