Masked Siamese Networks for Label-Efficient Learning

We propose Masked Siamese Networks (MSN), a self-supervised learning framework for learning image representations. Our approach matches the representation of an image view containing randomly masked patches to the representation of the original unmasked image. This self-supervised pre-training strategy is particularly scalable when applied to Vision Transformers since only the unmasked patches are processed by the network. As a result, MSNs improve the scalability of joint-embedding architectures, while producing representations of a high semantic level that perform competitively on low-shot image classification. For instance, on ImageNet-1K, with only 5,000 annotated images, our base MSN model achieves 72.4% top-1 accuracy, and with 1% of ImageNet-1K labels, we achieve 75.7% top-1 accuracy, setting a new state-of-the-art for self-supervised learning on this benchmark. Our code is publicly available.

A Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsBootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningBig Self-SupervisedModels are Strong…Big Self-Supervised Models are Strong Semi-Supervised LearnersMomentum Contrast forUnsupervised Visual…Momentum Contrast for Unsupervised Visual Representation LearningEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersAre Large-scale DatasetsNecessary for…Are Large-scale Datasets Necessary for Self-Supervised Pre-training?Semi-Supervised Learningof Visual Features by…Semi-Supervised Learning of Visual Features by Non-Parametrically Predicting View Assignments with Support SamplesExploring Simple SiameseRepresentation LearningExploring Simple Siamese Representation LearningSimMIM: A SimpleFramework for Masked…SimMIM: A Simple Framework for Masked Image ModelingMasked FeaturePrediction for…Masked Feature Prediction for Self-Supervised Visual Pre-TrainingBEiT: BERT Pre-Trainingof Image TransformersBEiT: BERT Pre-Training of Image TransformersMasked Autoencoders AreScalable Vision LearnersMasked Autoencoders Are Scalable Vision LearnersMasked Siamese ConvNetsMasked Siamese ConvNetsSupMAE: SupervisedMasked Autoencoders Are…SupMAE: Supervised Masked Autoencoders Are Efficient Vision LearnersContrastive MaskedAutoencoders are…Contrastive Masked Autoencoders are Stronger Vision LearnersSiamese Image Modelingfor Self-Supervised…Siamese Image Modeling for Self-Supervised Vision Representation LearningThe Hidden UniformCluster Prior in…The Hidden Uniform Cluster Prior in Self-Supervised LearningUnderstanding MaskedImage Modeling via…Understanding Masked Image Modeling via Learning Occlusion Invariant FeatureMAGE: MAsked GenerativeEncoder to Unify…MAGE: MAsked Generative Encoder to Unify Representation Learning and Image SynthesisArchitecture-AgnosticMasked Image Modeling -…Architecture-Agnostic Masked Image Modeling - From ViT back to CNNLayer GraftedPre-training: Bridging…Layer Grafted Pre-training: Bridging Contrastive Learning And Masked Image Modeling For Label-Efficient RepresentationsRevisiting FeaturePrediction for Learning…Revisiting Feature Prediction for Learning Visual Representations from VideoContrastive Tuning: ALittle Help to Make…Contrastive Tuning: A Little Help to Make Masked Autoencoders ForgetDiffusion-Based ActionRecognition Generalizes…Diffusion-Based Action Recognition Generalizes to Untrained DomainsMasked Siamese Networksfor Label-Efficient…Masked Siamese Networks for Label-Efficient LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.