Context Autoencoder for Self-Supervised Representation Learning

We present a novel masked image modeling (MIM) approach, context autoencoder (CAE), for self-supervised representation pretraining. We pretrain an encoder by making predictions in the encoded representation space. The pretraining tasks include two tasks: masked representation prediction - predict the representations for the masked patches, and masked patch reconstruction - reconstruct the masked patches. The network is an encoder-regressor-decoder architecture: the encoder takes the visible patches as input; the regressor predicts the representations of the masked patches, which are expected to be aligned with the representations computed from the encoder, using the representations of visible patches and the positions of visible and masked patches; the decoder reconstructs the masked patches from the predicted encoded representations. The CAE design encourages the separation of learning the encoder (representation) from completing the pertaining tasks: masked representation prediction and masked patch reconstruction tasks, and making predictions in the encoded representation space empirically shows the benefit to representation learning. We demonstrate the effectiveness of our CAE through superior transfer performance in downstream tasks: semantic segmentation, object detection and instance segmentation, and classification. The code will be available at https://github.com/Atten4Vis/CAE.

Bootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningAre Large-scale DatasetsNecessary for…Are Large-scale Datasets Necessary for Self-Supervised Pre-training?Emerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersSiT: Self-supervisedvIsion TransformerSiT: Self-supervised vIsion TransformerSimMIM: A SimpleFramework for Masked…SimMIM: A Simple Framework for Masked Image ModelingMasked FeaturePrediction for…Masked Feature Prediction for Self-Supervised Visual Pre-TrainingBEiT: BERT Pre-Trainingof Image TransformersBEiT: BERT Pre-Training of Image TransformersContrastive MaskedAutoencoders are…Contrastive Masked Autoencoders are Stronger Vision LearnersSiamese Image Modelingfor Self-Supervised…Siamese Image Modeling for Self-Supervised Vision Representation LearningCorrupted Image Modelingfor Self-Supervised…Corrupted Image Modeling for Self-Supervised Visual Pre-TrainingUnderstanding MaskedImage Modeling via…Understanding Masked Image Modeling via Learning Occlusion Invariant FeaturePeCo: PerceptualCodebook for BERT…PeCo: Perceptual Codebook for BERT Pre-training of Vision TransformersBootstrapped MaskedAutoencoders for Vision…Bootstrapped Masked Autoencoders for Vision BERT PretrainingMVP:Multimodality-Guided…MVP: Multimodality-Guided Visual Pre-trainingContrastive MaskedAutoencoders are…Contrastive Masked Autoencoders are Stronger Vision LearnersSiamese Image Modelingfor Self-Supervised…Siamese Image Modeling for Self-Supervised Vision Representation LearningFast-iTPN: IntegrallyPre-Trained Transformer…Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network With Token MigrationUnderstanding MaskedImage Modeling via…Understanding Masked Image Modeling via Learning Occlusion Invariant FeatureSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureDisjoint Masking WithJoint Distillation for…Disjoint Masking With Joint Distillation for Efficient Masked Image ModelingSERE: Exploring FeatureSelf-Relation for…SERE: Exploring Feature Self-Relation for Self-Supervised TransformerMAGE: MAsked GenerativeEncoder to Unify…MAGE: MAsked Generative Encoder to Unify Representation Learning and Image SynthesisHard Patches Mining forMasked Image ModelingHard Patches Mining for Masked Image ModelingMaskCLIP: MaskedSelf-Distillation…MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingContext Autoencoder forSelf-Supervised…Context Autoencoder for Self-Supervised Representation Learning過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。