Masked Autoencoders Are Scalable Vision Learners

This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core designs. First, we develop an asymmetric encoder-decoder architecture, with an encoder that operates only on the visible subset of patches (without mask tokens), along with a lightweight decoder that reconstructs the original image from the latent representation and mask tokens. Second, we find that masking a high proportion of the input image, e.g., 75%, yields a nontrivial and meaningful self-supervisory task. Coupling these two designs enables us to train large models efficiently and effectively: we accelerate training (by 3x or more) and improve accuracy. Our scalable approach allows for learning high-capacity models that generalize well: e.g., a vanilla ViT-Huge model achieves the best accuracy (87.8%) among methods that use only ImageNet-1K data. Transfer performance in downstream tasks outperforms supervised pre-training and shows promising scaling behavior.

Deep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionColorful ImageColorizationColorful Image ColorizationUnsupervised FeatureLearning via…Unsupervised Feature Learning via Non-Parametric Instance DiscriminationRepresentation Learningwith Contrastive…Representation Learning with Contrastive Predictive CodingSemantic Understandingof Scenes Through the…Semantic Understanding of Scenes Through the ADE20K DatasetCutMix: RegularizationStrategy to Train Stron…CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesFixing the train-testresolution discrepancyFixing the train-test resolution discrepancyAn Empirical Study ofTraining Self-Supervise…An Empirical Study of Training Self-Supervised Vision TransformersSelf-supervisedPretraining of Visual…Self-supervised Pretraining of Visual Features in the WildEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersBEiT: BERT Pre-Trainingof Image TransformersBEiT: BERT Pre-Training of Image TransformersVOLO: Vision Outlookerfor Visual RecognitionVOLO: Vision Outlooker for Visual RecognitionMasked Siamese Networksfor Label-Efficient…Masked Siamese Networks for Label-Efficient LearningAdaptFormer: AdaptingVision Transformers for…AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionJoint Feature Learningand Relation Modeling…Joint Feature Learning and Relation Modeling for Tracking: A One-Stream FrameworkAn Empirical Study ofRemote Sensing…An Empirical Study of Remote Sensing PretrainingCoCa: ContrastiveCaptioners are…CoCa: Contrastive Captioners are Image-Text Foundation ModelsSelf-Supervised VisionTransformers for Joint…Self-Supervised Vision Transformers for Joint SAR-Optical Representation LearningHow to Understand MaskedAutoencodersHow to Understand Masked AutoencodersSeasoning Model Soupsfor Robustness to…Seasoning Model Soups for Robustness to Adversarial and Natural Distribution ShiftsRevisiting FeaturePrediction for Learning…Revisiting Feature Prediction for Learning Visual Representations from VideoMultimodal WebNavigation with…Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsWhite-Box Transformersvia Sparse Rate…White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?DINOv2 Meets Text: AUnified Framework for…DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language AlignmentMasked Autoencoders AreScalable Vision LearnersMasked Autoencoders Are Scalable Vision LearnersEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.