Emerging Properties in Self-Supervised Vision Transformers

In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) [16] that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder [26], multi-crop training [9], and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.

Unsupervised FeatureLearning via…Unsupervised Feature Learning via Non-Parametric Instance DiscriminationUnsupervised Learning ofVisual Features by…Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsBootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsBarlow Twins:Self-Supervised Learnin…Barlow Twins: Self-Supervised Learning via Redundancy ReductionOnlineBag-of-Visual-Words…Online Bag-of-Visual-Words Generation for Unsupervised Representation LearningExploring Simple SiameseRepresentation LearningExploring Simple Siamese Representation LearningSelf-supervisedPretraining of Visual…Self-supervised Pretraining of Visual Features in the WildTraining data-efficientimage transformers &…Training data-efficient image transformers & distillation through attentionAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleSEED: Self-supervisedDistillation For Visual…SEED: Self-supervised Distillation For Visual RepresentationSeed the Views:Hierarchical Semantic…Seed the Views: Hierarchical Semantic Alignment for Contrastive Representation LearningThe Challenges ofContinuous…The Challenges of Continuous Self-Supervised LearningC3-DINO: JointContrastive and…C3-DINO: Joint Contrastive and Non-Contrastive Self-Supervised Learning for Speaker VerificationSelf-Supervised VisionTransformers for Joint…Self-Supervised Vision Transformers for Joint SAR-Optical Representation LearningSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureMulti-Mode OnlineKnowledge Distillation…Multi-Mode Online Knowledge Distillation for Self-Supervised Visual Representation LearningA Unified View of MaskedImage ModelingA Unified View of Masked Image ModelingWild Face Anti-SpoofingChallenge 2023…Wild Face Anti-Spoofing Challenge 2023: Benchmark and ResultsYet Another TrafficClassifier: A Masked…Yet Another Traffic Classifier: A Masked Autoencoder Based Traffic Transformer with Multi-Level Flow RepresentationDiGA: Distil toGeneralize and then…DiGA: Distil to Generalize and then Adapt for Domain Adaptive Semantic SegmentationUnsupervised ObjectLocalization in the Era…Unsupervised Object Localization in the Era of Self-Supervised ViTs: A SurveyA Closer Look atBenchmarking…A Closer Look at Benchmarking Self-Supervised Pre-training with Image ClassificationTraining ahigh-performance retina…Training a high-performance retinal foundation model with half-the-data and 400 times less computeEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision Transformers過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。