Unsupervised Learning of Visual Features by Contrasting Cluster Assignments

Unsupervised image representations have significantly reduced the gap with supervised pretraining, notably with the recent achievements of contrastive learning methods. These contrastive methods typically work online and rely on a large number of explicit pairwise feature comparisons, which is computationally challenging. In this paper, we propose an online algorithm, SwAV, that takes advantage of contrastive methods without requiring to compute pairwise comparisons. Specifically, our method simultaneously clusters the data while enforcing consistency between cluster assignments produced for different augmentations (or views) of the same image, instead of comparing features directly as in contrastive learning. Simply put, we use a swapped prediction mechanism where we predict the cluster assignment of a view from the representation of another view. Our method can be trained with large and small batches and can scale to unlimited amounts of data. Compared to previous contrastive methods, our method is more memory efficient since it does not require a large memory bank or a special momentum network. In addition, we also propose a new data augmentation strategy, multi-crop, that uses a mix of views with different resolutions in place of two full-resolution views, without increasing the memory or compute requirements much. We validate our findings by achieving 75.3% top-1 accuracy on ImageNet with ResNet-50, as well as surpassing supervised pretraining on all the considered transfer tasks.

Colorful ImageColorizationColorful Image ColorizationUnsupervised FeatureLearning via…Unsupervised Feature Learning via Non-Parametric Instance DiscriminationData-Efficient ImageRecognition with…Data-Efficient Image Recognition with Contrastive Predictive CodingRevisitingSelf-Supervised Visual…Revisiting Self-Supervised Visual Representation LearningMomentum Contrast forUnsupervised Visual…Momentum Contrast for Unsupervised Visual Representation LearningA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsSelf-Supervised Learningof Pretext-Invariant…Self-Supervised Learning of Pretext-Invariant RepresentationsImproved Baselines withMomentum Contrastive…Improved Baselines with Momentum Contrastive LearningSupervised ContrastiveLearningSupervised Contrastive LearningContrastive MultiviewCodingContrastive Multiview CodingSelf-Supervised VisualFeature Learning With…Self-Supervised Visual Feature Learning With Deep Neural Networks: A SurveyPrototypical ContrastiveLearning of Unsupervise…Prototypical Contrastive Learning of Unsupervised RepresentationsDivide and Contrast:Self-supervised Learnin…Divide and Contrast: Self-supervised Learning from Uncurated DataExploring Cross-ImagePixel Contrast for…Exploring Cross-Image Pixel Contrast for Semantic SegmentationSelf-SupervisedLearning: Generative or…Self-Supervised Learning: Generative or ContrastiveReview onself-supervised image…Review on self-supervised image recognition using deep neural networksUniMoCo: Unsupervised,Semi-Supervised and…UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation LearningGeographicalKnowledge-Driven…Geographical Knowledge-Driven Representation Learning for Remote Sensing ImagesS2-BNN: Bridging the GapBetween Self-Supervised…S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution CalibrationOn the Importance ofAsymmetry for Siamese…On the Importance of Asymmetry for Siamese Representation LearningLearning Where to Learnin Cross-View…Learning Where to Learn in Cross-View Self-Supervised LearningSelf-SupervisedRepresentation Learning…Self-Supervised Representation Learning: Introduction, advances, and challengesSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureObservation, Analysis,and Solution: Exploring…Observation, Analysis, and Solution: Exploring Strong Lightweight Vision Transformers via Masked Image Modeling Pre-TrainingUnsupervised Learning ofVisual Features by…Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.