Self-supervised Learning is More Robust to Dataset Imbalance

Self-supervised learning (SSL) is a scalable way to learn general visual representations since it learns without labels. However, large-scale unlabeled datasets in the wild often have long-tailed label distributions, where we know little about the behavior of SSL. In this work, we systematically investigate self-supervised learning under dataset imbalance. First, we find out via extensive experiments that off-the-shelf self-supervised representations are already more robust to class imbalance than supervised representations. The performance gap between balanced and imbalanced pre-training with SSL is significantly smaller than the gap with supervised learning, across sample sizes, for both in-domain and, especially, out-of-domain evaluation. Second, towards understanding the robustness of SSL, we hypothesize that SSL learns richer features from frequent data: it may learn label-irrelevant-but-transferable features that help classify the rare classes and downstream tasks. In contrast, supervised learning has no incentive to learn features irrelevant to the labels from frequent examples. We validate this hypothesis with semi-synthetic experiments and theoretical analyses on a simplified setting. Third, inspired by the theoretical insights, we devise a re-weighted regularization technique that consistently improves the SSL representation quality on imbalanced datasets with several evaluation criteria, closing the small gap between balanced and imbalanced datasets with the same number of examples.

SMOTE: SyntheticMinority Over-sampling…SMOTE: Synthetic Minority Over-sampling TechniqueADASYN: Adaptivesynthetic sampling…ADASYN: Adaptive synthetic sampling approach for imbalanced learningDropout: a simple way toprevent neural networks…Dropout: a simple way to prevent neural networks from overfittingDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionLearning from imbalanceddata: open challenges…Learning from imbalanced data: open challenges and future directionsA systematic study ofthe class imbalance…A systematic study of the class imbalance problem in convolutional neural networksUnsupervisedRepresentation Learning…Unsupervised Representation Learning by Predicting Image RotationsBootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningPredicting What YouAlready Know Helps…Predicting What You Already Know Helps: Provable Self-Supervised LearningProvable Guarantees forSelf-Supervised Deep…Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossUnderstandingself-supervised Learnin…Understanding self-supervised Learning Dynamics without Contrastive PairsWhen Does ContrastiveVisual Representation…When Does Contrastive Visual Representation Learning Work?UnderstandingCross-Domain Few-Shot…Understanding Cross-Domain Few-Shot Learning Based on Domain Similarity and Few-Shot DifficultyAssessing the State ofSelf-Supervised Human…Assessing the State of Self-Supervised Human Activity Recognition Using WearablesArCL: EnhancingContrastive Learning…ArCL: Enhancing Contrastive Learning with Augmentation-Robust RepresentationsGeometric ContrastiveLearningGeometric Contrastive LearningSelf-Supervised RemoteSensing Feature…Self-Supervised Remote Sensing Feature Learning: Learning Paradigms, Challenges, and Future WorksSCALE: OnlineSelf-Supervised Lifelon…SCALE: Online Self-Supervised Lifelong Learning without Prior KnowledgeDive into the details ofself-supervised learnin…Dive into the details of self-supervised learning for medical image analysisLearning Imbalanced Datawith Vision TransformersLearning Imbalanced Data with Vision TransformersClass-ConditionalSharpness-Aware…Class-Conditional Sharpness-Aware Minimization for Deep Long-Tailed RecognitionA Survey of Methods forAddressing Class…A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language ProcessingDUEL: DuplicateElimination on Active…DUEL: Duplicate Elimination on Active Memory for Self-Supervised Class-Imbalanced LearningDelving Deep intoSimplicity Bias for…Delving Deep into Simplicity Bias for Long-Tailed Image RecognitionSelf-supervised Learningis More Robust to…Self-supervised Learning is More Robust to Dataset ImbalanceEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.