Understanding self-supervised Learning Dynamics without Contrastive Pairs

While contrastive approaches of self-supervised learning (SSL) learn representations by minimizing the distance between two augmented views of the same data point (positive pairs) and maximizing views from different data points (negative pairs), recent \\emph{non-contrastive} SSL (e.g., BYOL and SimSiam) show remarkable performance {\\it without} negative pairs, with an extra learnable predictor and a stop-gradient operation. A fundamental question arises: why do these methods not collapse into trivial representations? We answer this question via a simple theoretical study and propose a novel approach, DirectPred, that \\emph{directly} sets the linear predictor based on the statistics of its inputs, without gradient training. On ImageNet, it performs comparably with more complex two-layer non-linear predictors that employ BatchNorm and outperforms a linear predictor by $2.5\\%$ in 300-epoch training (and $5\\%$ in 60-epoch). DirectPred is motivated by our theoretical study of the nonlinear learning dynamics of non-contrastive SSL in simple linear networks. Our study yields conceptual insights into how non-contrastive SSL methods learn, how they avoid representational collapse, and how multiple factors, like predictor networks, stop-gradients, exponential moving averages, and weight decay all come into play. Our simple theory recapitulates the results of real-world ablation studies in both STL-10 and ImageNet. Code is released https://github.com/facebookresearch/luckmatters/tree/master/ssl.

ImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionAlgorithmicRegularization in…Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically BalancedLearning Representationsby Maximizing Mutual…Learning Representations by Maximizing Mutual Information Across ViewsBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingBootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsImproved Baselines withMomentum Contrastive…Improved Baselines with Momentum Contrastive LearningContrastive MultiviewCodingContrastive Multiview CodingExploring Simple SiameseRepresentation LearningExploring Simple Siamese Representation LearningPredicting What YouAlready Know Helps…Predicting What You Already Know Helps: Provable Self-Supervised LearningLearning Multiple Layersof Features from Tiny…Learning Multiple Layers of Features from Tiny ImagesTowards DemystifyingRepresentation Learning…Towards Demystifying Representation Learning with Non-contrastive Self-supervisionBarlow Twins:Self-Supervised Learnin…Barlow Twins: Self-Supervised Learning via Redundancy ReductionOn Feature Decorrelationin Self-Supervised…On Feature Decorrelation in Self-Supervised LearningVICReg:Variance-Invariance-Cov…VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningUnderstandingDimensional Collapse in…Understanding Dimensional Collapse in Contrastive Self-supervised LearningCrafting BetterContrastive Views for…Crafting Better Contrastive Views for Siamese Representation LearningOn the Importance ofAsymmetry for Siamese…On the Importance of Asymmetry for Siamese Representation LearningLarge-ScaleRepresentation Learning…Large-Scale Representation Learning on Graphs via BootstrappingUnderstanding Collapsein Non-contrastive…Understanding Collapse in Non-contrastive Siamese Representation LearningCan Pretext-BasedSelf-Supervised Learnin…Can Pretext-Based Self-Supervised Learning Be Boosted by Downstream Data? A Theoretical AnalysisHow Does SimSiam AvoidCollapse Without…How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive LearningThe Power of Contrastfor Feature Learning: A…The Power of Contrast for Feature Learning: A Theoretical AnalysisUnderstandingself-supervised Learnin…Understanding self-supervised Learning Dynamics without Contrastive PairsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.