VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

Recent self-supervised methods for image representation learning are based on maximizing the agreement between embedding vectors from different views of the same image. A trivial solution is obtained when the encoder outputs constant vectors. This collapse problem is often avoided through implicit biases in the learning architecture, that often lack a clear justification or interpretation. In this paper, we introduce VICReg (Variance-Invariance-Covariance Regularization), a method that explicitly avoids the collapse problem with a simple regularization term on the variance of the embeddings along each dimension individually. VICReg combines the variance term with a decorrelation mechanism based on redundancy reduction and covariance regularization, and achieves results on par with the state of the art on several downstream tasks. In addition, we show that incorporating our new variance term into other methods helps stabilize the training and leads to performance improvements.

Faster R-CNN: TowardsReal-Time Object…Faster R-CNN: Towards Real-Time Object Detection with Region Proposal NetworksDeep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionAggregated ResidualTransformations for Dee…Aggregated Residual Transformations for Deep Neural NetworksUnsupervised FeatureLearning via…Unsupervised Feature Learning via Non-Parametric Instance DiscriminationA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsImproved Baselines withMomentum Contrastive…Improved Baselines with Momentum Contrastive LearningContrastive MultiviewCodingContrastive Multiview CodingBarlow Twins:Self-Supervised Learnin…Barlow Twins: Self-Supervised Learning via Redundancy ReductionPrototypical ContrastiveLearning of Unsupervise…Prototypical Contrastive Learning of Unsupervised RepresentationsWhitening forSelf-Supervised…Whitening for Self-Supervised Representation LearningOnlineBag-of-Visual-Words…Online Bag-of-Visual-Words Generation for Unsupervised Representation LearningUnderstandingself-supervised Learnin…Understanding self-supervised Learning Dynamics without Contrastive PairsAre Large-scale DatasetsNecessary for…Are Large-scale Datasets Necessary for Self-Supervised Pre-training?VIbCReg:Variance-Invariance-bet…VIbCReg: Variance-Invariance-better-Covariance Regularization for Self-Supervised Learning on Time SeriesUnderstandingDimensional Collapse in…Understanding Dimensional Collapse in Contrastive Self-supervised LearningOn the Importance ofAsymmetry for Siamese…On the Importance of Asymmetry for Siamese Representation LearningDual Temperature HelpsContrastive Learning…Dual Temperature Helps Contrastive Learning Without Many Negative Samples: Towards Understanding and Simplifying MoCoVision Models Are MoreRobust And Fair When…Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without SupervisionMasked Siamese Networksfor Label-Efficient…Masked Siamese Networks for Label-Efficient LearningSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureUnderstanding MaskedImage Modeling via…Understanding Masked Image Modeling via Learning Occlusion Invariant FeatureMulti-Mode OnlineKnowledge Distillation…Multi-Mode Online Knowledge Distillation for Self-Supervised Visual Representation LearningVNE: An Effective Methodfor Improving Deep…VNE: An Effective Method for Improving Deep Representation by Manipulating Eigenvalue DistributionMeasuringSelf-Supervised…Measuring Self-Supervised Representation Quality for Downstream Classification Using Discriminative FeaturesVICReg:Variance-Invariance-Cov…VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.