A Theoretical Analysis of Contrastive Unsupervised Representation Learning

Recent empirical works have successfully used unlabeled data to learn feature representations that are broadly useful in downstream classification tasks. Several of these methods are reminiscent of the well-known word2vec embedding algorithm: leveraging availability of pairs of semantically "similar" data points and "negative samples," the learner forces the inner product of representations of similar pairs with each other to be higher on average than with negative samples. The current paper uses the term contrastive learning for such algorithms and presents a theoretical framework for analyzing them by introducing latent classes and hypothesizing that semantically similar points are sampled from the same latent class. This framework allows us to show provable guarantees on the performance of the learned representations on the average classification task that is comprised of a subset of the same set of latent classes. Our generalization bound also shows that learned representations can reduce (labeled) sample complexity on downstream tasks. We conduct controlled experiments in both the text and image domains to support the theory.

Combining Labeled andUnlabeled Data with…Combining Labeled and Unlabeled Data with Co-TrainingNoise-contrastiveestimation: A new…Noise-contrastive estimation: A new estimation principle for unnormalized statistical modelsLearning Word Vectorsfor Sentiment AnalysisLearning Word Vectors for Sentiment AnalysisFoundations of MachineLearningFoundations of Machine LearningDistributedRepresentations of Word…Distributed Representations of Words and Phrases and their CompositionalityGlove: Global Vectorsfor Word RepresentationGlove: Global Vectors for Word RepresentationVery Deep ConvolutionalNetworks for Large-Scal…Very Deep Convolutional Networks for Large-Scale Image RecognitionUnsupervised Learning ofVisual Representations…Unsupervised Learning of Visual Representations Using VideosPrototypical Networksfor Few-shot LearningPrototypical Networks for Few-shot LearningUnsupervised Learning ofSentence Embeddings…Unsupervised Learning of Sentence Embeddings Using Compositional n-Gram FeaturesBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingLearning Multiple Layersof Features from Tiny…Learning Multiple Layers of Features from Tiny ImagesProvable Guarantees forGradient-Based…Provable Guarantees for Gradient-Based Meta-LearningInformation Leakage inEmbedding ModelsInformation Leakage in Embedding ModelsProvable Guarantees forSelf-Supervised Deep…Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossTheoretical Analysis ofSelf-Training with Deep…Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataSupervised ContrastiveLearning for Pre-traine…Supervised Contrastive Learning for Pre-trained Language Model Fine-tuningFor self-supervisedlearning, Rationality…For self-supervised learning, Rationality implies generalization, provablyPre-training MolecularGraph Representation…Pre-training Molecular Graph Representation with 3D GeometryUnsupervisedRepresentation Learning…Unsupervised Representation Learning for Time Series with Temporal Neighborhood CodingDivide and Contrast:Self-supervised Learnin…Divide and Contrast: Self-supervised Learning from Uncurated DataSelf-SupervisedLearning: Generative or…Self-Supervised Learning: Generative or ContrastiveSelf-SupervisedRepresentation Learning…Self-Supervised Representation Learning: Introduction, advances, and challengesRecent advances andclinical applications o…Recent advances and clinical applications of deep learning in medical image analysisA Theoretical Analysisof Contrastive…A Theoretical Analysis of Contrastive Unsupervised Representation Learning過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。