Exploring Simple Siamese Representation Learning

Siamese networks have become a common structure in various recent models for unsupervised visual representation learning. These models maximize the similarity between two augmentations of one image, subject to certain conditions for avoiding collapsing solutions. In this paper, we report surprising empirical results that simple Siamese networks can learn meaningful representations even using none of the following: (i) negative sample pairs, (ii) large batches, (iii) momentum encoders. Our experiments show that collapsing solutions do exist for the loss and structure, but a stop-gradient operation plays an essential role in preventing collapsing. We provide a hypothesis on the implication of stop-gradient, and further show proof-of-concept experiments verifying it. Our "SimSiam" method achieves competitive results on ImageNet and downstream tasks. We hope this simple baseline will motivate people to rethink the roles of Siamese architectures for unsupervised representation learning. Code is made available. 1

Deep Residual Learningfor Image RecognitionDeep Residual Learning for Image RecognitionUnsupervised FeatureLearning via…Unsupervised Feature Learning via Non-Parametric Instance DiscriminationLearning Representationsby Maximizing Mutual…Learning Representations by Maximizing Mutual Information Across ViewsData-Efficient ImageRecognition with…Data-Efficient Image Recognition with Contrastive Predictive CodingMomentum Contrast forUnsupervised Visual…Momentum Contrast for Unsupervised Visual Representation LearningUnsupervised Learning ofVisual Features by…Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsA Simple Framework forContrastive Learning of…A Simple Framework for Contrastive Learning of Visual RepresentationsBootstrap Your OwnLatent - A New Approach…Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningImproved Baselines withMomentum Contrastive…Improved Baselines with Momentum Contrastive LearningSelf-Supervised Learningof Pretext-Invariant…Self-Supervised Learning of Pretext-Invariant RepresentationsContrastive MultiviewCodingContrastive Multiview CodingSelf-labelling viasimultaneous clustering…Self-labelling via simultaneous clustering and representation learningUnderstandingself-supervised Learnin…Understanding self-supervised Learning Dynamics without Contrastive PairsUniMoCo: Unsupervised,Semi-Supervised and…UniMoCo: Unsupervised, Semi-Supervised and Full-Supervised Visual Representation LearningMean Shift forSelf-Supervised LearningMean Shift for Self-Supervised LearningAdCo: AdversarialContrast for Efficient…AdCo: Adversarial Contrast for Efficient Learning of Unsupervised Representations From Self-Trained Negative AdversariesEmerging Properties inSelf-Supervised Vision…Emerging Properties in Self-Supervised Vision TransformersMomentum^2 Teacher:Momentum Teacher with…Momentum^2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised LearningProvable Guarantees forSelf-Supervised Deep…Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossSelf-SupervisedLearning: Generative or…Self-Supervised Learning: Generative or ContrastiveContrastive Learning ofImage Representations…Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyLearning Where to Learnin Cross-View…Learning Where to Learn in Cross-View Self-Supervised LearningUnsupervisedRepresentation for…Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive LearningSelf-Supervised Learningfrom Images with a…Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureExploring Simple SiameseRepresentation LearningExploring Simple Siamese Representation Learning過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。