Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering. In this publication, we present Sentence-BERT (SBERT), a modification of the pretrained BERT network that use siamese and triplet network structures to derive semantically meaningful sentence embeddings that can be compared using cosine-similarity. This reduces the effort for finding the most similar pair from 65 hours with BERT / RoBERTa to about 5 seconds with SBERT, while maintaining the accuracy from BERT. We evaluate SBERT and SRoBERTa on common STS tasks and transfer learning tasks, where it outperforms other state-of-the-art sentence embeddings methods.

Glove: Global Vectorsfor Word RepresentationGlove: Global Vectors for Word RepresentationA large annotated corpusfor learning natural…A large annotated corpus for learning natural language inferenceFaceNet: A UnifiedEmbedding for Face…FaceNet: A Unified Embedding for Face Recognition and ClusteringSemEval-2016 Task 1:Semantic Textual…SemEval-2016 Task 1: Semantic Textual Similarity, Monolingual and Cross-Lingual EvaluationSentEval: An EvaluationToolkit for Universal…SentEval: An Evaluation Toolkit for Universal Sentence RepresentationsQualitative SpatialReasoning over Question…Qualitative Spatial Reasoning over Questions (Short Paper)Universal SentenceEncoderUniversal Sentence EncoderTransformer-BasedFeature Learning for…Transformer-Based Feature Learning for Algorithm Selection in Combinatorial OptimisationBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingXLNet: GeneralizedAutoregressive…XLNet: Generalized Autoregressive Pretraining for Language UnderstandingRoBERTa: A RobustlyOptimized BERT…RoBERTa: A Robustly Optimized BERT Pretraining ApproachBERTScore: EvaluatingText Generation with…BERTScore: Evaluating Text Generation with BERTA SemanticallyConsistent and…A Semantically Consistent and Syntactically Variational Encoder-Decoder Framework for Paraphrase GenerationSelf-Guided ContrastiveLearning for BERT…Self-Guided Contrastive Learning for BERT Sentence RepresentationsExploring the Role ofBERT Token…Exploring the Role of BERT Token Representations to Explain Sentence Probing ResultsDistilling RelationEmbeddings from…Distilling Relation Embeddings from Pretrained Language ModelsPractical Cross-ModalManifold Alignment for…Practical Cross-Modal Manifold Alignment for Robotic Grounded Language LearningExtracting LatentSteering Vectors from…Extracting Latent Steering Vectors from Pretrained Language ModelsSpinning LanguageModels: Risks of…Spinning Language Models: Risks of Propaganda-As-A-Service and CountermeasuresExposing the Limits ofVideo-Text Models…Exposing the Limits of Video-Text Models through Contrast SetsGlobal evidence ofexpressed sentiment…Global evidence of expressed sentiment alterations during the COVID-19 pandemicSiamese BERT-Based Modelfor Web Search Relevanc…Siamese BERT-Based Model for Web Search Relevance Ranking Evaluated on a New Czech DatasetToolLLM: FacilitatingLarge Language Models t…ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsFine-TuningRetrieval-Augmented…Fine-Tuning Retrieval-Augmented Generation with an Auto-Regressive Language Model for Sentiment Analysis in Financial ReviewsSentence-BERT: SentenceEmbeddings using Siames…Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。