A Simple but Tough-to-Beat Baseline for Sentence Embeddings

The success of neural network methods for computing word embeddings has motivated methods for generating semantic embeddings of longer pieces of text, such as sentences and paragraphs. Surprisingly, Wieting et al (ICLR'16) showed that such complicated methods are outperformed, especially in out-of-domain (transfer learning) settings, by simpler methods involving mild retraining of word embeddings and basic linear regression. The method of Wieting et al. requires retraining with a substantial labeled dataset such as Paraphrase Database (Ganitkevitch et al., 2013). The current paper goes further, showing that the following completely unsupervised sentence embedding is a formidable baseline: Use word embeddings computed using one of the popular methods on unlabeled corpus like Wikipedia, represent the sentence by a weighted average of the word vectors, and then modify them a bit using PCA/SVD. This weighting improves performance by about 10% to 30% in textual similarity tasks, and beats sophisticated supervised methods including RNN's and LSTM's. It even improves Wieting et al.'s embeddings. This simple method should be used as the baseline to beat in future, especially when labeled training data is scarce or nonexistent. The paper also gives a theoretical explanation of the success of the above unsupervised method using a latent variable generative model for sentences, which is a simple extension of the model in Arora et al. (TACL'16) with new smoothing terms that allow for words occurring out of context, as well as high probabilities for words like and, not in all contexts.

Unsupervised Learning ofSentence Representation…Unsupervised Learning of Sentence Representations using Convolutional Neural NetworksUnsupervised Learning ofSentence Embeddings…Unsupervised Learning of Sentence Embeddings Using Compositional n-Gram FeaturesAll-but-the-Top: Simpleand Effective…All-but-the-Top: Simple and Effective Postprocessing for Word RepresentationsXNLI: EvaluatingCross-lingual Sentence…XNLI: Evaluating Cross-lingual Sentence RepresentationsExploring SemanticProperties of Sentence…Exploring Semantic Properties of Sentence EmbeddingsHierarchical OptimalTransport for Document…Hierarchical Optimal Transport for Document RepresentationA Review of AutomatedSpeech and Language…A Review of Automated Speech and Language Features for Assessment of Cognitive and Thought DisordersWhat makes a goodconversation? How…What makes a good conversation? How controllable attributes affect human judgmentsQuizBot: ADialogue-based Adaptive…QuizBot: A Dialogue-based Adaptive Learning System for Factual KnowledgeTransparency andaccountability in AI…Transparency and accountability in AI decision support: Explaining and visualizing convolutional neural networks for text informationMGAT: Multimodal GraphAttention Network for…MGAT: Multimodal Graph Attention Network for RecommendationFew-shot TextClassification with…Few-shot Text Classification with Distributional SignaturesA Simple butTough-to-Beat Baseline…A Simple but Tough-to-Beat Baseline for Sentence Embeddings中心の論文この論文を引用する論文古い新しい

カタログにはまだ過去の参考文献がありません。

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。