Enriching Word Vectors with Subword Information

Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. This is a limitation, especially for languages with large vocabularies and many rare words. In this paper, we propose a new approach based on the skipgram model, where each word is represented as a bag of character $n$-grams. A vector representation is associated to each character $n$-gram; words being represented as the sum of these representations. Our method is fast, allowing to train models on large corpora quickly and allows us to compute word representations for words that did not appear in the training data. We evaluate our word representations on nine different languages, both on word similarity and analogy tasks. By comparing to recently proposed morphological word representations, we show that our vectors achieve state-of-the-art performance on these tasks.

openalex_id:w2962784628openalex_id:w2962784628Distributional StructureDistributional StructureFactored Neural LanguageModelsFactored Neural Language ModelsDistributedRepresentations of Word…Distributed Representations of Words and Phrases and their CompositionalityEfficient Estimation ofWord Representations in…Efficient Estimation of Word Representations in Vector SpaceBetter WordRepresentations with…Better Word Representations with Recursive Neural Networks for MorphologyCompositional Morphologyfor Word Representation…Compositional Morphology for Word Representations and Language ModellingCo-learning of WordRepresentations and…Co-learning of Word Representations and Morpheme RepresentationsUnsupervised MorphologyInduction Using Word…Unsupervised Morphology Induction Using Word EmbeddingsFinding Function inForm: Compositional…Finding Function in Form: Compositional Character Models for Open Vocabulary Word RepresentationCharagram: EmbeddingWords and Sentences via…Charagram: Embedding Words and Sentences via Character n-gramsCharacter-Aware NeuralLanguage ModelsCharacter-Aware Neural Language ModelsMimicking WordEmbeddings using Subwor…Mimicking Word Embeddings using Subword RNNsLearning to ComposeDomain-Specific…Learning to Compose Domain-Specific Transformations for Data AugmentationAnalogical Reasoning onChinese Morphological…Analogical Reasoning on Chinese Morphological and Semantic RelationsUnsupervisedMultilingual Word…Unsupervised Multilingual Word EmbeddingsPublicly AvailableClinicalPublicly Available ClinicalEnriching WordEmbeddings with Global…Enriching Word Embeddings with Global Information and Testing on Highly Inflected LanguageSECNLP: A survey ofembeddings in clinical…SECNLP: A survey of embeddings in clinical natural language processingTowards a real-timeprocessing framework…Towards a real-time processing framework based on improved distributed recurrent neural network variants with fastText for social big data analyticsAre We ConsistentlyBiased? Multidimensiona…Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word VectorsSimAlign: High QualityWord Alignments without…SimAlign: High Quality Word Alignments without Parallel Training Data using Static and Contextualized EmbeddingsSurvey of Low-ResourceMachine TranslationSurvey of Low-Resource Machine TranslationPerformance Evaluationof Vector Embeddings…Performance Evaluation of Vector Embeddings with Retrieval-Augmented GenerationEnriching Word Vectorswith Subword InformationEnriching Word Vectors with Subword Information過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。