Supervised word mover's distance

Recently, a new document metric called the word mover's distance (WMD) has been proposed with unprecedented results on kNN-based document classification. The WMD elevates high-quality word embeddings to a document metric by formulating the distance between two documents as an optimal transport problem between the embedded words. However, the document distances are entirely un-supervised and lack a mechanism to incorporate supervision when available. In this paper we propose an efficient technique to learn a supervised metric, which we call the Supervised-WMD (S-WMD) metric. The supervised training minimizes the stochastic leave-one-out nearest neighbor classification error on a per-document level by updating an affine transformation of the underlying word embedding space and a word-imporance weight vector. As the gradient of the original WMD distance would result in an inefficient nested optimization problem, we provide an arbitrarily close approximation that results in a practical and efficient update rule. We evaluate S-WMD on eight real-world text classification tasks on which it consistently outperforms almost all of our 26 competitive baselines.

Term-WeightingApproaches in Automatic…Term-Weighting Approaches in Automatic Text RetrievalIndexing by LatentSemantic AnalysisIndexing by Latent Semantic AnalysisLatent DirichletAllocationLatent Dirichlet AllocationInformation-theoreticmetric learningInformation-theoretic metric learningVisualizing Data usingt-SNEVisualizing Data using t-SNEDistance Metric Learningfor Large Margin Neares…Distance Metric Learning for Large Margin Nearest Neighbor ClassificationSupervised Earth Mover'sDistance Learning and…Supervised Earth Mover's Distance Learning and Its Computer Vision ApplicationsGround Metric LearningGround Metric LearningLearningSentiment-Specific Word…Learning Sentiment-Specific Word Embedding for Twitter Sentiment ClassificationNeural Word Embedding asImplicit Matrix…Neural Word Embedding as Implicit Matrix FactorizationFrom Word Embeddings ToDocument DistancesFrom Word Embeddings To Document DistancesLearning with aWasserstein LossLearning with a Wasserstein LossEarth Mover's DistanceMinimization for…Earth Mover's Distance Minimization for Unsupervised Bilingual Lexicon InductionWord Mover's Embedding:From Word2Vec to…Word Mover's Embedding: From Word2Vec to Document EmbeddingContext Mover's Distance& Barycenters: Optimal…Context Mover's Distance & Barycenters: Optimal Transport of Contexts for Building RepresentationsHierarchical OptimalTransport for Document…Hierarchical Optimal Transport for Document RepresentationClassifying ExtremelyShort Texts by…Classifying Extremely Short Texts by Exploiting Semantic Centroids in Word Mover's Distance SpaceSentence Mover'sSimilarity: Automatic…Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence TextsComputational OptimalTransportComputational Optimal TransportSemantics-assistedWasserstein Learning fo…Semantics-assisted Wasserstein Learning for Topic and Word EmbeddingsMeasurement of TextSimilarity: A SurveyMeasurement of Text Similarity: A SurveyWMDecompose: A Frameworkfor Leveraging the…WMDecompose: A Framework for Leveraging the Interpretable Properties of Word Mover's Distance in Sociocultural AnalysisOutlier-Robust OptimalTransportOutlier-Robust Optimal TransportRe-evaluating WordMover's DistanceRe-evaluating Word Mover's DistanceSupervised word mover'sdistanceSupervised word mover's distanceEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.