Hierarchical Optimal Transport for Document Representation

© 2019 Neural information processing systems foundation. All rights reserved. The ability to measure similarity between documents enables intelligent summarization and analysis of large corpora. Past distances between documents suffer from either an inability to incorporate semantic similarities between words or from scalability issues. As an alternative, we introduce hierarchical optimal transport as a meta-distance between documents, where documents are modeled as distributions over topics, which themselves are modeled as distributions over words. We then solve an optimal transport problem on the smaller topic space to compute a similarity score. We give conditions on the topics under which this construction defines a distance, and we relate it to the word mover's distance. We evaluate our technique for k-NN classification and show better interpretability and scalability with comparable performance to current methods at a fraction of the cost.

The Hungarian Method forthe Assignment ProblemThe Hungarian Method for the Assignment ProblemIndexing by LatentSemantic AnalysisIndexing by Latent Semantic AnalysisLatent DirichletAllocationLatent Dirichlet AllocationHierarchical DirichletProcessesHierarchical Dirichlet ProcessesVisualizing Data usingt-SNEVisualizing Data using t-SNEScikit-learn: MachineLearning in PythonScikit-learn: Machine Learning in PythonDistributedRepresentations of Word…Distributed Representations of Words and Phrases and their CompositionalityStochastic VariationalInferenceStochastic Variational InferenceGlove: Global Vectorsfor Word RepresentationGlove: Global Vectors for Word RepresentationFrom Word Embeddings ToDocument DistancesFrom Word Embeddings To Document DistancesSupervised word mover'sdistanceSupervised word mover's distanceWord Mover's Embedding:From Word2Vec to…Word Mover's Embedding: From Word2Vec to Document EmbeddingHierarchical OptimalTransport for Multimoda…Hierarchical Optimal Transport for Multimodal Distribution AlignmentScalable NearestNeighbor Search for…Scalable Nearest Neighbor Search for Optimal TransportSemantics-assistedWasserstein Learning fo…Semantics-assisted Wasserstein Learning for Topic and Word EmbeddingsGeometric DatasetDistances via Optimal…Geometric Dataset Distances via Optimal TransportTransporting Labels viaHierarchical Optimal…Transporting Labels via Hierarchical Optimal Transport for Semi-Supervised LearningLearning Autoencoderswith Relational…Learning Autoencoders with Relational RegularizationOutlier-Robust OptimalTransportOutlier-Robust Optimal TransportInterpretablecontrastive word mover'…Interpretable contrastive word mover's embeddingDifferentiableHierarchical Optimal…Differentiable Hierarchical Optimal Transport for Robust Multi-View LearningEfficient OptimalTransport Algorithm by…Efficient Optimal Transport Algorithm by Accelerated Gradient DescentHOTNAS: HierarchicalOptimal Transport for…HOTNAS: Hierarchical Optimal Transport for Neural Architecture SearchEarth Movers in The BigData Era: A Review of…Earth Movers in The Big Data Era: A Review of Optimal Transport in Machine LearningHierarchical OptimalTransport for Document…Hierarchical Optimal Transport for Document Representation過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。