Learning with a Wasserstein Loss

Learning to predict multi-label outputs is challenging, but in many problems there is a natural metric on the outputs that can be used to improve predictions. In this paper we develop a loss function for multi-label learning, based on the Wasserstein distance. The Wasserstein distance provides a natural notion of dissimilarity for probability measures. Although optimizing with respect to the exact Wasserstein distance is costly, recent work has described a regularized approximation that is efficiently computed. We describe an efficient learning algorithm based on this regularization, as well as a novel extension of the Wasserstein distance from probability measures to unnormalized measures. We also describe a statistical learning bound for the loss. The Wasserstein loss can encourage smoothness of the predictions with respect to a chosen metric on the output space. We demonstrate this property on a real-data tag prediction problem, using the Yahoo Flickr Creative Commons dataset, outperforming a baseline that doesn't use the metric.

Nonlinear totalvariation based noise…Nonlinear total variation based noise removal algorithmsIntroduction to linearoptimizationIntroduction to linear optimizationThe Earth Mover'sDistance as a Metric fo…The Earth Mover's Distance as a Metric for Image RetrievalRademacher and GaussianComplexities: Risk…Rademacher and Gaussian Complexities: Risk Bounds and Structural ResultsApproximate earthmover's distance in…Approximate earth mover's distance in linear timeOptimal Transport: Oldand NewOptimal Transport: Old and NewFast and robust EarthMover's DistancesFast and robust Earth Mover's DistancesDistributedRepresentations of Word…Distributed Representations of Words and Phrases and their CompositionalityWasserstein Propagationfor Semi-Supervised…Wasserstein Propagation for Semi-Supervised LearningFast Computation ofWasserstein BarycentersFast Computation of Wasserstein BarycentersImageNet Large ScaleVisual Recognition…ImageNet Large Scale Visual Recognition ChallengeFully ConvolutionalNetworks for Semantic…Fully Convolutional Networks for Semantic SegmentationNon-Negative MatrixFactorization with…Non-Negative Matrix Factorization with Sinkhorn DistanceMapping Estimation forDiscrete Optimal…Mapping Estimation for Discrete Optimal TransportSupervised word mover'sdistanceSupervised word mover's distanceSquared Earth Mover'sDistance-based Loss for…Squared Earth Mover's Distance-based Loss for Training Deep Neural NetworksGromov-WassersteinAveraging of Kernel and…Gromov-Wasserstein Averaging of Kernel and Distance MatricesScaling algorithms forunbalanced optimal…Scaling algorithms for unbalanced optimal transport problemsA Transportation LpDistance for Signal…A Transportation Lp Distance for Signal AnalysisLabel DistributionLearning by Optimal…Label Distribution Learning by Optimal TransportWasserstein DiscriminantAnalysisWasserstein Discriminant AnalysisUnbalanced optimaltransport: Dynamic and…Unbalanced optimal transport: Dynamic and Kantorovich formulationsSmooth and SparseOptimal TransportSmooth and Sparse Optimal TransportComplex ObjectClassification: A…Complex Object Classification: A Multi-Modal Multi-Instance Multi-Label Deep Network with Optimal TransportLearning with aWasserstein LossLearning with a Wasserstein Loss過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。