MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance

A robust evaluation metric has a profound impact on the development of text generation systems. A desirable metric compares system output against references based on their semantics rather than surface forms. In this paper we investigate strategies to encode system and reference texts to devise a metric that shows a high correlation with human judgment of text quality. We validate our new metric, namely MoverScore, on a number of text generation tasks including summarization, machine translation, image captioning, and data-to-text generation, where the outputs are produced by a variety of neural and non-neural systems. Our findings suggest that metrics combining contextualized representations with a distance measure perform the best. Such metrics also demonstrate strong generalization capability across tasks. For ease-of-use we make our metrics available as web service.

Evaluating ContentSelection in…Evaluating Content Selection in Summarization: The Pyramid MethodBetter SummarizationEvaluation with Word…Better Summarization Evaluation with Word Embeddings for ROUGECIDEr: Consensus-basedImage Description…CIDEr: Consensus-based Image Description EvaluationLearning to Score SystemSummaries for Better…Learning to Score System Summaries for Better Content Selection EvaluationMEANT 2.0: Accuratesemantic MT evaluation…MEANT 2.0: Accurate semantic MT evaluation for any output languageResults of the WMT17Metrics Shared TaskResults of the WMT17 Metrics Shared TaskRUSE: Regressor UsingSentence Embeddings for…RUSE: Regressor Using Sentence Embeddings for Automatic Machine Translation EvaluationThe price of debiasingautomatic metrics in…The price of debiasing automatic metrics in natural language evalautionSentence Mover'sSimilarity: Automatic…Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence TextsBetter Rewards YieldBetter Summaries…Better Rewards Yield Better Summaries: Learning to Summarise Without ReferencesPutting Evaluation inContext: Contextual…Putting Evaluation in Context: Contextual Embeddings Improve Machine Translation EvaluationBERTScore: EvaluatingText Generation with…BERTScore: Evaluating Text Generation with BERTFill in the BLANC:Human-free quality…Fill in the BLANC: Human-free quality estimation of document summariesAn Anchor-BasedAutomatic Evaluation…An Anchor-Based Automatic Evaluation Metric for Document SummarizationBetter than Average:Paired Evaluation of NL…Better than Average: Paired Evaluation of NLP SystemsCSDS: A Fine-GrainedChinese Dataset for…CSDS: A Fine-Grained Chinese Dataset for Customer Service Dialogue SummarizationPlay the Shannon GameWith Language Models: A…Play the Shannon Game With Language Models: A Human-Free Approach to Summary EvaluationFrugalScore: LearningCheaper, Lighter and…FrugalScore: Learning Cheaper, Lighter and Faster Evaluation Metrics for Automatic Text GenerationReproducibility Issuesfor BERT-based…Reproducibility Issues for BERT-based Evaluation MetricsEvaluating EvaluationMetrics: A Framework fo…Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement TheoryDATScore: EvaluatingTranslation with Data…DATScore: Evaluating Translation with Data Augmented TranslationsIs ChatGPT a Good NLGEvaluator? A Preliminar…Is ChatGPT a Good NLG Evaluator? A Preliminary StudyG-Eval: NLG Evaluationusing GPT-4 with Better…G-Eval: NLG Evaluation using GPT-4 with Better Human AlignmentEnd to End UrduAbstractive Text…End to End Urdu Abstractive Text Summarization With Dataset and Improvement in Evaluation MetricMoverScore: TextGeneration Evaluating…MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。