The Hitchhiker's Guide to Testing Statistical Significance in Natural Language Processing

Statistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental. In this opinion/theoretical paper we discuss the role of statistical significance testing in Natural Language Processing (NLP) research. We establish the fundamental concepts of significance testing and discuss the specific aspects of NLP tasks, experimental setups and evaluation measures that affect the choice of significance tests in NLP research. Based on this discussion, we propose a simple practical protocol for statistical significance test selection in NLP setups and accompany this protocol with a brief survey of the most relevant tests. We then survey recent empirical papers published in ACL and TACL during 2017 and show that while our community assigns great value to experimental results, statistical significance testing is often ignored or misused. We conclude with a brief discussion of open issues that should be properly addressed so that this important tool can be applied in NLP research in a statistically sound manner 1 .

Note on the SamplingError of the Difference…Note on the Sampling Error of the Difference Between Correlated Proportions or PercentagesAn Analysis of VarianceTest for Normality…An Analysis of Variance Test for Normality (Complete Samples)An Introduction to theBootstrapAn Introduction to the BootstrapAn Introduction to theBootstrapAn Introduction to the BootstrapBleu: a Method forAutomatic Evaluation of…Bleu: a Method for Automatic Evaluation of Machine TranslationStatistical SignificanceTests for Machine…Statistical Significance Tests for Machine Translation EvaluationROUGE: A Package forAutomatic Evaluation of…ROUGE: A Package for Automatic Evaluation of SummariesOn Some Pitfalls inAutomatic Evaluation an…On Some Pitfalls in Automatic Evaluation and Significance Testing for MTMETEOR: An AutomaticMetric for MT Evaluatio…METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human JudgmentsEuroparl: A ParallelCorpus for Statistical…Europarl: A Parallel Corpus for Statistical Machine TranslationAn EmpiricalInvestigation of…An Empirical Investigation of Statistical Significance in NLPWhat's in a p-value inNLP?What's in a p-value in NLP?Bayes Test of Precision,Recall, and F1 Measure…Bayes Test of Precision, Recall, and F1 Measure for Comparison of Two Natural Language Processing ModelsBest practices for thehuman evaluation of…Best practices for the human evaluation of automatically generated textA Semi-supervisedApproach to Generate th…A Semi-supervised Approach to Generate the Code-Mixed Text using Pre-trained Encoder and Transfer LearningMixText:Linguistically-Informed…MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationFact or Fiction:Verifying Scientific…Fact or Fiction: Verifying Scientific ClaimsReinforced Multi-taskApproach for Multi-hop…Reinforced Multi-task Approach for Multi-hop Question GenerationTo Ship or Not to Ship:An Extensive Evaluation…To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine TranslationEvaluation Examples arenot Equally Informative…Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?What happens if youtreat ordinal ratings a…What happens if you treat ordinal ratings as interval data? Human evaluations in NLP are even more under-powered than you thinkBetter than Average:Paired Evaluation of NL…Better than Average: Paired Evaluation of NLP SystemsOf Human Criteria andAutomatic Metrics: A…Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story Generationdeep-significance - Easyand Meaningful…deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural NetworksThe Hitchhiker's Guideto Testing Statistical…The Hitchhiker's Guide to Testing Statistical Significance in Natural Language ProcessingEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.