With Little Power Comes Great Responsibility

Despite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community. Underpowered experiments make it more difficult to discern the difference between statistical noise and meaningful model improvements, and increase the chances of exaggerated findings. By metaanalyzing a set of existing NLP papers and datasets, we characterize typical power for a variety of settings and conclude that underpowered experiments are common in the NLP literature. In particular, for several tasks in the popular GLUE benchmark, small test sets mean that most attempted comparisons to state of the art models will not be adequately powered. Similarly, based on reasonable assumptions, we find that the most typical experimental design for human rating studies will be underpowered to detect small model differences, of the sort that are frequently studied. For machine translation, we find that typical test sets of 2000 sentences have approximately 75% power to detect differences of 1 BLEU point. To improve the situation going forward, we give an overview of best practices for power analysis in NLP and release a series of notebooks to assist with future power analyses. 1

Statistical Comparisonsof Classifiers over…Statistical Comparisons of Classifiers over Multiple Data SetsRandom effects structurefor confirmatory…Random effects structure for confirmatory hypothesis testing: Keep it maximalWhat's in a p-value inNLP?What's in a p-value in NLP?Sentence Encoders onSTILTs: Supplementary…Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data TasksThe Hitchhiker's Guideto Testing Statistical…The Hitchhiker's Guide to Testing Statistical Significance in Natural Language ProcessingTransformer-BasedFeature Learning for…Transformer-Based Feature Learning for Algorithm Selection in Combinatorial OptimisationBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingShow Your Work: ImprovedReporting of…Show Your Work: Improved Reporting of Experimental ResultsBest practices for thehuman evaluation of…Best practices for the human evaluation of automatically generated textExploring the Limits ofTransfer Learning with…Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerBERTs of a feather donot generalize together…BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performanceGreen AIGreen AIFLEX: UnifyingEvaluation for Few-Shot…FLEX: Unifying Evaluation for Few-Shot NLPAutomatic TextEvaluation through the…Automatic Text Evaluation through the Lens of Wasserstein BarycentersChallenges for cognitivedecoding using deep…Challenges for cognitive decoding using deep learning methodsThe MultiBERTs: BERTReproductions for…The MultiBERTs: BERT Reproductions for Robustness AnalysisThe Worst of BothWorlds: A Comparative…The Worst of Both Worlds: A Comparative Analysis of Errors in Learning from Data in Psychology and Machine Learningdeep-significance - Easyand Meaningful…deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural NetworksGuiding Large LanguageModels via Directional…Guiding Large Language Models via Directional Stimulus PromptingCan Small and SyntheticBenchmarks Drive…Can Small and Synthetic Benchmarks Drive Modeling Innovation? A Retrospective Study of Question Answering Modeling ApproachesDocument-Level MachineTranslation with Large…Document-Level Machine Translation with Large Language ModelsAdding Error Bars toEvals: A Statistical…Adding Error Bars to Evals: A Statistical Approach to Language Model EvaluationsDéjà Vu: MultilingualLLM Evaluation through…Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation EvaluationA Sober Look at Progressin Language Model…A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to ReproducibilityWith Little Power ComesGreat ResponsibilityWith Little Power Comes Great ResponsibilityEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.