Energy and Policy Considerations for Deep Learning in NLP

Recent progress in hardware and methodology for training neural networks has ushered in a new generation of large networks trained on abundant data. These models have obtained notable gains in accuracy across many NLP tasks. However, these accuracy improvements depend on the availability of exceptionally large computational resources that necessitate similarly substantial energy consumption. As a result these models are costly to train and develop, both financially, due to the cost of hardware and electricity or cloud compute time, and environmentally, due to the carbon footprint required to fuel modern tensor processing hardware. In this paper we bring this issue to the attention of NLP researchers by quantifying the approximate financial and environmental costs of training a variety of recently successful neural network models for NLP. Based on these findings, we propose actionable recommendations to reduce costs and improve equity in NLP research and practice.

Algorithms forhyper-parameter…Algorithms for hyper-parameter optimizationPractical BayesianOptimization of Machine…Practical Bayesian Optimization of Machine Learning AlgorithmsRandom Search forHyper-Parameter…Random Search for Hyper-Parameter OptimizationEffective Approaches toAttention-based Neural…Effective Approaches to Attention-based Neural Machine TranslationAn Analysis of DeepNeural Network Models…An Analysis of Deep Neural Network Models for Practical ApplicationsEvaluating the EnergyEfficiency of Deep…Evaluating the Energy Efficiency of Deep Convolutional Neural Networks on CPUs and GPUsBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingThe Evolved TransformerThe Evolved TransformerAccelerating Sparse DNNModels without…Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise SparsityCarbon Emissions andLarge Neural Network…Carbon Emissions and Large Neural Network TrainingSparse Training viaBoosting Pruning…Sparse Training via Boosting Pruning Plasticity with NeuroregenerationRevisiting the TrainLoss: an Efficient…Revisiting the Train Loss: an Efficient Performance Estimator for Neural Architecture SearchImproving ClassifierTraining Efficiency for…Improving Classifier Training Efficiency for Automatic Cyberbullying Detection with Feature DensityChallenges in DeployingMachine Learning: A…Challenges in Deploying Machine Learning: A Survey of Case StudiesGLaM: Efficient Scalingof Language Models with…GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsDoReMi: Optimizing DataMixtures Speeds Up…DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingCompute-Efficient DeepLearning: Algorithmic…Compute-Efficient Deep Learning: Algorithmic Trends and OpportunitiesPower Hungry Processing:Watts Driving the Cost…Power Hungry Processing: Watts Driving the Cost of AI Deployment?Dendrites endowartificial neural…Dendrites endow artificial neural networks with accurate, robust and parameter-efficient learningDon't be lazy: CompletePenables…Don't be lazy: CompleteP enables compute-efficient deep transformersEnergy and PolicyConsiderations for Deep…Energy and Policy Considerations for Deep Learning in NLP過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。