No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models

The computation necessary for training Transformer-based language models has skyrocketed in recent years. This trend has motivated research on efficient training algorithms designed to improve training, validation, and downstream performance faster than standard training. In this work, we revisit three categories of such algorithms: dynamic architectures (layer stacking, layer dropping), batch selection (selective backprop, RHO loss), and efficient optimizers (Lion, Sophia). When pre-training BERT and T5 with a fixed computation budget using such methods, we find that their training, validation, and downstream gains vanish compared to a baseline with a fully-decayed learning rate. We define an evaluation protocol that enables computation to be done on arbitrary machines by mapping all computation time to a reference machine which we call reference system time. We discuss the limitations of our proposed protocol and release our code to encourage rigorous research in efficient training procedures: https://github.com/JeanKaddour/NoTrainNoGain.

Scaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsScaling Language Models:Methods, Analysis &…Scaling Language Models: Methods, Analysis & Insights from Training GopherTraining Compute-OptimalLarge Language ModelsTraining Compute-Optimal Large Language ModelsChain of ThoughtPrompting Elicits…Chain of Thought Prompting Elicits Reasoning in Large Language ModelsStop Wasting My Time!Saving Days of ImageNet…Stop Wasting My Time! Saving Days of ImageNet and BERT Training with Latest Weight AveragingThe MiniPile Challengefor Data-Efficient…The MiniPile Challenge for Data-Efficient Language ModelsCramming: Training aLanguage Model on a…Cramming: Training a Language Model on a Single GPU in One DayUnderstanding theEffectiveness of Early…Understanding the Effectiveness of Early Weight Averaging for Training Large Language ModelsDoReMi: Optimizing DataMixtures Speeds Up…DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingOn Efficient Training ofLarge-Scale Deep…On Efficient Training of Large-Scale Deep Learning Models: A Literature ReviewSophia: A ScalableStochastic Second-order…Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-trainingA Pretrainer's Guide toTraining Data: Measurin…A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & ToxicityChallenges andApplications of Large…Challenges and Applications of Large Language ModelsSparks of Large AudioModels: A Survey and…Sparks of Large Audio Models: A Survey and OutlookMosaicBERT: ABidirectional Encoder…MosaicBERT: A Bidirectional Encoder Optimized for Fast PretrainingThe Languini Kitchen:Enabling Language…The Languini Kitchen: Enabling Language Modelling Research at Different Scales of ComputeSheared LLaMA:Accelerating Language…Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningA Large-ScaleExploration of…A Large-Scale Exploration of μ-TransferButterfly Effects of SGDNoise: Error…Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and AutoregressionKnowledge Fusion ofLarge Language ModelsKnowledge Fusion of Large Language ModelsSOLAR 10.7B: ScalingLarge Language Models…SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-ScalingFantastic PretrainingOptimizers and Where to…Fantastic Pretraining Optimizers and Where to Find ThemMetadata ConditioningAccelerates Language…Metadata Conditioning Accelerates Language Model Pre-trainingAdaptive DataOptimization: Dynamic…Adaptive Data Optimization: Dynamic Sample Selection with Scaling LawsNo Train No Gain:Revisiting Efficient…No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.