Training Compute-Optimal Large Language Models

We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.

Adam: A Method forStochastic OptimizationAdam: A Method for Stochastic OptimizationTriviaQA: A Large ScaleDistantly Supervised…TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading ComprehensionSentencePiece: A simpleand language independen…SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text ProcessingScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsThe Pile: An 800GBDataset of Diverse Text…The Pile: An 800GB Dataset of Diverse Text for Language ModelingScaling Language Models:Methods, Analysis &…Scaling Language Models: Methods, Analysis & Insights from Training GopherMeasuring MassiveMultitask Language…Measuring Massive Multitask Language UnderstandingUsing DeepSpeed andMegatron to Train…Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language ModelTruthfulQA: MeasuringHow Models Mimic Human…TruthfulQA: Measuring How Models Mimic Human FalsehoodsUnified Scaling Laws forRouted Language ModelsUnified Scaling Laws for Routed Language ModelsSwitch Transformers:Scaling to Trillion…Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityImproving LanguageModels by Retrieving…Improving Language Models by Retrieving from Trillions of TokensNo Train No Gain:Revisiting Efficient…No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language ModelsReproducible ScalingLaws for Contrastive…Reproducible Scaling Laws for Contrastive Language-Image LearningPlanBench: An ExtensibleBenchmark for Evaluatin…PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about ChangeTo Repeat or Not ToRepeat: Insights from…To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisLarge Language ModelAlignment: A SurveyLarge Language Model Alignment: A SurveyThe Dawn of LMMs:Preliminary Exploration…The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)MAP-Neo: Highly Capableand Transparent…MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model SeriesNemotron-4 15B TechnicalReportNemotron-4 15B Technical ReportLarge Language Modelsfor Software…Large Language Models for Software Engineering: A Systematic Literature ReviewBASE TTS: Lessons frombuilding a…BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of dataLearning Multi-LevelFeatures with Matryoshk…Learning Multi-Level Features with Matryoshka Sparse AutoencodersTraining on the TestTask Confounds…Training on the Test Task Confounds Evaluation and EmergenceTraining Compute-OptimalLarge Language ModelsTraining Compute-Optimal Large Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.