Investigating Continual Pretraining in Large Language Models: Insights and Implications

Continual learning (CL) in large language models (LLMs) is an evolving domain that focuses on developing efficient and sustainable training strategies to adapt models to emerging knowledge and achieve robustness in dynamic environments. Our primary emphasis is on continual domain-adaptive pretraining, a process designed to equip LLMs with the ability to integrate new information from various domains while retaining previously learned knowledge. Since existing works concentrate mostly on continual fine-tuning for a limited selection of downstream tasks or training domains, we introduce a new benchmark designed to measure the adaptability of LLMs to changing pretraining data landscapes. We further examine the impact of model size on learning efficacy and forgetting, as well as how the progression and similarity of emerging domains affect the knowledge transfer within these models. Our findings uncover several key insights: (i) continual pretraining consistently improves <1.5B models studied in this work and is also superior to domain adaptation, (ii) larger models always achieve better perplexity than smaller ones when continually pretrained on the same corpus, (iii) smaller models are particularly sensitive to continual pretraining, showing the most significant rates of both learning and forgetting, (iv) continual pretraining boosts downstream task performance of GPT-2 family, (v) continual pretraining enables LLMs to specialize better when the sequence of domains shows semantic similarity while randomizing training domains leads to better transfer and final performance otherwise. We posit that our research establishes a new benchmark for CL in LLMs, providing a more realistic evaluation of knowledge retention and transfer across diverse domains.

SQuAD: 100, 000+Questions for Machine…SQuAD: 100, 000+ Questions for Machine Comprehension of TextBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingSentence-BERT: SentenceEmbeddings using Siames…Sentence-BERT: Sentence Embeddings using Siamese BERT-NetworksBERT Post-Training forReview Reading…BERT Post-Training for Review Reading Comprehension and Aspect-based Sentiment AnalysisScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsDon't Stop Pretraining:Adapt Language Models t…Don't Stop Pretraining: Adapt Language Models to Domains and TasksLifelong Pretraining:Continually Adapting…Lifelong Pretraining: Continually Adapting Language Models to Emerging CorporaContinual Pre-Trainingof Large Language…Continual Pre-Training of Large Language Models: How to (re)warm your model?Progressive Prompts:Continual Learning for…Progressive Prompts: Continual Learning for Language ModelsOrthogonal SubspaceLearning for Language…Orthogonal Subspace Learning for Language Model Continual LearningContinual Learning forLarge Language Models…Continual Learning for Large Language Models: A SurveyAn Empirical Study ofCatastrophic Forgetting…An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuningReuse, Don't Retrain: ARecipe for Continued…Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language ModelsEfficient ContinualPre-training by…Efficient Continual Pre-training by Mitigating the Stability GapCMR Scaling Law:Predicting Critical…CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsAurora-M: The First OpenSource Multilingual…Aurora-M: The First Open Source Multilingual Language Model Red-teamed according to the U.S. Executive OrderThe Future of ContinualLearning in the Era of…The Future of Continual Learning in the Era of Foundation Models: Three Key DirectionsContinual Learning ofLarge Language Models…Continual Learning of Large Language Models: A Comprehensive SurveyEnhancingDomain-Specific Encoder…Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without ThemTowards LifelongLearning of Large…Towards Lifelong Learning of Large Language Models: A SurveyScaling Agents viaContinual Pre-trainingScaling Agents via Continual Pre-trainingFrom Tokens to Words: Onthe Inner Lexicon of…From Tokens to Words: On the Inner Lexicon of LLMsBeyond Benchmarks: ANovel Framework for…Beyond Benchmarks: A Novel Framework for Domain-Specific LLM Evaluation and Knowledge MappingTele-LLMs: A Series ofSpecialized Large…Tele-LLMs: A Series of Specialized Large Language Models for TelecommunicationsInvestigating ContinualPretraining in Large…Investigating Continual Pretraining in Large Language Models: Insights and Implications過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。