A Theory for Emergence of Complex Skills in Language Models

A major driver of AI products today is the fact that new skills emerge in language models when their parameter set and training corpora are scaled up. This phenomenon is poorly understood, and a mechanistic explanation via mathematical analysis of gradient-based training seems difficult. The current paper takes a different approach, analysing emergence using the famous (and empirical) Scaling Laws of LLMs and a simple statistical framework. Contributions include: (a) A statistical framework that relates cross-entropy loss of LLMs to competence on the basic skills that underlie language tasks. (b) Mathematical analysis showing that the Scaling Laws imply a strong form of inductive bias that allows the pre-trained model to learn very efficiently. We informally call this {\em slingshot generalization} since naively viewed it appears to give competence levels at skills that violate usual generalization theory. (c) A key example of slingshot generalization, that competence at executing tasks involving $k$-tuples of skills emerges essentially at the same scaling and same rate as competence on the elementary skills themselves.

Deep Learning Scaling isPredictable, EmpiricallyDeep Learning Scaling is Predictable, EmpiricallyScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsEmergent Abilities ofLarge Language ModelsEmergent Abilities of Large Language ModelsTraining Compute-OptimalLarge Language ModelsTraining Compute-Optimal Large Language ModelsPaLM 2 Technical ReportPaLM 2 Technical ReportPaLM: Scaling LanguageModeling with PathwaysPaLM: Scaling Language Modeling with PathwaysGPT-4 Technical ReportGPT-4 Technical ReportBeyond the ImitationGame: Quantifying and…Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language modelsExplaining NeuralScaling LawsExplaining Neural Scaling LawsThe Quantization Modelof Neural ScalingThe Quantization Model of Neural ScalingCompositional AbilitiesEmerge Multiplicatively…Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskA Theoretical Insightinto Attack and Defense…A Theoretical Insight into Attack and Defense of Gradient Leakage in TransformerA Dynamical Model ofNeural Scaling LawsA Dynamical Model of Neural Scaling LawsAn exactly solvablemodel for emergence and…An exactly solvable model for emergence and scaling laws in the multitask sparse parity problemSKILL-MIX: a Flexibleand Expandable Family o…SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI ModelsTowards a theory of howthe structure of…Towards a theory of how the structure of language is acquired by deep neural networksCompositionalCapabilities of…Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable TasksA Percolation Model ofEmergence: Analyzing…A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal LanguageSkill-Targeted AdaptiveTrainingSkill-Targeted Adaptive TrainingAnalyzing Neural ScalingLaws in Two-Layer…Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data SpectraScaling andrenormalization in…Scaling and renormalization in high-dimensional regressionA Theory for Emergenceof Complex Skills in…A Theory for Emergence of Complex Skills in Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.