Progress measures for grokking via mechanistic interpretability

Neural networks often exhibit emergent behavior, where qualitatively new capabilities arise from scaling up the amount of parameters, training data, or training steps. One approach to understanding emergence is to find continuous \textit{progress measures} that underlie the seemingly discontinuous qualitative changes. We argue that progress measures can be found via mechanistic interpretability: reverse-engineering learned behaviors into their individual components. As a case study, we investigate the recently-discovered phenomenon of ``grokking'' exhibited by small transformers trained on modular addition tasks. We fully reverse engineer the algorithm learned by these networks, which uses discrete Fourier transforms and trigonometric identities to convert addition to rotation about a circle. We confirm the algorithm by analyzing the activations and weights and by performing ablations in Fourier space. Based on this understanding, we define progress measures that allow us to study the dynamics of training and split training into three continuous phases: memorization, circuit formation, and cleanup. Our results show that grokking, rather than being a sudden shift, arises from the gradual amplification of structured mechanisms encoded in the weights, followed by the later removal of memorizing components.

Data Structures forStatistical Computing i…Data Structures for Statistical Computing in PythonFixing Weight DecayRegularization in AdamFixing Weight Decay Regularization in AdamThe Lottery TicketHypothesis: Finding…The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural NetworksArray Programming withNumPyArray Programming with NumPyGrokking: GeneralizationBeyond Overfitting on…Grokking: Generalization Beyond Overfitting on Small Algorithmic DatasetsThe Slingshot Mechanism:An Empirical Study of…The Slingshot Mechanism: An Empirical Study of Adaptive Optimizers and the Grokking PhenomenonTowards UnderstandingGrokking: An Effective…Towards Understanding Grokking: An Effective Theory of Representation LearningHidden Progress in DeepLearning: SGD Learns…Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitChain of ThoughtPrompting Elicits…Chain of Thought Prompting Elicits Reasoning in Large Language ModelsEmergent Abilities ofLarge Language ModelsEmergent Abilities of Large Language ModelsThe Effects of RewardMisspecification…The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallThe Quantization Modelof Neural ScalingThe Quantization Model of Neural ScalingA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaSeeing Is Believing:Brain-Inspired Modular…Seeing Is Believing: Brain-Inspired Modular Training for Mechanistic InterpretabilityThe semantic landscapeparadigm for neural…The semantic landscape paradigm for neural networksGrokking as theTransition from Lazy to…Grokking as the Transition from Lazy to Rich Training DynamicsGrokking GroupMultiplication with…Grokking Group Multiplication with CosetsCompositionalCapabilities of…Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable TasksWhy Do You Grok? ATheoretical Analysis of…Why Do You Grok? A Theoretical Analysis of Grokking Modular AdditionPre-trained LargeLanguage Models Use…Pre-trained Large Language Models Use Fourier Features to Compute AdditionAlternating GradientFlows: A Theory of…Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksPhysics of LanguageModels: Part 1, Learnin…Physics of Language Models: Part 1, Learning Hierarchical Language StructuresProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretability過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。