Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking

Fine-tuning on generalized tasks such as instruction following, code generation, and mathematics has been shown to enhance language models' performance on a range of tasks. Nevertheless, explanations of how such fine-tuning influences the internal computations in these models remain elusive. We study how fine-tuning affects the internal mechanisms implemented in language models. As a case study, we explore the property of entity tracking, a crucial facet of language comprehension, where models fine-tuned on mathematics have substantial performance gains. We identify the mechanism that enables entity tracking and show that (i) in both the original model and its fine-tuned versions primarily the same circuit implements entity tracking. In fact, the entity tracking circuit of the original model on the fine-tuned versions performs better than the full original model. (ii) The circuits of all the models implement roughly the same functionality: Entity tracking is performed by tracking the position of the correct entity in both the original model and its fine-tuned versions. (iii) Performance boost in the fine-tuned models is primarily attributed to its improved ability to handle the augmented positional information. To uncover these findings, we employ: Patch Patching, DCM, which automatically detects model components responsible for specific semantics, and CMAP, a new approach for patching activations across models to reveal improved mechanisms. Our findings suggest that fine-tuning enhances, rather than fundamentally alters, the mechanistic operation of the model.

BERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingRevealing the DarkSecrets of BERTRevealing the Dark Secrets of BERTToy Models ofSuperpositionToy Models of SuperpositionScalingInstruction-Finetuned…Scaling Instruction-Finetuned Language ModelsFine-Tuning can DistortPretrained Features and…Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionEntity Tracking inLanguage ModelsEntity Tracking in Language ModelsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsJudging LLM-as-a-Judgewith MT-Bench and…Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsEmergent WorldRepresentations…Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskProgress measures forgrokking via mechanisti…Progress measures for grokking via mechanistic interpretabilityA Toy Model ofUniversality: Reverse…A Toy Model of Universality: Reverse Engineering How Networks Learn Group OperationsMechanisticallyanalyzing the effects o…Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksHow do Language ModelsBind Entities in…How do Language Models Bind Entities in Context?Refusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionWhat Makes and BreaksSafety Fine-tuning? A…What Makes and Breaks Safety Fine-tuning? A Mechanistic StudyRepresentationalAnalysis of Binding in…Representational Analysis of Binding in Large Language ModelsSparse Feature Circuits:Discovering and Editing…Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsMonitoring Latent WorldStates in Language…Monitoring Latent World States in Language Models with Propositional ProbesNeuroplasticity andCorruption in Model…Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object IdentificationTargeted LatentAdversarial Training…Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMsLanguage Models useLookbacks to Track…Language Models use Lookbacks to Track BeliefsActivation SpaceInterventions Can Be…Activation Space Interventions Can Be Transferred Between Large Language ModelsLocate, Steer, andImprove: A Practical…Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language ModelsFine-Tuning EnhancesExisting Mechanisms: A…Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。