Task-Specific Skill Localization in Fine-tuned Language Models

Pre-trained language models can be fine-tuned to solve diverse NLP tasks, including in few-shot settings. Thus fine-tuning allows the model to quickly pick up task-specific ``skills,'' but there has been limited study of where these newly-learnt skills reside inside the massive model. This paper introduces the term skill localization for this problem and proposes a solution. Given the downstream task and a model fine-tuned on that task, a simple optimization is used to identify a very small subset of parameters ($\sim0.01$% of model parameters) responsible for ($>95$%) of the model's performance, in the sense that grafting the fine-tuned values for just this tiny subset onto the pre-trained model gives performance almost as well as the fine-tuned model. While reminiscent of recent works on parameter-efficient fine-tuning, the novel aspects here are that: (i) No further re-training is needed on the subset (unlike, say, with lottery tickets). (ii) Notable improvements are seen over vanilla fine-tuning with respect to calibration of predictions in-distribution ($40$-$90$% error reduction) as well as the quality of predictions out-of-distribution (OOD). In models trained on multiple tasks, a stronger notion of skill localization is observed, where the sparse regions corresponding to different tasks are almost disjoint, and their overlap (when it happens) is a proxy for task similarity. Experiments suggest that localization via grafting can assist certain forms of continual learning.

GLUE: A Multi-TaskBenchmark and Analysis…GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language UnderstandingRoBERTa: A RobustlyOptimized BERT…RoBERTa: A Robustly Optimized BERT Pretraining ApproachMulti-Task Deep NeuralNetworks for Natural…Multi-Task Deep Neural Networks for Natural Language UnderstandingCompressing BERT:Studying the Effects of…Compressing BERT: Studying the Effects of Weight Pruning on Transfer LearningPrefix-Tuning:Optimizing Continuous…Prefix-Tuning: Optimizing Continuous Prompts for GenerationAdapterFusion:Non-Destructive Task…AdapterFusion: Non-Destructive Task Composition for Transfer LearningSuper Tickets inPre-Trained Language…Super Tickets in Pre-Trained Language Models: From Model Compression to Improving GeneralizationBitFit: SimpleParameter-efficient…BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-modelsDoes Localization InformEditing? Surprising…Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsDiscovering LatentKnowledge in Language…Discovering Latent Knowledge in Language Models Without SupervisionLarge Models areParsimonious Learners…Large Models are Parsimonious Learners: Activation Sparsity in Trained TransformersA Kernel-Based View ofLanguage Model…A Kernel-Based View of Language Model Fine-TuningThe Expressibility ofPolynomial based…The Expressibility of Polynomial based Attention SchemeConvergence of Two-LayerRegression with…Convergence of Two-Layer Regression with Nonlinear UnitsLocal Convergence ofApproximate Newton…Local Convergence of Approximate Newton Method for Two Layer Nonlinear RegressionUnmasking Transformers:A Theoretical Approach…Unmasking Transformers: A Theoretical Approach to Data Recovery via Attention WeightsContinual InstructionTuning for Large…Continual Instruction Tuning for Large Multimodal ModelsModel Tailor: MitigatingCatastrophic Forgetting…Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language ModelsAttention is NaturallySparse with Gaussian…Attention is Naturally Sparse with Gaussian Distributed InputFine-Tuning EnhancesExisting Mechanisms: A…Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingFunction Vectors inLarge Language ModelsFunction Vectors in Large Language ModelsSuperiority of Softmax:Unveiling the…Superiority of Softmax: Unveiling the Performance Edge Over Linear AttentionThe UnreasonableIneffectiveness of the…The Unreasonable Ineffectiveness of the Deeper LayersA Fast OptimizationView: Reformulating…A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication TimeTask-Specific SkillLocalization in…Task-Specific Skill Localization in Fine-tuned Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。