A Kernel-Based View of Language Model Fine-Tuning

It has become standard to solve NLP tasks by fine-tuning pre-trained language models (LMs), especially in low-data settings. There is minimal theoretical understanding of empirical success, e.g., why fine-tuning a model with $10^8$ or more parameters on a couple dozen training points does not result in overfitting. We investigate whether the Neural Tangent Kernel (NTK) - which originated as a model to study the gradient descent dynamics of infinitely wide networks with suitable random initialization - describes fine-tuning of pre-trained LMs. This study was inspired by the decent performance of NTK for computer vision tasks (Wei et al., 2022). We extend the NTK formalism to Adam and use Tensor Programs (Yang, 2020) to characterize conditions under which the NTK lens may describe fine-tuning updates to pre-trained language models. Extensive experiments on 14 NLP tasks validate our theory and show that formulating the downstream task as a masked word prediction problem through prompting often induces kernel-based dynamics during fine-tuning. Finally, we use this kernel view to propose an explanation for the success of parameter-efficient subspace-based fine-tuning methods.

A large annotated corpusfor learning natural…A large annotated corpus for learning natural language inferenceTensor Programs II:Neural Tangent Kernel…Tensor Programs II: Neural Tangent Kernel for Any ArchitectureTensor Programs III:Neural Matrix LawsTensor Programs III: Neural Matrix LawsPre-train, Prompt, andPredict: A Systematic…Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language ProcessingBitFit: SimpleParameter-efficient…BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-modelsTensor Programs V:Tuning Large Neural…Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferExploring the Limits ofLarge Scale Pre-trainingExploring the Limits of Large Scale Pre-trainingAn Analysis of Attentionvia the Lens of…An Analysis of Attention via the Lens of Exchangeability and Latent Variable ModelsFine-Tuning LanguageModels with Just Forwar…Fine-Tuning Language Models with Just Forward PassesTask-Specific SkillLocalization in…Task-Specific Skill Localization in Fine-tuned Language ModelsA review of deeplearning techniques for…A review of deep learning techniques for speech processingTRAK: Attributing ModelBehavior at ScaleTRAK: Attributing Model Behavior at ScaleTask Arithmetic in theTangent Space: Improved…Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsThe Journey, Not theDestination: How Data…The Journey, Not the Destination: How Data Guides Diffusion ModelsScaling Down to ScaleUp: A Guide to…Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-TuningSample basedExplanations via…Sample based Explanations via Generalized RepresentersSparsity May Cry: Let UsFail (Current) Sparse…Sparsity May Cry: Let Us Fail (Current) Sparse Neural Networks Together!What and How doesIn-Context Learning…What and How does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and GeneralizationSmall-to-LargeGeneralization: Data…Small-to-Large Generalization: Data Influences Models Consistently Across ScaleA Kernel-Based View ofLanguage Model…A Kernel-Based View of Language Model Fine-Tuning過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。