Making Pre-trained Language Models Better Few-shot Learners

The recent GPT-3 model (Brown et al., 2020) achieves remarkable few-shot performance solely by leveraging a natural-language prompt and a few task demonstrations as input context. Inspired by their findings, we study few-shot learning in a more practical scenario, where we use smaller language models for which fine-tuning is computationally efficient. We present LM-BFF--better few-shot fine-tuning of language models--a suite of simple and complementary techniques for fine-tuning language models on a small number of annotated examples. Our approach includes (1) prompt-based fine-tuning together with a novel pipeline for automating prompt generation; and (2) a refined strategy for dynamically and selectively incorporating demonstrations into each context. Finally, we present a systematic evaluation for analyzing few-shot performance on a range of NLP tasks, including classification and regression. Our experiments demonstrate that our methods combine to dramatically outperform standard fine-tuning procedures in this low resource setting, achieving up to 30% absolute improvement, and 11% on average across all tasks. Our approach makes minimal assumptions on task resources and domain expertise, and hence constitutes a strong task-agnostic method for few-shot learning.

Qualitative SpatialReasoning over Question…Qualitative Spatial Reasoning over Questions (Short Paper)Sentence Encoders onSTILTs: Supplementary…Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data TasksSentence-BERT: SentenceEmbeddings using Siames…Sentence-BERT: Sentence Embeddings using Siamese BERT-NetworksTransformer-BasedFeature Learning for…Transformer-Based Feature Learning for Algorithm Selection in Combinatorial OptimisationBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingCommonsense KnowledgeMining from Pretrained…Commonsense Knowledge Mining from Pretrained ModelsAutomaticallyIdentifying Words That…Automatically Identifying Words That Can Serve as Labels for Few-Shot Text ClassificationFine-Tuning PretrainedLanguage Models: Weight…Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early StoppingLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersMixout: EffectiveRegularization to…Mixout: Effective Regularization to Finetune Large-scale Pretrained Language ModelsFactual Probing Is[MASK]: Learning vs…Factual Probing Is [MASK]: Learning vs. Learning to RecallRevisiting Few-sampleBERT Fine-tuningRevisiting Few-sample BERT Fine-tuningP-Tuning v2: PromptTuning Can Be Comparabl…P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and TasksAMMUS : A Survey ofTransformer-based…AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language ProcessingFantastically OrderedPrompts and Where to…Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityIDPG: AnInstance-Dependent…IDPG: An Instance-Dependent Prompt Generation MethodSTT: Soft TemplateTuning for Few-Shot…STT: Soft Template Tuning for Few-Shot AdaptationPrompt-learning forFine-grained Entity…Prompt-learning for Fine-grained Entity TypingVisual Prompting:Modifying Pixel Space t…Visual Prompting: Modifying Pixel Space to Adapt Pre-trained ModelsLogic Against Bias:Textual Entailment…Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence ReasoningLarge Language ModelsAre Human-Level Prompt…Large Language Models Are Human-Level Prompt EngineersReliable Gradient-freeand Likelihood-free…Reliable Gradient-free and Likelihood-free Prompt TuningVita-CLIP: Video andtext adaptive CLIP via…Vita-CLIP: Video and text adaptive CLIP via Multimodal PromptingMaking Text EmbeddersFew-Shot LearnersMaking Text Embedders Few-Shot LearnersMaking Pre-trainedLanguage Models Better…Making Pre-trained Language Models Better Few-shot Learners過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。