Training language models to follow instructions with human feedback

Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.

Fine-Tuning LanguageModels from Human…Fine-Tuning Language Models from Human PreferencesLearning to summarizefrom human feedbackLearning to summarize from human feedbackLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersRecursively SummarizingBooks with Human…Recursively Summarizing Books with Human FeedbackScaling Language Models:Methods, Analysis &…Scaling Language Models: Methods, Analysis & Insights from Training GopherA General LanguageAssistant as a…A General Language Assistant as a Laboratory for AlignmentProcess for AdaptingLanguage Models to…Process for Adapting Language Models to Society (PALMS) with Values-Targeted DatasetsAlignment of LanguageAgentsAlignment of Language AgentsFinetuned LanguageModels Are Zero-Shot…Finetuned Language Models Are Zero-Shot LearnersTruthfulQA: MeasuringHow Models Mimic Human…TruthfulQA: Measuring How Models Mimic Human FalsehoodsLaMDA: Language Modelsfor Dialog ApplicationsLaMDA: Language Models for Dialog ApplicationsMultitask PromptedTraining Enables…Multitask Prompted Training Enables Zero-Shot Task GeneralizationPlanBench: An ExtensibleBenchmark for Evaluatin…PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about ChangeFoundation models forgeneralist medical…Foundation models for generalist medical artificial intelligenceA Chinese Prompt AttackDataset for LLMs with…A Chinese Prompt Attack Dataset for LLMs with Evil ContentImproving FactualConsistency for…Improving Factual Consistency for Knowledge-Grounded Dialogue Systems via Knowledge Enhancement and AlignmentIs ChatGPT the UltimateProgramming Assistant -…Is ChatGPT the Ultimate Programming Assistant - How far is it?Exploring the Responsesof Large Language Model…Exploring the Responses of Large Language Models to Beginner Programmers' Help RequestsCollectiveConstitutional AI…Collective Constitutional AI: Aligning a Language Model with Public InputAlpaPICO: Extraction ofPICO Frames from…AlpaPICO: Extraction of PICO Frames from Clinical Trial Documents Using LLMsA matter of principle?AI alignment as the fai…A matter of principle? AI alignment as the fair treatment of claimsHow Good Is ChatGPT inGiving Advice on Your…How Good Is ChatGPT in Giving Advice on Your Visualization Design?Artificial intelligencefor modelling infectiou…Artificial intelligence for modelling infectious disease epidemicsVLA-RL: TowardsMasterful and General…VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningTraining language modelsto follow instructions…Training language models to follow instructions with human feedback過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。