Understanding the Effects of RLHF on LLM Generalisation and Diversity

Large language models (LLMs) fine-tuned with reinforcement learning from human feedback (RLHF) have been used in some of the most widely deployed AI models to date, such as OpenAI's ChatGPT or Anthropic's Claude. While there has been significant work developing these methods, our understanding of the benefits and downsides of each stage in RLHF is still limited. To fill this gap, we present an extensive analysis of how each stage of the process (i.e. supervised fine-tuning (SFT), reward modelling, and RLHF) affects two key properties: out-of-distribution (OOD) generalisation and output diversity. OOD generalisation is crucial given the wide range of real-world scenarios in which these models are being used, while output diversity refers to the model's ability to generate varied outputs and is important for a variety of use cases. We perform our analysis across two base models on both summarisation and instruction following tasks, the latter being highly relevant for current LLM use cases. We find that RLHF generalises better than SFT to new inputs, particularly as the distribution shift between train and test becomes larger. However, RLHF significantly reduces output diversity compared to SFT across a variety of measures, implying a tradeoff in current LLM fine-tuning methods between generalisation and diversity. Our results provide guidance on which fine-tuning method should be used depending on the application, and show that more research is needed to improve the tradeoff between generalisation and diversity.

Bleu: a Method forAutomatic Evaluation of…Bleu: a Method for Automatic Evaluation of Machine TranslationTraining language modelsto follow instructions…Training language models to follow instructions with human feedbackScalingInstruction-Finetuned…Scaling Instruction-Finetuned Language ModelsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsAlpacaFarm: A SimulationFramework for Methods…AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackRRHF: Rank Responses toAlign Language Models…RRHF: Rank Responses to Align Language Models with Human Feedback without tearsDirect PreferenceOptimization: Your…Direct Preference Optimization: Your Language Model is Secretly a Reward ModelInstruction Tuning withGPT-4Instruction Tuning with GPT-4OpenAssistantConversations -…OpenAssistant Conversations - Democratizing Large Language Model AlignmentLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsGPT-4 Technical ReportGPT-4 Technical ReportChain of HindsightAligns Language Models…Chain of Hindsight Aligns Language Models with FeedbackReMax: A Simple,Effective, and Efficien…ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsTowards UnderstandingSycophancy in Language…Towards Understanding Sycophancy in Language ModelsFrom Language Modelingto Instruction…From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction TuningEvaluating the Diversityand Quality of LLM…Evaluating the Diversity and Quality of LLM Generated ContentEntropic DistributionMatching in Supervised…Entropic Distribution Matching in Supervised Fine-tuning of LLMs: Less Overfitting and Better DiversityKL-RegularizedReinforcement Learning…KL-Regularized Reinforcement Learning is Designed to Mode CollapseThe Price of Format:Diversity Collapse in…The Price of Format: Diversity Collapse in LLMsExploring Data ScalingTrends and Effects in…Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human FeedbackIPO: Your Language Modelis Secretly a Preferenc…IPO: Your Language Model is Secretly a Preference ClassifierDIVE: DiversifiedIterative…DIVE: Diversified Iterative Self-ImprovementOutcome-basedExploration for LLM…Outcome-based Exploration for LLM ReasoningA Survey on PersonalizedAlignment - The Missing…A Survey on Personalized Alignment - The Missing Piece for Large Language Models in Real-World ApplicationsUnderstanding theEffects of RLHF on LLM…Understanding the Effects of RLHF on LLM Generalisation and DiversityEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.