Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment training improves performance on almost all NLP evaluations, and is fully compatible with training for specialized skills such as python coding and summarization. We explore an iterated online mode of training, where preference models and RL policies are updated on a weekly cadence with fresh human feedback data, efficiently improving our datasets and models. Finally, we investigate the robustness of RLHF training, and identify a roughly linear relation between the RL reward and the square root of the KL divergence between the policy and its initialization. Alongside our main results, we perform peripheral analyses on calibration, competing objectives, and the use of OOD detection, compare our models with human writers, and provide samples from our models using prompts appearing in recent related work.

Recipes for Safety inOpen-domain ChatbotsRecipes for Safety in Open-domain ChatbotsEvaluating LargeLanguage Models Trained…Evaluating Large Language Models Trained on CodeA Simple Fix toMahalanobis Distance fo…A Simple Fix to Mahalanobis Distance for Improving Near-OOD DetectionTraining language modelsto follow instructions…Training language models to follow instructions with human feedbackLarge Language ModelAlignment: A SurveyLarge Language Model Alignment: A SurveyRLAIF vs. RLHF: ScalingReinforcement Learning…RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackSafe RLHF: SafeReinforcement Learning…Safe RLHF: Safe Reinforcement Learning from Human FeedbackDirect Language ModelAlignment from Online A…Direct Language Model Alignment from Online AI FeedbackTraining DiffusionModels with…Training Diffusion Models with Reinforcement LearningRule Based Rewards forLanguage Model SafetyRule Based Rewards for Language Model SafetyReward Model EnsemblesHelp Mitigate…Reward Model Ensembles Help Mitigate OveroptimizationRegularizing HiddenStates Enables Learning…Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsCollectiveConstitutional AI…Collective Constitutional AI: Aligning a Language Model with Public InputUnpacking DPO and PPO:Disentangling Best…Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference FeedbackFundamental Limitationsof Alignment in Large…Fundamental Limitations of Alignment in Large Language ModelsSafeChain: Safety ofLanguage Models with…SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning CapabilitiesTraining a Helpful andHarmless Assistant with…Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。