Nash Learning from Human Feedback

Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Typically, RLHF involves the initial step of learning a reward model from human feedback, often expressed as preferences between pairs of text generations produced by a pre-trained LLM. Subsequently, the LLM's policy is fine-tuned by optimizing it to maximize the reward model through a reinforcement learning algorithm. However, an inherent limitation of current reward models is their inability to fully represent the richness of human preferences and their dependency on the sampling distribution. In this study, we introduce an alternative pipeline for the fine-tuning of LLMs using pairwise human feedback. Our approach entails the initial learning of a preference model, which is conditioned on two inputs given a prompt, followed by the pursuit of a policy that consistently generates responses preferred over those generated by any competing policy, thus defining the Nash equilibrium of this preference model. We term this approach Nash learning from human feedback (NLHF). In the context of a tabular policy representation, we present a novel algorithmic solution, Nash-MD, founded on the principles of mirror descent. This algorithm produces a sequence of policies, with the last iteration converging to the regularized Nash equilibrium. Additionally, we explore parametric representations of policies and introduce gradient descent algorithms for deep-learning architectures. To demonstrate the effectiveness of our approach, we present experimental results involving the fine-tuning of a LLM for a text summarization task. We believe NLHF offers a compelling avenue for preference learning and policy optimization with the potential of advancing the field of aligning LLMs with human preferences.

Direct NashOptimization: Teaching…Direct Nash Optimization: Teaching Language Models to Self-Improve with General PreferencesPreference Fine-Tuningof LLMs Should Leverage…Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy DataREBEL: ReinforcementLearning via Regressing…REBEL: Reinforcement Learning via Regressing Relative RewardsKTO: Model Alignment asProspect Theoretic…KTO: Model Alignment as Prospect Theoretic OptimizationReMax: A Simple,Effective, and Efficien…ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsGroup Robust PreferenceOptimization in…Group Robust Preference Optimization in Reward-free RLHFExploratory PreferenceOptimization: Harnessin…Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHFSelf-Play PreferenceOptimization for…Self-Play Preference Optimization for Language Model AlignmentStatisticalImpossibility and…Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash EquilibriumIterative Nash PolicyOptimization: Aligning…Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret LearningCorrecting the Mythos ofKL-Regularization…Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference OptimizationOn the Algorithmic Biasof Aligning Large…On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching RegularizationNash Learning from HumanFeedbackNash Learning from Human FeedbackFocus paperCiting papersOlderNewer

No earlier referenced papers in the catalog yet.

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.