A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

While alignment algorithms are now commonly used to tune pre-trained language models towards a user's preferences, we lack explanations for the underlying mechanisms in which models become ``aligned'', thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm, direct preference optimization (DPO), and the mechanisms by which it reduces toxicity. Namely, we first study how toxicity is represented and elicited in a pre-trained language model, GPT2-medium. We then apply DPO with a carefully crafted pairwise dataset to reduce toxicity. We examine how the resulting model averts toxic outputs, and find that capabilities learned from pre-training are not removed, but rather bypassed. We use this insight to demonstrate a simple method to un-align the model, reverting it back to its toxic behavior.

Bridging Nonlinearitiesand Stochastic…Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear UnitsBERT Rediscovers theClassical NLP PipelineBERT Rediscovers the Classical NLP PipelineTransformer Feed-ForwardLayers Are Key-Value…Transformer Feed-Forward Layers Are Key-Value MemoriesShadow Alignment: TheEase of Subverting…Shadow Alignment: The Ease of Subverting Safely-Aligned Language ModelsInference-TimeIntervention: Eliciting…Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsDissecting Recall ofFactual Associations in…Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsMechanisticallyanalyzing the effects o…Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksLinearity of RelationDecoding in Transformer…Linearity of Relation Decoding in Transformer Language ModelsFunction Vectors inLarge Language ModelsFunction Vectors in Large Language ModelsFine-tuning AlignedLanguage Models…Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!DiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)Assessing theBrittleness of Safety…Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionEight Methods toEvaluate Robust…Eight Methods to Evaluate Robust Unlearning in LLMsNeuron-Level KnowledgeAttribution in Large…Neuron-Level Knowledge Attribution in Large Language ModelsPreference Tuning ForToxicity Mitigation…Preference Tuning For Toxicity Mitigation Generalizes Across LanguagesModel Editing as aRobust and Denoised…Model Editing as a Robust and Denoised variant of DPO: A Case Study on ToxicityAn AdversarialPerspective on Machine…An Adversarial Perspective on Machine Unlearning for AI SafetyTargeted LatentAdversarial Training…Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMsAblation is Not Enoughto Emulate DPO: How…Ablation is Not Enough to Emulate DPO: How Neuron Dynamics Drive Toxicity ReductionICLR: In-ContextLearning of…ICLR: In-Context Learning of RepresentationsWeak-to-StrongJailbreaking on Large…Weak-to-Strong Jailbreaking on Large Language ModelsUnderstanding JailbreakSuccess: A Study of…Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language ModelsA MechanisticUnderstanding of…A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。