Refusal in Language Models Is Mediated by a Single Direction

Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.

The Geometry of Truth:Emergent Linear…The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False DatasetsActivation Addition:Steering Language Model…Activation Addition: Steering Language Models Without OptimizationShadow Alignment: TheEase of Subverting…Shadow Alignment: The Ease of Subverting Safely-Aligned Language ModelsSteering Llama 2 viaContrastive Activation…Steering Llama 2 via Contrastive Activation AdditionThe LinearRepresentation…The Linear Representation Hypothesis and the Geometry of Large Language ModelsMechanisticallyanalyzing the effects o…Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksA MechanisticUnderstanding of…A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityA StrongREJECT for EmptyJailbreaksA StrongREJECT for Empty JailbreaksAssessing theBrittleness of Safety…Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsJailbreakBench: An OpenRobustness Benchmark fo…JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language ModelsRemoving RLHFProtections in GPT-4 vi…Removing RLHF Protections in GPT-4 via Fine-TuningJailbreaking LeadingSafety-Aligned LLMs wit…Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksSteering Without SideEffects: Improving…Steering Without Side Effects: Improving Post-Deployment Control of Language ModelsBEEAR: Embedding-basedAdversarial Removal of…BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language ModelsObfuscated ActivationsBypass LLM Latent-Space…Obfuscated Activations Bypass LLM Latent-Space DefensesAutomatic Pseudo-HarmfulPrompt Generation for…Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language ModelsSurgical, Cheap, andFlexible: Mitigating…Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector AblationTargeted LatentAdversarial Training…Targeted Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMsThe Geometry of Refusalin Large Language…The Geometry of Refusal in Large Language Models: Concept Cones and Representational IndependenceTowards UnderstandingSafety Alignment: A…Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsREINFORCE AdversarialAttacks on Large…REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic ObjectiveAblation is Not Enoughto Emulate DPO: How…Ablation is Not Enough to Emulate DPO: How Neuron Dynamics Drive Toxicity ReductionJust Enough Shifts:Mitigating Over-Refusal…Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-TuningThe PersistentVulnerability of Aligne…The Persistent Vulnerability of Aligned AI SystemsRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single Direction過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。