Subliminal Learning: Language models transmit behavioral traits via hidden signals in data

We study subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model with some trait T (such as liking owls or being misaligned) generates a dataset consisting solely of number sequences. Remarkably, a "student" model trained on this dataset learns T. This occurs even when the data is filtered to remove references to T. We observe the same effect when training on code or reasoning traces generated by the same teacher model. However, we do not observe the effect when the teacher and student have different base models. To help explain our findings, we prove a theoretical result showing that subliminal learning occurs in all neural networks under certain conditions, and demonstrate subliminal learning in a simple MLP classifier. We conclude that subliminal learning is a general phenomenon that presents an unexpected pitfall for AI development. Distillation could propagate unintended traits, even when developers try to prevent this via data filtering.

Gradient-based learningapplied to document…Gradient-based learning applied to document recognitionDistilling the Knowledgein a Neural NetworkDistilling the Knowledge in a Neural NetworkTraining Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsSleeper Agents: TrainingDeceptive LLMs that…Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingSycophancy toSubterfuge…Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language ModelsAlignment faking inlarge language modelsAlignment faking in large language modelsThought Crime: Backdoorsand Emergent…Thought Crime: Backdoors and Emergent Misalignment in Reasoning ModelsPersona Features ControlEmergent MisalignmentPersona Features Control Emergent MisalignmentEmergent Misalignment:Narrow finetuning can…Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsModel Organisms forEmergent MisalignmentModel Organisms for Emergent MisalignmentMonitoring ReasoningModels for Misbehavior…Monitoring Reasoning Models for Misbehavior and the Risks of Promoting ObfuscationDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningWeird Generalization andInductive Backdoors: Ne…Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMsSchool of Reward Hacks:Hacking harmless tasks…School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMsNatural EmergentMisalignment from Rewar…Natural Emergent Misalignment from Reward Hacking in Production RLInoculation Prompting:Eliciting traits from…Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timeTowards UnderstandingSubliminal Learning…Towards Understanding Subliminal Learning: When and How Hidden Biases TransferNarrow Finetuning LeavesClearly Readable Traces…Narrow Finetuning Leaves Clearly Readable Traces in Activation DifferencesConditionalmisalignment: common…Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggersReinforcement LearningAmplifies Emergent…Reinforcement Learning Amplifies Emergent Misalignment from Harmless RewardsSubliminal Effects inYour Data: A General…Subliminal Effects in Your Data: A General Mechanism via Log-LinearitySubliminal Learning is aLoRA ArtifactSubliminal Learning is a LoRA ArtifactFinding RELIEF: ShapingReasoning Behavior…Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief EngineeringPhantom Transfer:Data-level Defences are…Phantom Transfer: Data-level Defences are Insufficient Against Data PoisoningSubliminal Learning:Language models transmi…Subliminal Learning: Language models transmit behavioral traits via hidden signals in data過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。