Representation Engineering: A Top-Down Approach to AI Transparency

In this paper, we identify and characterize the emerging area of representation engineering (RepE), an approach to enhancing the transparency of AI systems that draws on insights from cognitive neuroscience. RepE places population-level representations, rather than neurons or circuits, at the center of analysis, equipping us with novel methods for monitoring and manipulating high-level cognitive phenomena in deep neural networks (DNNs). We provide baselines and an initial analysis of RepE techniques, showing that they offer simple yet effective solutions for improving our understanding and control of large language models. We showcase how these methods can provide traction on a wide range of safety-relevant problems, including honesty, harmlessness, power-seeking, and more, demonstrating the promise of top-down transparency research. We hope that this work catalyzes further exploration of RepE and fosters advancements in the transparency and safety of AI systems.

Deep InsideConvolutional Networks…Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency MapsStriving for Simplicity:The All Convolutional…Striving for Simplicity: The All Convolutional NetUnderstanding NeuralNetworks Through Deep…Understanding Neural Networks Through Deep VisualizationUnsupervisedRepresentation Learning…Unsupervised Representation Learning with Deep Convolutional Generative Adversarial NetworksSmoothGrad: removingnoise by adding noiseSmoothGrad: removing noise by adding noiseLearning to GenerateReviews and Discovering…Learning to Generate Reviews and Discovering SentimentWhat Does BERT Look At?An Analysis of BERT's…What Does BERT Look At? An Analysis of BERT's AttentionLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersUnderstanding the Roleof Individual Units in…Understanding the Role of Individual Units in a Deep Neural NetworkIn-context Learning andInduction HeadsIn-context Learning and Induction HeadsActivation Addition:Steering Language Model…Activation Addition: Steering Language Models Without OptimizationEliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionFundamental Limitationsof Alignment in Large…Fundamental Limitations of Alignment in Large Language ModelsIntroduction to AISafety, Ethics, and…Introduction to AI Safety, Ethics, and SocietyEmergent World Modelsand Latent Variable…Emergent World Models and Latent Variable Estimation in Chess-Playing Language ModelsBEEAR: Embedding-basedAdversarial Removal of…BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language ModelsPersona Features ControlEmergent MisalignmentPersona Features Control Emergent MisalignmentInterpretation MeetsSafety: A Survey on…Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM SafetyConvergent LinearRepresentations of…Convergent Linear Representations of Emergent MisalignmentLLMs Know More Than TheyShow: On the Intrinsic…LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsActivation SpaceInterventions Can Be…Activation Space Interventions Can Be Transferred Between Large Language ModelsDo I Know This Entity?Knowledge Awareness and…Do I Know This Entity? Knowledge Awareness and Hallucinations in Language ModelsThe Assistant Axis:Situating and…The Assistant Axis: Situating and Stabilizing the Default Persona of Language ModelsRepresentationEngineering: A Top-Down…Representation Engineering: A Top-Down Approach to AI Transparency過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。