Mechanistic Interpretability for AI Safety - A Review

Understanding AI systems' inner workings is critical for ensuring value alignment and safety. This review explores mechanistic interpretability: reverse engineering the computational mechanisms and representations learned by neural networks into human-understandable algorithms and concepts to provide a granular, causal understanding. We establish foundational concepts such as features encoding knowledge within neural activations and hypotheses about their representation and computation. We survey methodologies for causally dissecting model behaviors and assess the relevance of mechanistic interpretability to AI safety. We examine benefits in understanding, control, alignment, and risks such as capability gains and dual-use concerns. We investigate challenges surrounding scalability, automation, and comprehensive interpretation. We advocate for clarifying concepts, setting standards, and scaling techniques to handle complex models and behaviors and expand to domains such as vision and reinforcement learning. Mechanistic interpretability could help prevent catastrophic outcomes as AI systems become more powerful and inscrutable.

Sparse AutoencodersReveal Universal Featur…Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language ModelsInterpretabilityresearch of deep…Interpretability research of deep learning: A literature surveyLarge Language ModelSafety: A Holistic…Large Language Model Safety: A Holistic SurveyAutomaticallyInterpreting Millions o…Automatically Interpreting Millions of Features in Large Language ModelsMulti-Step Reasoningwith Large Language…Multi-Step Reasoning with Large Language Models, a SurveyAgentic Large LanguageModels, a SurveyAgentic Large Language Models, a SurveyLLM Social SimulationsAre a Promising Researc…LLM Social Simulations Are a Promising Research MethodThe Geometry ofConcepts: Sparse…The Geometry of Concepts: Sparse Autoencoder Feature StructureInterpretation MeetsSafety: A Survey on…Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM SafetyUnderstanding the RepeatCurse in Large Language…Understanding the Repeat Curse in Large Language Models from a Feature PerspectiveWhen Truth IsOverridden: Uncovering…When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsLocate, Steer, andImprove: A Practical…Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language ModelsMechanisticInterpretability for AI…Mechanistic Interpretability for AI Safety - A Review中心の論文この論文を引用する論文古い新しい

カタログにはまだ過去の参考文献がありません。

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。