Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

The last decade of machine learning has seen drastic increases in scale and capabilities. Deep neural networks (DNNs) are increasingly being deployed in the real world. However, they are difficult to analyze, raising concerns about using them without a rigorous understanding of how they function. Effective tools for interpreting them will be important for building more trustworthy AI by helping to identify problems, fix bugs, and improve basic understanding. In particular, "inner" interpretability techniques, which focus on explaining the internal components of DNNs, are well-suited for developing a mechanistic understanding, guiding manual modifications, and reverse engineering solutions. Much recent work has focused on DNN interpretability, and rapid progress has thus far made a thorough systematization of methods difficult. In this survey, we review over 300 works with a focus on inner interpretability tools. We introduce a taxonomy that classifies methods by what part of the network they help to explain (weights, neurons, subnetworks, or latent representations) and whether they are implemented during (intrinsic) or after (post hoc) training. To our knowledge, we are also the first to survey a number of connections between interpretability research and work in adversarial robustness, continual learning, modularity, network compression, and studying the human visual system. We discuss key challenges and argue that the status quo in interpretability research is largely unproductive. Finally, we highlight the importance of future work that emphasizes diagnostics, debugging, adversaries, and benchmarking in order to make interpretability tools more useful to engineers in practical applications.

Dropout: a simple way toprevent neural networks…Dropout: a simple way to prevent neural networks from overfittingImageNet Large ScaleVisual Recognition…ImageNet Large Scale Visual Recognition ChallengeOn the importance ofsingle directions for…On the importance of single directions for generalizationRevisiting theImportance of Individua…Revisiting the Importance of Individual Units in CNNs via AblationLifelong Learning withDynamically Expandable…Lifelong Learning with Dynamically Expandable NetworksIdentifying andControlling Important…Identifying and Controlling Important Neurons in Neural Machine TranslationUnderstanding the Roleof Individual Units in…Understanding the Role of Individual Units in a Deep Neural NetworkZoom In: An Introductionto CircuitsZoom In: An Introduction to CircuitsConcept BottleneckModelsConcept Bottleneck ModelsDetecting Modularity inDeep Neural NetworksDetecting Modularity in Deep Neural NetworksExemplary Natural ImagesExplain CNN Activations…Exemplary Natural Images Explain CNN Activations Better than State-of-the-Art Feature VisualizationLearning Multiple Layersof Features from Tiny…Learning Multiple Layers of Features from Tiny ImagesOne Thing to Fool themAll: Generating…One Thing to Fool them All: Generating Interpretable, Universal, and Physically-Realizable Adversarial FeaturesDiagnostics for DeepNeural Networks with…Diagnostics for Deep Neural Networks with Automated Copy/Paste AttacksHow does GPT-2 computegreater-than?…How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallRed Teaming Deep NeuralNetworks with Feature…Red Teaming Deep Neural Networks with Feature Synthesis ToolsTowards a MechanisticInterpretation of…Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsExplaining AI throughmechanistic…Explaining AI through mechanistic interpretabilityA Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsExplainable ArtificialIntelligence (XAI) 2.0…Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research DirectionsFoundational Challengesin Assuring Alignment…Foundational Challenges in Assuring Alignment and Safety of Large Language ModelsExplainable ArtificialIntelligence: A Survey…Explainable Artificial Intelligence: A Survey of Needs, Techniques, Applications, and Future DirectionOpen Problems inMechanistic…Open Problems in Mechanistic InterpretabilityToward Transparent AI: ASurvey on Interpreting…Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。