Axiomatic Attribution for Deep Networks

We study the problem of attributing the prediction of a deep network to its input features, a problem previously studied by several other works. We identify two fundamental axioms---Sensitivity and Implementation Invariance that attribution methods ought to satisfy. We show that they are not satisfied by most known attribution methods, which we consider to be a fundamental weakness of those methods. We use the axioms to guide the design of a new attribution method called Integrated Gradients. Our method requires no modification to the original network and is extremely simple to implement; it just needs a few calls to the standard gradient operator. We apply this method to a couple of image models, a couple of text models and a chemistry model, demonstrating its ability to debug networks, to extract rules from a network, and to enable users to engage with models better.

Building a LargeAnnotated Corpus of…Building a Large Annotated Corpus of English: The Penn TreebankHow to ExplainIndividual…How to Explain Individual Classification DecisionsVisualizing andUnderstanding…Visualizing and Understanding Convolutional NetworksStriving for Simplicity:The All Convolutional…Striving for Simplicity: The All Convolutional NetUnderstanding NeuralNetworks Through Deep…Understanding Neural Networks Through Deep VisualizationNeural MachineTranslation by Jointly…Neural Machine Translation by Jointly Learning to Align and TranslateImageNet Large ScaleVisual Recognition…ImageNet Large Scale Visual Recognition ChallengeUnderstanding Deep ImageRepresentations by…Understanding Deep Image Representations by Inverting ThemNot Just a Black Box:Learning Important…Not Just a Black Box: Learning Important Features Through Propagating Activation Differences"Why Should I TrustYou?": Explaining the…"Why Should I Trust You?": Explaining the Predictions of Any ClassifierDevelopment andValidation of a Deep…Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus PhotographsLearning ImportantFeatures Through…Learning Important Features Through Propagating Activation DifferencesWhy should you trust myinterpretation?…Why should you trust my interpretation? Understanding uncertainty in LIME predictionsAnalyzing FederatedLearning through an…Analyzing Federated Learning through an Adversarial LensLearning data-drivendiscretizations for…Learning data-driven discretizations for partial differential equationsRelative AttributingPropagation…Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural NetworksReinforcement LearningInterpretation Methods…Reinforcement Learning Interpretation Methods: A SurveyInterpretingSuper-Resolution…Interpreting Super-Resolution Networks With Local Attribution MapsExplainable MachineLearning with Prior…Explainable Machine Learning with Prior Knowledge: An OverviewOne Explanation is NotEnough: Structured…One Explanation is Not Enough: Structured Attention Graphs for Image ClassificationExplaining a Series ofModels by Propagating…Explaining a Series of Models by Propagating Local Feature AttributionsPost-hocInterpretability for…Post-hoc Interpretability for Neural NLP: A SurveyToward Transparent AI: ASurvey on Interpreting…Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural NetworksIntegrated DecisionGradients: Compute Your…Integrated Decision Gradients: Compute Your Attributions Where the Model Makes Its DecisionAxiomatic Attributionfor Deep NetworksAxiomatic Attribution for Deep Networks過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。