Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

The interpretation of deep learning models is a challenge due to their size, complexity, and often opaque internal state. In addition, many systems, such as image classifiers, operate on low-level features rather than high-level concepts. To address these challenges, we introduce Concept Activation Vectors (CAVs), which provide an interpretation of a neural net's internal state in terms of human-friendly concepts. The key idea is to view the high-dimensional internal state of a neural net as an aid, not an obstacle. We show how to use CAVs as part of a technique, Testing with CAVs (TCAV), that uses directional derivatives to quantify the degree to which a user-defined concept is important to a classification result--for example, how sensitive a prediction of "zebra" is to the presence of stripes. Using the domain of image classification as a testing ground, we describe how CAVs may be used to explore hypotheses and generate insights for a standard image classification network as well as a medical application.

Going Deeper withConvolutionsGoing Deeper with ConvolutionsRethinking the InceptionArchitecture for…Rethinking the Inception Architecture for Computer Vision"Why Should I TrustYou?": Explaining the…"Why Should I Trust You?": Explaining the Predictions of Any ClassifierNetwork Dissection:Quantifying…Network Dissection: Quantifying Interpretability of Deep Visual RepresentationsAxiomatic Attributionfor Deep NetworksAxiomatic Attribution for Deep NetworksUnderstanding Black-boxPredictions via…Understanding Black-box Predictions via Influence FunctionsSmoothGrad: removingnoise by adding noiseSmoothGrad: removing noise by adding noiseReal Time Image Saliencyfor Black Box…Real Time Image Saliency for Black Box ClassifiersA Roadmap for a RigorousScience of…A Roadmap for a Rigorous Science of InterpretabilityUnderstandingintermediate layers…Understanding intermediate layers using linear classifier probesThe (Un)reliability ofSaliency MethodsThe (Un)reliability of Saliency MethodsInterpretation of NeuralNetworks Is FragileInterpretation of Neural Networks Is FragileExplainable AI forMedical Imaging…Explainable AI for Medical Imaging: Knowledge MattersRegression ConceptVectors for…Regression Concept Vectors for Bidirectional Explanations in HistopathologyGAN Dissection:Visualizing and…GAN Dissection: Visualizing and Understanding Generative Adversarial NetworksDesigning Theory-DrivenUser-Centric Explainabl…Designing Theory-Driven User-Centric Explainable AIFrom local explanationsto global understanding…From local explanations to global understanding with explainable AI for treesExplainable AI: A Reviewof Machine Learning…Explainable AI: A Review of Machine Learning Interpretability MethodsCause and Effect:Concept-based…Cause and Effect: Concept-based Explanation of Neural NetworksExplaining the black-boxsmoothly - A…Explaining the black-box smoothly - A counterfactual approachAre ExplanationsHelpful? A Comparative…Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-MakingExplainable artificialintelligence: an…Explainable artificial intelligence: an analytical reviewAcquisition of chessknowledge in AlphaZeroAcquisition of chess knowledge in AlphaZeroConsiderations whenlearning additive…Considerations when learning additive explanations for black-box modelsInterpretability BeyondFeature Attribution…Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。