Understanding intermediate layers using linear classifier probes

Neural network models have a reputation for being black boxes. We propose to monitor the features at every layer of a model and measure how suitable they are for classification. We use linear classifiers, which we refer to as "probes", trained entirely independently of the model itself. This helps us better understand the roles and dynamics of the intermediate layers. We demonstrate how this can be used to develop a better intuition about models and to diagnose potential problems. We apply this technique to the popular models Inception v3 and Resnet-50. Among other things, we observe experimentally that the linear separability of features increase monotonically along the depth of the model.

Intriguing properties ofneural networksIntriguing properties of neural networksSVCCA: Singular VectorCanonical Correlation…SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and InterpretabilityUnderstanding deeplearning requires…Understanding deep learning requires rethinking generalizationFractalNet: Ultra-DeepNeural Networks without…FractalNet: Ultra-Deep Neural Networks without ResidualsExplaining RecurrentNeural Network…Explaining Recurrent Neural Network Predictions in Sentiment AnalysisResidual ConnectionsEncourage Iterative…Residual Connections Encourage Iterative InferenceUnderstanding SyntheticGradients and Decoupled…Understanding Synthetic Gradients and Decoupled Neural InterfacesSanity Checks forSaliency MapsSanity Checks for Saliency MapsOn the importance ofsingle directions for…On the importance of single directions for generalizationShallowing DeepNetworks: Layer-Wise…Shallowing Deep Networks: Layer-Wise Pruning Based on Feature RepresentationsEvaluating RecurrentNeural Network…Evaluating Recurrent Neural Network ExplanationsInformation-TheoreticProbing for Linguistic…Information-Theoretic Probing for Linguistic StructureDiscriminativeRegression Machine: A…Discriminative Regression Machine: A Classifier for High-Dimensional Data or Imbalanced DataNeural Networks asKernel Learners: The…Neural Networks as Kernel Learners: The Silent Alignment EffectToward Transparent AI: ASurvey on Interpreting…Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural NetworksLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeEmergent World Modelsand Latent Variable…Emergent World Models and Latent Variable Estimation in Chess-Playing Language ModelsGeneral agents needworld modelsGeneral agents need world modelsUnderstandingintermediate layers…Understanding intermediate layers using linear classifier probes過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。