Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias

Common methods for interpreting neural models in natural language processing typically examine either their structure or their behavior, but not both. We propose a methodology grounded in the theory of causal mediation analysis for interpreting which parts of a model are causally implicated in its behavior. It enables us to analyze the mechanisms by which information flows from input to output through various model components, known as mediators. We apply this methodology to analyze gender bias in pre-trained Transformer language models. We study the role of individual neurons and attention heads in mediating gender bias across three datasets designed to gauge a model's sensitivity to gender bias. Our mediation analysis reveals that gender bias effects are (i) sparse, concentrated in a small part of the network; (ii) synergistic, amplified or repressed by different components; and (iii) decomposable into effects flowing directly from the input and indirectly through the mediators.

Direct and IndirectEffectsDirect and Indirect EffectsA general approach tocausal mediation…A general approach to causal mediation analysis.Social Bias in ElicitedNatural Language…Social Bias in Elicited Natural Language InferencesLearning Gender-NeutralWord EmbeddingsLearning Gender-Neutral Word EmbeddingsGender Bias inCoreference Resolution…Gender Bias in Coreference Resolution: Evaluation and Debiasing MethodsGender Bias inCoreference ResolutionGender Bias in Coreference ResolutionLipstick on a Pig:Debiasing Methods Cover…Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove ThemIt's All in the Name:Mitigating Gender Bias…It's All in the Name: Mitigating Gender Bias with Name-Based Counterfactual Data SubstitutionAssessing Social andIntersectional Biases i…Assessing Social and Intersectional Biases in Contextualized Word RepresentationsHuggingFace'sTransformers…HuggingFace's Transformers: State-of-the-art Natural Language ProcessingInvestigating GenderBias in Language Models…Investigating Gender Bias in Language Models Using Causal Mediation AnalysisReducing Sentiment Biasin Language Models via…Reducing Sentiment Bias in Language Models via Counterfactual EvaluationLanguage (Technology) isPower: A Critical Surve…Language (Technology) is Power: A Critical Survey of "Bias" in NLPModular RepresentationUnderlies Systematic…Modular Representation Underlies Systematic Generalization in Neural Natural Language Inference ModelsBERTology Meets Biology:Interpreting Attention…BERTology Meets Biology: Interpreting Attention in Protein Language ModelsDouble-Hard Debias:Tailoring Word…Double-Hard Debias: Tailoring Word Embeddings for Gender Bias MitigationVisualizing Transformersfor NLP: A Brief SurveyVisualizing Transformers for NLP: A Brief SurveyA Survey on Gender Biasin Natural Language…A Survey on Gender Bias in Natural Language ProcessingCausal Analysis ofSyntactic Agreement…Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsWhen Bert Forgets How ToPOS: Amnesic Probing of…When Bert Forgets How To POS: Amnesic Probing of Linguistic Properties and MLM PredictionsAn InterpretabilityIllusion for BERTAn Interpretability Illusion for BERTTowards a ComprehensiveUnderstanding and…Towards a Comprehensive Understanding and Accurate Evaluation of Societal Biases in Pre-Trained TransformersTheories of "Gender" inNLP Bias ResearchTheories of "Gender" in NLP Bias ResearchHow to use and interpretactivation patchingHow to use and interpret activation patchingCausal MediationAnalysis for…Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender BiasEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.