Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection

The ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models. We present Iterative Null-space Projection (INLP), a novel method for removing information from neural representations. Our method is based on repeated training of linear classifiers that predict a certain property we aim to remove, followed by projection of the representations on their null-space. By doing so, the classifiers become oblivious to that target property, making it hard to linearly separate the data according to it. While applicable for multiple uses, we evaluate our method on bias and fairness use-cases, and show that our method is able to mitigate bias in word embeddings, as well as to increase fairness in a setting of multi-class classification.

Trends & Controversies:Support Vector MachinesTrends & Controversies: Support Vector MachinesVisualizing Data usingt-SNEVisualizing Data using t-SNEScikit-learn: MachineLearning in PythonScikit-learn: Machine Learning in PythonUnsupervised FeatureLearning and Deep…Unsupervised Feature Learning and Deep Learning: A Review and New PerspectivesSimLex-999: EvaluatingSemantic Models With…SimLex-999: Evaluating Semantic Models With (Genuine) Similarity EstimationEquality of Opportunityin Supervised LearningEquality of Opportunity in Supervised LearningSemantics derivedautomatically from…Semantics derived automatically from language corpora contain human-like biasesLanguage Models asKnowledge Bases?Language Models as Knowledge Bases?BERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingUnderstandingUndesirable Word…Understanding Undesirable Word Embedding AssociationsA Structural Probe forFinding Syntax in Word…A Structural Probe for Finding Syntax in Word RepresentationsBERT Rediscovers theClassical NLP PipelineBERT Rediscovers the Classical NLP PipelineOSCaR: OrthogonalSubspace Correction and…OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word EmbeddingsThe Low-DimensionalLinear Geometry of…The Low-Dimensional Linear Geometry of Contextualized Word RepresentationsCEBaB: Estimating theCausal Effects of…CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorTheories of "Gender" inNLP Bias ResearchTheories of "Gender" in NLP Bias ResearchGold Doesn't AlwaysGlitter: Spectral…Gold Doesn't Always Glitter: Spectral Removal of Linear and Nonlinear Guarded Attribute InformationLEACE: Perfect linearconcept erasure in…LEACE: Perfect linear concept erasure in closed formRefusal in LanguageModels Is Mediated by a…Refusal in Language Models Is Mediated by a Single DirectionCan SensitiveInformation Be Deleted…Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction AttacksImprovingInstruction-Following i…Improving Instruction-Following in Language Models through Activation SteeringFinding AlignmentsBetween Interpretable…Finding Alignments Between Interpretable Causal Variables and Distributed Neural RepresentationsA Primer on the InnerWorkings of…A Primer on the Inner Workings of Transformer-based Language ModelsRobust LLM safeguardingvia refusal feature…Robust LLM safeguarding via refusal feature adversarial trainingNull It Out: GuardingProtected Attributes by…Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。