著者: Neel Nanda , Andrew Lee , Martin Wattenberg - Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP 2023 被引用: 172
How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023a). In this work, we provide evidence of a closely related linear representation of the board. In particular, we show that probing for "my colour" vs. "opponent's colour" may be a simple yet powerful way to interpret the model's internal state. This precise understanding of the internal representations allows us to control the model's behaviour with simple vector arithmetic. Linear representations enable significant interpretability progress, which we demonstrate with further exploration of how the world model is computed.
✨ ログイン状態を確認しています… PDF 被引用 BibTeX を表示 BibTeX を閉じる BibTeX を表示 引用
Efficient Estimation of Word Representations in… Efficient Estimation of Word Representations in Vector Space Toy Models of Superposition Toy Models of Superposition Locating and Editing Factual Associations in… Locating and Editing Factual Associations in GPT Emergent World Representations… Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task Finding Neurons in a Haystack: Case Studies… Finding Neurons in a Haystack: Case Studies with Sparse Probing The Geometry of Truth: Emergent Linear… The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets Inference-Time Intervention: Eliciting… Inference-Time Intervention: Eliciting Truthful Answers from a Language Model Does Localization Inform Editing? Surprising… Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models The Hydra Effect: Emergent Self-repair in… The Hydra Effect: Emergent Self-repair in Language Model Computations Discovering Latent Knowledge in Language… Discovering Latent Knowledge in Language Models Without Supervision Mass-Editing Memory in a Transformer Mass-Editing Memory in a Transformer Progress measures for grokking via mechanisti… Progress measures for grokking via mechanistic interpretability Language Models Represent Space and Time Language Models Represent Space and Time On the Origins of Linear Representations in Larg… On the Origins of Linear Representations in Large Language Models Towards Best Practices of Activation Patching… Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Refusal in Language Models Is Mediated by a… Refusal in Language Models Is Mediated by a Single Direction Learning Interpretable Concepts: Unifying… Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models Opening the Black Box of Large Language Models… Opening the Black Box of Large Language Models: Two Views on Holistic Interpretability Is This the Subspace You Are Looking for? An… Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching Not All Language Model Features Are… Not All Language Model Features Are One-Dimensionally Linear ICLR: In-Context Learning of… ICLR: In-Context Learning of Representations AxBench: Steering LLMs? Even Simple Baselines… AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders From Flat to Hierarchical: Extractin… From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit Towards Principled Evaluations of Sparse… Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control Emergent Linear Representations in Worl… Emergent Linear Representations in World Models of Self-Supervised Sequence Models 過去の参考文献 中心の論文 この論文を引用する論文 古い 新しい ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。