Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Understanding the internal representations of large language models (LLMs) can help explain models' behavior and verify their alignment with human values. Given the capabilities of LLMs in generating human-understandable text, we propose leveraging the model itself to explain its internal representations in natural language. We introduce a framework called Patchscopes and show how it can be used to answer a wide range of questions about an LLM's computation. We show that many prior interpretability methods based on projecting representations into the vocabulary space and intervening on the LLM computation can be viewed as instances of this framework. Moreover, several of their shortcomings such as failure in inspecting early layers or lack of expressivity can be mitigated by Patchscopes. Beyond unifying prior inspection techniques, Patchscopes also opens up new possibilities such as using a more capable model to explain the representations of a smaller model, and multihop reasoning error correction.

Eliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisLocalizing ModelBehavior with Path…Localizing Model Behavior with Path PatchingDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsDoes Localization InformEditing? Surprising…Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsMeasuring andManipulating Knowledge…Measuring and Manipulating Knowledge Representations in Language ModelsLinearity of RelationDecoding in Transformer…Linearity of Relation Decoding in Transformer Language ModelsJump to Conclusions:Short-Cutting…Jump to Conclusions: Short-Cutting Transformers With Linear TransformationsLanguage ModelsImplement Simple…Language Models Implement Simple Word2Vec-style Vector ArithmeticTowards Best Practicesof Activation Patching…Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsRAVEL: EvaluatingInterpretability Method…RAVEL: Evaluating Interpretability Methods on Disentangling Language Model RepresentationsBackward Lens:Projecting Language…Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceWho's asking? Userpersonas and the…Who's asking? User personas and the mechanics of latent misalignmentMechanisticUnderstanding and…Mechanistic Understanding and Mitigation of Language Model Non-Factual HallucinationsTowards UnifyingInterpretability and…Towards Unifying Interpretability and Control: Evaluation via InterventionLoFiT: LocalizedFine-tuning on LLM…LoFiT: Localized Fine-tuning on LLM RepresentationsOpen Problems inMechanistic…Open Problems in Mechanistic InterpretabilityDo Multilingual LLMsThink In English?Do Multilingual LLMs Think In English?The Semantic HubHypothesis: Language…The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and ModalitiesActivation Oracles:Training and Evaluating…Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation ExplainersSeparating Tongue fromThought: Activation…Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in TransformersRepresentationEngineering for…Representation Engineering for Large-Language Models: Survey and Research ChallengesPatchscopes: A UnifyingFramework for Inspectin…Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。