Eliciting Latent Predictions from Transformers with the Tuned Lens

We analyze transformers from the perspective of iterative inference, seeking to understand how model predictions are refined layer by layer. To do so, we train an affine probe for each block in a frozen pretrained model, making it possible to decode every hidden state into a distribution over the vocabulary. Our method, the tuned lens, is a refinement of the earlier "logit lens" technique, which yielded useful insights but is often brittle. We test our method on various autoregressive language models with up to 20B parameters, showing it to be more predictive, reliable and unbiased than the logit lens. With causal experiments, we show the tuned lens uses similar features to the model itself. We also find the trajectory of latent predictions can be used to detect malicious inputs with high accuracy. All code needed to reproduce our results can be found at https://github.com/AlignmentResearch/tuned-lens.

Understandingintermediate layers…Understanding intermediate layers using linear classifier probesEnriching Word Vectorswith Subword InformationEnriching Word Vectors with Subword InformationDeeBERT: Dynamic EarlyExiting for Acceleratin…DeeBERT: Dynamic Early Exiting for Accelerating BERT InferenceAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleToy Models ofSuperpositionToy Models of SuperpositionTransformer Feed-ForwardLayers Build Prediction…Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceTraining a Helpful andHarmless Assistant with…Training a Helpful and Harmless Assistant with Reinforcement Learning from Human FeedbackInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallAnalyzing Transformersin Embedding SpaceAnalyzing Transformers in Embedding SpaceEmergent WorldRepresentations…Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskHelping Cancer Patientsto Choose the Best…Helping Cancer Patients to Choose the Best Treatment: Towards Automated Data-Driven and Personalized Information Presentation of Cancer Treatment OptionsDistilBERT, a distilledversion of BERT…DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighterDoes Circuit AnalysisInterpretability Scale?…Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in ChinchillaRepresentationEngineering: A Top-Down…Representation Engineering: A Top-Down Approach to AI TransparencyA MechanisticInterpretation of…A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisCharacterizingMechanisms for Factual…Characterizing Mechanisms for Factual Recall in Language ModelsHow Alignment andJailbreak Work: Explain…How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden StatesJailbreakLens:Interpreting Jailbreak…JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and CircuitBackward Lens:Projecting Language…Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceObfuscated ActivationsBypass LLM Latent-Space…Obfuscated Activations Bypass LLM Latent-Space DefensesFrom Tokens to Words: Onthe Inner Lexicon of…From Tokens to Words: On the Inner Lexicon of LLMsLocate-then-edit forMulti-hop Factual Recal…Locate-then-edit for Multi-hop Factual Recall under Knowledge EditingInterpretation MeetsSafety: A Survey on…Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM SafetyLarge Language Modelsand Causal Inference in…Large Language Models and Causal Inference in Collaboration: A Comprehensive SurveyEliciting LatentPredictions from…Eliciting Latent Predictions from Transformers with the Tuned LensEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.