How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Pre-trained language models can be surprisingly adept at tasks they were not explicitly trained on, but how they implement these capabilities is poorly understood. In this paper, we investigate the basic mathematical abilities often acquired by pre-trained language models. Concretely, we use mechanistic interpretability techniques to explain the (limited) mathematical abilities of GPT-2 small. As a case study, we examine its ability to take in sentences such as "The war lasted from the year 1732 to the year 17", and predict valid two-digit end years (years > 32). We first identify a circuit, a small subset of GPT-2 small's computational graph that computes this task's output. Then, we explain the role of each circuit component, showing that GPT-2 small's final multi-layer perceptrons boost the probability of end years greater than the start year. Finally, we find related tasks that activate our circuit. Our results suggest that GPT-2 small computes greater-than using a complex but general mechanism that activates across diverse contexts.

Neural MachineTranslation of Rare…Neural Machine Translation of Rare Words with Subword UnitsDeep Contextualized WordRepresentationsDeep Contextualized Word RepresentationsUnder the Hood: UsingDiagnostic Classifiers…Under the Hood: Using Diagnostic Classifiers to Investigate and Improve how Language Models Track Agreement InformationBERT Rediscovers theClassical NLP PipelineBERT Rediscovers the Classical NLP PipelineWhat Does BERT Look At?An Analysis of BERT's…What Does BERT Look At? An Analysis of BERT's AttentionDo NLP Models KnowNumbers? Probing…Do NLP Models Know Numbers? Probing Numeracy in EmbeddingsThe emergence of numberand syntax units in LST…The emergence of number and syntax units in LSTM language modelsTransformer Feed-ForwardLayers Are Key-Value…Transformer Feed-Forward Layers Are Key-Value MemoriesWhen Bert Forgets How ToPOS: Amnesic Probing of…When Bert Forgets How To POS: Amnesic Probing of Linguistic Properties and MLM PredictionsNeuron-levelInterpretation of Deep…Neuron-level Interpretation of Deep NLP Models: A SurveyInterpretability in theWild: a Circuit for…Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallToward Transparent AI: ASurvey on Interpreting…Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural NetworksTowards AutomatedCircuit Discovery for…Towards Automated Circuit Discovery for Mechanistic InterpretabilityIdentifying and AdaptingTransformer-Components…Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language ModelAttribution PatchingOutperforms Automated…Attribution Patching Outperforms Automated Circuit DiscoveryTranscoders FindInterpretable LLM…Transcoders Find Interpretable LLM Feature CircuitsHow to use and interpretactivation patchingHow to use and interpret activation patchingTeaching Arithmetic toSmall TransformersTeaching Arithmetic to Small TransformersOpening the AI blackbox: program synthesis…Opening the AI black box: program synthesis via mechanistic interpretabilityLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimePre-trained LargeLanguage Models Use…Pre-trained Large Language Models Use Fourier Features to Compute AdditionGrokking GroupMultiplication with…Grokking Group Multiplication with CosetsWhat Do VLMs NOTICE? AMechanistic…What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and EvaluationAttention heads of largelanguage modelsAttention heads of large language modelsHow does GPT-2 computegreater-than?…How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。