Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

We introduce Inference-Time Intervention (ITI), a technique designed to enhance the "truthfulness" of large language models (LLMs). ITI operates by shifting model activations during inference, following a set of directions across a limited number of attention heads. This intervention significantly improves the performance of LLaMA models on the TruthfulQA benchmark. On an instruction-finetuned LLaMA called Alpaca, ITI improves its truthfulness from 32.5% to 65.1%. We identify a tradeoff between truthfulness and helpfulness and demonstrate how to balance it by tuning the intervention strength. ITI is minimally invasive and computationally inexpensive. Moreover, the technique is data efficient: while approaches like RLHF require extensive annotations, ITI locates truthful directions using only few hundred examples. Our findings suggest that LLMs may have an internal representation of the likelihood of something being true, even as they produce falsehoods on the surface.

Scaling Language Models:Methods, Analysis &…Scaling Language Models: Methods, Analysis & Insights from Training GopherA General LanguageAssistant as a…A General Language Assistant as a Laboratory for AlignmentWebGPT: Browser-assistedquestion-answering with…WebGPT: Browser-assisted question-answering with human feedbackExtracting LatentSteering Vectors from…Extracting Latent Steering Vectors from Pretrained Language ModelsTeaching language modelsto support answers with…Teaching language models to support answers with verified quotesLoRA: Low-RankAdaptation of Large…LoRA: Low-Rank Adaptation of Large Language ModelsSelf-critiquing modelsfor assisting human…Self-critiquing models for assisting human evaluatorsRed Teaming LanguageModels to Reduce Harms…Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons LearnedChain of ThoughtPrompting Elicits…Chain of Thought Prompting Elicits Reasoning in Large Language ModelsTraining a Helpful andHarmless Assistant with…Training a Helpful and Harmless Assistant with Reinforcement Learning from Human FeedbackLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsDiscovering LanguageModel Behaviors with…Discovering Language Model Behaviors with Model-Written EvaluationsEmergent LinearRepresentations in Worl…Emergent Linear Representations in World Models of Self-Supervised Sequence ModelsBackdoor ActivationAttack: Attack Large…Backdoor Activation Attack: Attack Large Language Models using Activation Steering for Safety-AlignmentLarge Language ModelAlignment: A SurveyLarge Language Model Alignment: A SurveyLearning InterpretableConcepts: Unifying…Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation ModelsTruth Forest: TowardMulti-Scale Truthfulnes…Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without TuningRAIN: Your LanguageModels Can Align…RAIN: Your Language Models Can Align Themselves without FinetuningCan AI Assistants KnowWhat They Don't Know?Can AI Assistants Know What They Don't Know?Weak-to-StrongGeneralization…Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionThe Dawn After the Dark:An Empirical Study on…The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsLies, Damned Lies, andDistributional Language…Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language ModelsA MechanisticUnderstanding of…A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityHaluEval-Wild:Evaluating…HaluEval-Wild: Evaluating Hallucinations of Language Models in the WildInference-TimeIntervention: Eliciting…Inference-Time Intervention: Eliciting Truthful Answers from a Language Model過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。