LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have demonstrated that LLMs' internal states encode information regarding the truthfulness of their outputs, and that this information can be utilized to detect errors. In this work, we show that the internal representations of LLMs encode much more information about truthfulness than previously recognized. We first discover that the truthfulness information is concentrated in specific tokens, and leveraging this property significantly enhances error detection performance. Yet, we show that such error detectors fail to generalize across datasets, implying that -- contrary to prior claims -- truthfulness encoding is not universal but rather multifaceted. Next, we show that internal representations can also be used for predicting the types of errors the model is likely to make, facilitating the development of tailored mitigation strategies. Lastly, we reveal a discrepancy between LLMs' internal encoding and external behavior: they may encode the correct answer, yet consistently generate an incorrect one. Taken together, these insights deepen our understanding of LLM errors from the model's internal perspective, which can guide future research on enhancing error analysis and mitigation.

Q^2: Evaluating FactualConsistency in…Q^2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringDiscovering LatentKnowledge in Language…Discovering Latent Knowledge in Language Models Without SupervisionThe Troubling Emergenceof Hallucination in…The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive RemediationsLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsCognitive Dissonance:Why Do Language Model…Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?Mistral 7BMistral 7BLook Before You Leap: AnExploratory Study of…Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language ModelsCan Large LanguageModels Faithfully…Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?Truth is Universal:Robust Detection of Lie…Truth is Universal: Robust Detection of Lies in LLMsOn Early Detection ofHallucinations in…On Early Detection of Hallucinations in Factual Question AnsweringInside-Out: HiddenFactual Knowledge in…Inside-Out: Hidden Factual Knowledge in LLMsConfidence ImprovesSelf-Consistency in LLMsConfidence Improves Self-Consistency in LLMsBeyond Token Probes:Hallucination Detection…Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViTNeural Message-Passingon Attention Graphs for…Neural Message-Passing on Attention Graphs for Hallucination DetectionLLM Microscope: WhatModel Internals Reveal…LLM Microscope: What Model Internals Reveal About Answer Correctness and Context UtilizationCalibrating VerbalUncertainty as a Linear…Calibrating Verbal Uncertainty as a Linear Feature to Reduce HallucinationsTrust Me, I'm Wrong:LLMs Hallucinate with…Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the AnswerA comprehensive taxonomyof hallucinations in…A comprehensive taxonomy of hallucinations in Large Language ModelsRepresentationEngineering for…Representation Engineering for Large-Language Models: Survey and Research ChallengesCoT-UQ: ImprovingResponse-wise…CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-ThoughtPersona Features ControlEmergent MisalignmentPersona Features Control Emergent MisalignmentLoki's Dance ofIllusions: A…Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language ModelsBeyond Next TokenProbabilities…Beyond Next Token Probabilities: Learnable, Fast Detection of Hallucinations and Data Contamination on LLM Output DistributionsPARALLAX: SeparatingGenuine Hallucination…PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction ArtifactsLLMs Know More Than TheyShow: On the Intrinsic…LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.