Detecting hallucinations in large language models using semantic entropy

Abstract Large language model (LLM) systems, such as ChatGPT 1 or Gemini 2 , can show impressive reasoning and question-answering capabilities but often ‘hallucinate’ false outputs and unsubstantiated answers 3,4 . Answering unreliably or without the necessary information prevents adoption in diverse fields, with problems including fabrication of legal precedents 5 or untrue facts in news articles 6 and even posing a risk to human life in medical domains such as radiology 7 . Encouraging truthfulness through supervision or reinforcement has been only partially successful 8 . Researchers need a general method for detecting hallucinations in LLMs that works even with new and unseen questions to which humans might not know the answer. Here we develop new methods grounded in statistics, proposing entropy-based uncertainty estimators for LLMs to detect a subset of hallucinations—confabulations—which are arbitrary and incorrect generations. Our method addresses the fact that one idea can be expressed in many ways by computing uncertainty at the level of meaning rather than specific sequences of words. Our method works across datasets and tasks without a priori knowledge of the task, requires no task-specific data and robustly generalizes to new tasks not seen before. By detecting when a prompt is likely to produce a confabulation, our method helps users understand when they must take extra care with LLMs and opens up new possibilities for using LLMs that are otherwise prevented by their unreliability.

On a Measure of theInformation Provided by…On a Measure of the Information Provided by an ExperimentAn overview of theBIOASQ large-scale…An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competitionInformation-BasedObjective Functions for…Information-Based Objective Functions for Active Data SelectionHallucinating FacesHallucinating FacesAleatory or epistemic?Does it matter?Aleatory or epistemic? Does it matter?Natural Questions: aBenchmark for Question…Natural Questions: a Benchmark for Question Answering ResearchRanking GeneratedSummaries by…Ranking Generated Summaries by Correctness: An Interesting but Challenging Application for Natural Language InferenceSummaC: Re-VisitingNLI-based Models for…SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in SummarizationGemini: A Family ofHighly Capable…Gemini: A Family of Highly Capable Multimodal ModelsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsChatGPT and Other LargeLanguage Models Are…ChatGPT and Other Large Language Models Are Double-edged SwordsThe RefinedWeb Datasetfor Falcon LLM…The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data OnlyA Framework to AssessClinical Safety and…A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text SummarisationThe Clinicians’ Guide toLarge Language Models…The Clinicians’ Guide to Large Language Models: A General Perspective With a Focus on HallucinationsArtificial intelligencefor modelling infectiou…Artificial intelligence for modelling infectious disease epidemicsMulti-model assuranceanalysis showing large…Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision supportFoundation models inbioinformaticsFoundation models in bioinformaticsWhat large languagemodels know and what…What large language models know and what people think they knowCtrlA: AdaptiveRetrieval-Augmented…CtrlA: Adaptive Retrieval-Augmented Generation via Probe-Guided ControlAI-driven platform forsystematic nomenclature…AI-driven platform for systematic nomenclature and intelligent knowledge acquisition of natural medicinal materialsRAGing ahead inrheumatology: new…RAGing ahead in rheumatology: new language model architectures to tame artificial intelligenceToward large reasoningmodels: A survey of…Toward large reasoning models: A survey of reinforced reasoning with large language modelsChatCNC: Conversationalmachine monitoring via…ChatCNC: Conversational machine monitoring via large language model and real-time data retrieval augmented generationLarge Language ModelsHallucination: A…Large Language Models Hallucination: A Comprehensive SurveyDetecting hallucinationsin large language model…Detecting hallucinations in large language models using semantic entropy過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。