Measuring short-form factuality in large language models

We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First, SimpleQA is challenging, as it is adversarially collected against GPT-4 responses. Second, responses are easy to grade, because questions are created such that there exists only a single, indisputable answer. Each answer in SimpleQA is graded as either correct, incorrect, or not attempted. A model with ideal behavior would get as many questions correct as possible while not attempting the questions for which it is not confident it knows the correct answer. SimpleQA is a simple, targeted evaluation for whether models "know what they know," and our hope is that this benchmark will remain relevant for the next few generations of frontier models. SimpleQA can be found at https://github.com/openai/simple-evals.

Language Models (Mostly)Know What They KnowLanguage Models (Mostly) Know What They KnowTeaching Models toExpress Their…Teaching Models to Express Their Uncertainty in WordsSelf-ConsistencyImproves Chain of…Self-Consistency Improves Chain of Thought Reasoning in Language ModelsEvaluatingHallucinations in…Evaluating Hallucinations in Chinese Large Language ModelsFreshLLMs: RefreshingLarge Language Models…FreshLLMs: Refreshing Large Language Models with Search Engine AugmentationLong-form factuality inlarge language modelsLong-form factuality in large language modelsFact, Fetch, and Reason:A Unified Evaluation of…Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented GenerationO1 Replication Journey -Part 2: Surpassing…O1 Replication Journey - Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?SimpleQA Verified: AReliable Factuality…SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric KnowledgeMemento: Fine-tuning LLMAgents without…Memento: Fine-tuning LLM Agents without Fine-tuning LLMsDR Tulu: ReinforcementLearning with Evolving…DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep ResearchSearch Arena: AnalyzingSearch-Augmented LLMsSearch Arena: Analyzing Search-Augmented LLMsA Comprehensive Surveyon Trustworthiness in…A Comprehensive Survey on Trustworthiness in Reasoning with Large Language ModelsOlmo 3Olmo 3MiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionBeyond Binary Rewards:Training LMs to Reason…Beyond Binary Rewards: Training LMs to Reason About Their UncertaintyCWM: An Open-Weights LLMfor Research on Code…CWM: An Open-Weights LLM for Research on Code Generation with World ModelsDeepSearchQA: Bridgingthe Comprehensiveness…DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research AgentsHallucinations UndermineTrust; Metacognition is…Hallucinations Undermine Trust; Metacognition is a Way ForwardMeasuring short-formfactuality in large…Measuring short-form factuality in large language modelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.