Do Large Language Model Benchmarks Test Reliability?

When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus on measuring their reliability. To understand the potential ramifications of this gap, we investigate how well current benchmarks quantify model reliability. We find that pervasive label errors can compromise these evaluations, obscuring lingering model failures and hiding unreliable behavior. Motivated by this gap in the evaluation of reliability, we then propose the concept of so-called platinum benchmarks, i.e., benchmarks carefully curated to minimize label errors and ambiguity. As a first attempt at constructing such benchmarks, we revise examples from fifteen existing popular benchmarks. We evaluate a wide range of models on these platinum benchmarks and find that, indeed, frontier LLMs still exhibit failures on simple tasks such as elementary-level math word problems. Analyzing these failures further reveals previously unidentified patterns of problems on which frontier models consistently struggle. We provide code at https://github.com/MadryLab/platinum-benchmarks

HotpotQA: A Dataset forDiverse, Explainable…HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question AnsweringTraining Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsMeasuring MassiveMultitask Language…Measuring Massive Multitask Language UnderstandingGPQA: A Graduate-LevelGoogle-Proof Q&A…GPQA: A Graduate-Level Google-Proof Q&A BenchmarkRetrieval-AugmentedGeneration for Large…Retrieval-Augmented Generation for Large Language Models: A SurveyGPT-4 Technical ReportGPT-4 Technical ReportThe Llama 3 Herd ofModelsThe Llama 3 Herd of ModelsDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportGemini 1.5: Unlockingmultimodal understandin…Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextDeepSeek-Coder: When theLarge Language Model…DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code IntelligenceDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningAre We Done with MMLU?Are We Done with MMLU?Kimi K2: Open AgenticIntelligenceKimi K2: Open Agentic IntelligenceOlmo 3Olmo 3Signal and Noise: AFramework for Reducing…Signal and Noise: A Framework for Reducing Uncertainty in Language Model EvaluationThe Leaderboard IllusionThe Leaderboard IllusionAntidistillationSamplingAntidistillation SamplingRethinking Reflection inPre-TrainingRethinking Reflection in Pre-TrainingThe Illusion ofDiminishing Returns…The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsFantastic Bugs and Whereto Find Them in AI…Fantastic Bugs and Where to Find Them in AI BenchmarksEvaluating LLM MetricsThrough Real-World…Evaluating LLM Metrics Through Real-World CapabilitiesFrom KMMLU-Redux toKMMLU-Pro: A…From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM EvaluationBridging the Gap BetweenPromise and Performance…Bridging the Gap Between Promise and Performance for Microscaling FP4 QuantizationLet's Verify MathQuestions Step by StepLet's Verify Math Questions Step by StepDo Large Language ModelBenchmarks Test…Do Large Language Model Benchmarks Test Reliability?過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。