GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality. We analyze the potential for frontier models, when paired with human oversight, to perform GDPval tasks cheaper and faster than unaided experts. We also demonstrate that increased reasoning effort, increased task context, and increased scaffolding improves model performance on GDPval. Finally, we open-source a gold subset of 220 tasks and provide a public automated grading service at evals.openai.com to facilitate future research in understanding real-world model capabilities.

Measuring MassiveMultitask Language…Measuring Massive Multitask Language UnderstandingGPQA: A Graduate-LevelGoogle-Proof Q&A…GPQA: A Graduate-Level Google-Proof Q&A BenchmarkAgentBench: EvaluatingLLMs as AgentsAgentBench: Evaluating LLMs as AgentsLLM Evaluators Recognizeand Favor Their Own…LLM Evaluators Recognize and Favor Their Own GenerationsClio: Privacy-PreservingInsights into Real-Worl…Clio: Privacy-Preserving Insights into Real-World AI UseSWE-Lancer: Can FrontierLLMs Earn 1 Million fro…SWE-Lancer: Can Frontier LLMs Earn 1 Million from Real-World Freelance Software Engineering?Humanity's Last ExamHumanity's Last ExamRemote Labor Index:Measuring AI Automation…Remote Labor Index: Measuring AI Automation of Remote WorkAI-Trader: BenchmarkingAutonomous Agents in…AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial MarketsUpBench: A DynamicallyEvolving Real-World…UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AIPRBench: Large-ScaleExpert Rubrics for…PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional ReasoningGLM-5: from Vibe Codingto Agentic EngineeringGLM-5: from Vibe Coding to Agentic EngineeringKimi K2.5: VisualAgentic IntelligenceKimi K2.5: Visual Agentic IntelligenceJobBench: Aligning AgentWork With Human WillJobBench: Aligning Agent Work With Human WillBigFinanceBench: AWorkflow-Grounded…BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research AgentsFrontierFinance: ALong-Horizon…FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial TasksThe 2025 AI Agent Index:Documenting Technical…The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI SystemsSafePro: Evaluating theSafety of…SafePro: Evaluating the Safety of Professional-Level AI AgentsOfficeQA Pro: AnEnterprise Benchmark fo…OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded ReasoningGDPval: Evaluating AIModel Performance on…GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.