GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa .

Chain of ThoughtPrompting Elicits…Chain of Thought Prompting Elicits Reasoning in Large Language ModelsAutomatic PromptOptimization with…Automatic Prompt Optimization with "Gradient Descent" and Beam SearchOptimizing Instructionsand Demonstrations for…Optimizing Instructions and Demonstrations for Multi-Stage Language Model ProgramsLarge Language Models asOptimizersLarge Language Models as OptimizersDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsLLM-based Optimizationof Compound AI Systems…LLM-based Optimization of Compound AI Systems: A SurveyLLMs Are In-ContextReinforcement LearnersLLMs Are In-Context Reinforcement LearnersGeneralizing VerifiableInstruction FollowingGeneralizing Verifiable Instruction FollowingSearch-R1: Training LLMsto Reason and Leverage…Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningQwen3 Technical ReportQwen3 Technical ReportDynamic Cheatsheet:Test-Time Learning with…Dynamic Cheatsheet: Test-Time Learning with Adaptive MemoryMulti-module GRPO:Composing Policy…Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model ProgramsPromptBridge:Cross-Model Prompt…PromptBridge: Cross-Model Prompt Transfer for Large Language ModelsLet the Barbarians In:How AI Can Accelerate…Let the Barbarians In: How AI Can Accelerate Systems Performance ResearchAstra: A Multi-AgentSystem for GPU Kernel…Astra: A Multi-Agent System for GPU Kernel Performance OptimizationC-Evolve:Consensus-based…C-Evolve: Consensus-based Evolution for Prompt GroupsMaestro: Joint Graph &Config Optimization for…Maestro: Joint Graph & Config Optimization for Reliable AI AgentsCode-enabled languagemodels can outperform…Code-enabled language models can outperform reasoning models on diverse tasksReinforcement LearningImproves Traversal of…Reinforcement Learning Improves Traversal of Hierarchical Knowledge in LLMsp1: Better PromptOptimization with Fewer…p1: Better Prompt Optimization with Fewer PromptsGlia: A Human-InspiredAI for Automated System…Glia: A Human-Inspired AI for Automated Systems Design and OptimizationPACEvolve: EnablingLong-Horizon…PACEvolve: Enabling Long-Horizon Progress-Aware Consistent EvolutionLost in the Noise: HowReasoning Models Fail…Lost in the Noise: How Reasoning Models Fail with Contextual DistractorsEvolutionary SystemPrompt Learning for…Evolutionary System Prompt Learning for Reinforcement Learning in LLMsGEPA: Reflective PromptEvolution Can Outperfor…GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。