DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the $\textbf{D}$ecoupled Clip and $\textbf{D}$ynamic s$\textbf{A}$mpling $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{DAPO}$) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL.

GPT-4 Technical ReportGPT-4 Technical ReportDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsQwen2.5 Technical ReportQwen2.5 Technical ReportDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportDiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)HybridFlow: A Flexibleand Efficient RLHF…HybridFlow: A Flexible and Efficient RLHF FrameworkKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsREINFORCE++: A Simpleand Efficient Approach…REINFORCE++: A Simple and Efficient Approach for Aligning Large Language ModelsDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningVinePPO: Unlocking RLPotential For LLM…VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit AssignmentWhat's Behind PPO'sCollapse in Long-CoT?…What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the SecretProcess Reinforcementthrough Implicit RewardsProcess Reinforcement through Implicit RewardsThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningSimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningFirst Return,Entropy-Eliciting…First Return, Entropy-Eliciting ExploreOn the Design ofKL-Regularized Policy…On the Design of KL-Regularized Policy Gradient Algorithms for LLM ReasoningCoRT: Code-integratedReasoning within…CoRT: Code-integrated Reasoning within ThinkingVLA-RL: TowardsMasterful and General…VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningHow to Train a Leader:Hierarchical Reasoning…How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMsGHPO: Adaptive Guidancefor Stable and Efficien…GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement LearningExplore Data Left Behindin Reinforcement…Explore Data Left Behind in Reinforcement Learning for Reasoning Language ModelsLook Back to ReasonForward: Revisitable…Look Back to Reason Forward: Revisitable Memory for Long-Context LLM AgentsSkywork-R1V3 TechnicalReportSkywork-R1V3 Technical ReportLow-probability TokensSustain Exploration in…Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable RewardDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at Scale過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。