ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Recent advances in reasoning-centric language models have highlighted reinforcement learning (RL) as a promising method for aligning models with verifiable rewards. However, it remains contentious whether RL truly expands a model's reasoning capabilities or merely amplifies high-reward outputs already latent in the base model's distribution, and whether continually scaling up RL compute reliably leads to improved reasoning performance. In this work, we challenge prevailing assumptions by demonstrating that prolonged RL (ProRL) training can uncover novel reasoning strategies that are inaccessible to base models, even under extensive sampling. We introduce ProRL, a novel training methodology that incorporates KL divergence control, reference policy resetting, and a diverse suite of tasks. Our empirical analysis reveals that RL-trained models consistently outperform base models across a wide range of pass@k evaluations, including scenarios where base models fail entirely regardless of the number of attempts. We further show that reasoning boundary improvements correlates strongly with task competence of base model and training duration, suggesting that RL can explore and populate new regions of solution space over time. These findings offer new insights into the conditions under which RL meaningfully expands reasoning boundaries in language models and establish a foundation for future work on long-horizon RL for reasoning. We release model weights to support further research: https://huggingface.co/nvidia/Nemotron-Research-Reasoning-Qwen-1.5B

Concrete Problems in AISafetyConcrete Problems in AI SafetyOpenAI o1 System CardOpenAI o1 System CardDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsQuiet-STaR: LanguageModels Can Teach…Quiet-STaR: Language Models Can Teach Themselves to Think Before SpeakingDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningCognitive Behaviors thatEnable Self-Improving…Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsSFT Memorizes, RLGeneralizes: A…SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingLanguage Models Learn toMislead Humans via RLHFLanguage Models Learn to Mislead Humans via RLHFProcess Reinforcementthrough Implicit RewardsProcess Reinforcement through Implicit RewardsQuestA: ExpandingReasoning Capacity in…QuestA: Expanding Reasoning Capacity in LLMs via Question AugmentationBroRL: ScalingReinforcement Learning…BroRL: Scaling Reinforcement Learning via Broadened ExplorationNudging the Boundariesof LLM ReasoningNudging the Boundaries of LLM ReasoningRLVE: Scaling UpReinforcement Learning…RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable EnvironmentsJointly ReinforcingDiversity and Quality i…Jointly Reinforcing Diversity and Quality in Language Model GenerationsOMEGA: Can LLMs ReasonOutside the Box in Math…OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative GeneralizationA Survey ofReinforcement Learning…A Survey of Reinforcement Learning for Large Reasoning ModelsRethinking EntropyInterventions in RLVR…Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveMaximum LikelihoodReinforcement LearningMaximum Likelihood Reinforcement LearningPOPE: Learning to Reasonon Hard Problems via…POPE: Learning to Reason on Hard Problems via Privileged On-Policy ExplorationUnveiling ImplicitAdvantage Symmetry: Why…Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty AdaptationCan Prompt Difficulty beOnline Predicted for…Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?ProRL: ProlongedReinforcement Learning…ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。