Reinforcement Learning for Reasoning in Large Language Models with One Training Example

We show that reinforcement learning with verifiable reward using one training example (1-shot RLVR) is effective in incentivizing the math reasoning capabilities of large language models (LLMs). Applying RLVR to the base model Qwen2.5-Math-1.5B, we identify a single example that elevates model performance on MATH500 from 36.0% to 73.6% (8.6% improvement beyond format correction), and improves the average performance across six common mathematical reasoning benchmarks from 17.6% to 35.7% (7.0% non-format gain). This result matches the performance obtained using the 1.2k DeepScaleR subset (MATH500: 73.6%, average: 35.9%), which contains the aforementioned example. Furthermore, RLVR with only two examples even slightly exceeds these results (MATH500: 74.8%, average: 36.6%). Similar substantial improvements are observed across various models (Qwen2.5-Math-7B, Llama3.2-3B-Instruct, DeepSeek-R1-Distill-Qwen-1.5B), RL algorithms (GRPO and PPO), and different math examples. In addition, we identify some interesting phenomena during 1-shot RLVR, including cross-category generalization, increased frequency of self-reflection, and sustained test performance improvement even after the training accuracy has saturated, a phenomenon we term post-saturation generalization. Moreover, we verify that the effectiveness of 1-shot RLVR primarily arises from the policy gradient loss, distinguishing it from the "grokking" phenomenon. We also show the critical role of promoting exploration (e.g., by incorporating entropy loss with an appropriate coefficient) in 1-shot RLVR training. We also further discuss related observations about format correction, label robustness and prompt modification. These findings can inspire future work on RLVR efficiency and encourage a re-examination of recent progress and the underlying mechanisms in RLVR. All resources are open source at https://github.com/ypwang61/One-Shot-RLVR.

VAPO: Efficient andReliable Reinforcement…VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning TasksDoes ReinforcementLearning Really…Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Absolute Zero:Reinforced Self-play…Absolute Zero: Reinforced Self-play Reasoning with Zero DataTTRL: Test-TimeReinforcement LearningTTRL: Test-Time Reinforcement LearningSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildLight-R1: CurriculumSFT, DPO and RL for Lon…Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and BeyondUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveConcise Reasoning viaReinforcement LearningConcise Reasoning via Reinforcement LearningWhat's Behind PPO'sCollapse in Long-CoT?…What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the SecretDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleCognitive Behaviors thatEnable Self-Improving…Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsRight Question isAlready Half the Answer…Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationAct Only When It Pays:Efficient Reinforcement…Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningAbsolute Zero:Reinforced Self-play…Absolute Zero: Reinforced Self-play Reasoning with Zero DataProsperity beforeCollapse: How Far Can…Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?Nudging the Boundariesof LLM ReasoningNudging the Boundaries of LLM ReasoningReinforcement Learningvs. Distillation…Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM ReasoningRESTRAIN: From SpuriousVotes to Signals -…RESTRAIN: From Spurious Votes to Signals - Self-Driven RL with Self-PenalizationCan Prompt Difficulty beOnline Predicted for…Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?Efficient ReinforcementFinetuning via Adaptive…Efficient Reinforcement Finetuning via Adaptive Curriculum LearningRL-PLUS: CounteringCapability Boundary…RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy OptimizationUnbiased Dynamic Pruningfor Efficient…Unbiased Dynamic Pruning for Efficient Group-Based Policy OptimizationUnlocking Exploration inRLVR: Uncertainty-aware…Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper ReasoningReinforcement Learningfor Reasoning in Large…Reinforcement Learning for Reasoning in Large Language Models with One Training Example過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。