Understanding R1-Zero-Like Training: A Critical Perspective

DeepSeek-R1-Zero has shown that reinforcement learning (RL) at scale can directly enhance the reasoning capabilities of LLMs without supervised fine-tuning. In this work, we critically examine R1-Zero-like training by analyzing its two core components: base models and RL. We investigate a wide range of base models, including DeepSeek-V3-Base, to understand how pretraining characteristics influence RL performance. Our analysis reveals that DeepSeek-V3-Base already exhibit ''Aha moment'', while Qwen2.5 base models demonstrate strong reasoning capabilities even without prompt templates, suggesting potential pretraining biases. Additionally, we identify an optimization bias in Group Relative Policy Optimization (GRPO), which artificially increases response length (especially for incorrect outputs) during training. To address this, we introduce Dr. GRPO, an unbiased optimization method that improves token efficiency while maintaining reasoning performance. Leveraging these insights, we present a minimalist R1-Zero recipe that achieves 43.3% accuracy on AIME 2024 with a 7B base model, establishing a new state-of-the-art. Our code is available at https://github.com/sail-sg/understand-r1-zero.

Measuring MathematicalProblem Solving With th…Measuring Mathematical Problem Solving With the MATH DatasetQwen2.5-Math TechnicalReport: Toward…Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-ImprovementDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsOlympiadBench: AChallenging Benchmark…OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsOpenRLHF: AnEasy-to-use, Scalable…OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF FrameworkBack to Basics:Revisiting REINFORCE…Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsTÜLU 3: PushingFrontiers in Open…TÜLU 3: Pushing Frontiers in Open Language Model Post-TrainingQwen2.5 Technical ReportQwen2.5 Technical ReportDiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)DeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningHybridFlow: A Flexibleand Efficient RLHF…HybridFlow: A Flexible and Efficient RLHF FrameworkProcess Reinforcementthrough Implicit RewardsProcess Reinforcement through Implicit RewardsMiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionGroup-in-Group PolicyOptimization for LLM…Group-in-Group Policy Optimization for LLM Agent TrainingLearning to Reason underOff-Policy GuidanceLearning to Reason under Off-Policy GuidanceThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningDCPO: Dynamic ClippingPolicy OptimizationDCPO: Dynamic Clipping Policy OptimizationSpurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVRLearn the Ropes, ThenTrust the Wins…Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningVLA-RL: TowardsMasterful and General…VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningSimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningVLM-R1: A Stable andGeneralizable R1-style…VLM-R1: A Stable and Generalizable R1-style Large Vision-Language ModelGoedel-Prover-V2:Scaling Formal Theorem…Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-CorrectionOn the Design ofKL-Regularized Policy…On the Design of KL-Regularized Policy Gradient Algorithms for LLM ReasoningUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.