SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

DeepSeek-R1 has shown that long chain-of-thought (CoT) reasoning can naturally emerge through a simple reinforcement learning (RL) framework with rule-based rewards, where the training may directly start from the base models-a paradigm referred to as zero RL training. Most recent efforts to reproduce zero RL training have primarily focused on the Qwen2.5 model series, which may not be representative as we find the base models already exhibit strong instruction-following and self-reflection abilities. In this work, we investigate zero RL training across 10 diverse base models, spanning different families and sizes including LLama3-8B, Mistral-7B/24B, DeepSeek-Math-7B, Qwen2.5-math-7B, and all Qwen2.5 models from 0.5B to 32B. Leveraging several key design strategies-such as adjusting format reward and controlling query difficulty-we achieve substantial improvements in both reasoning accuracy and response length across most settings. However, by carefully monitoring the training dynamics, we observe that different base models exhibit distinct patterns during training. For instance, the increased response length does not always correlate with the emergence of certain cognitive behaviors such as verification (i.e., the "aha moment"). Notably, we observe the "aha moment" for the first time in small models not from the Qwen family. We share the key designs that enable successful zero RL training, along with our findings and practices. To facilitate further research, we open-source the code, models, and analysis tools.

DeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportOlympiadBench: AChallenging Benchmark…OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsMath-Shepherd: Verifyand Reinforce LLMs…Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsDo NOT Think That Muchfor 2+3=? On the…Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMsOpenAI o1 System CardOpenAI o1 System CardLogic-RL: Unleashing LLMReasoning with…Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement LearningDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleCognitive Behaviors thatEnable Self-Improving…Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsDemystifying LongChain-of-Thought…Demystifying Long Chain-of-Thought Reasoning in LLMsDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningAbsolute Zero:Reinforced Self-play…Absolute Zero: Reinforced Self-play Reasoning with Zero DataCurriculum ReinforcementLearning from Easy to…Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningOn the Mechanism ofReasoning Pattern…On the Mechanism of Reasoning Pattern Selection in Reinforcement Learning for Language ModelsSpurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVRARM: Adaptive ReasoningModelARM: Adaptive Reasoning ModelMiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionSwS: Self-awareWeakness-driven Problem…SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM ReasoningCoRT: Code-integratedReasoning within…CoRT: Code-integrated Reasoning within ThinkingReinforcement Learningvs. Distillation…Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM ReasoningThinkPrune: Pruning LongChain-of-Thought of LLM…ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement LearningF-GRPO: Don't Let YourPolicy Learn the Obviou…F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the RareSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。