Demystifying Long Chain-of-Thought Reasoning in LLMs

Scaling inference compute enhances reasoning in large language models (LLMs), with long chains-of-thought (CoTs) enabling strategies like backtracking and error correction. Reinforcement learning (RL) has emerged as a crucial method for developing these capabilities, yet the conditions under which long CoTs emerge remain unclear, and RL training requires careful design choices. In this study, we systematically investigate the mechanics of long CoT reasoning, identifying the key factors that enable models to generate long CoT trajectories. Through extensive supervised fine-tuning (SFT) and RL experiments, we present four main findings: (1) While SFT is not strictly necessary, it simplifies training and improves efficiency; (2) Reasoning capabilities tend to emerge with increased training compute, but their development is not guaranteed, making reward shaping crucial for stabilizing CoT length growth; (3) Scaling verifiable reward signals is critical for RL. We find that leveraging noisy, web-extracted solutions with filtering mechanisms shows strong potential, particularly for out-of-distribution (OOD) tasks such as STEM reasoning; and (4) Core abilities like error correction are inherently present in base models, but incentivizing these skills effectively for complex tasks via RL demands significant compute, and measuring their emergence requires a nuanced approach. These insights provide practical guidance for optimizing training strategies to enhance long CoT reasoning in LLMs. Our code is available at: https://github.com/eddycmu/demystify-long-cot.

On the resemblance andcontainment of documentsOn the resemblance and containment of documentsTraining Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsReinforced Self-Training(ReST) for Language…Reinforced Self-Training (ReST) for Language ModelingScaling Relationship onLearning Mathematical…Scaling Relationship on Learning Mathematical Reasoning with Large Language ModelsGPT-4 Technical ReportGPT-4 Technical ReportBeyond Human Data:Scaling Self-Training…Beyond Human Data: Scaling Self-Training for Problem-Solving with Language ModelsREINFORCE++: A Simpleand Efficient Approach…REINFORCE++: A Simple and Efficient Approach for Aligning Large Language ModelsSelf-Training ElicitsConcise Reasoning in…Self-Training Elicits Concise Reasoning in Large Language ModelsTowards Concise andAdaptive Thinking in…Towards Concise and Adaptive Thinking in Large Reasoning Models: A SurveyStop Overthinking: ASurvey on Efficient…Stop Overthinking: A Survey on Efficient Reasoning for Large Language ModelsLong-ShortChain-of-Thought Mixtur…Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language ModelsUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveSmall Models Struggle toLearn from Strong…Small Models Struggle to Learn from Strong ReasonersLearn to ReasonEfficiently with…Learn to Reason Efficiently with Adaptive Length-based Reward ShapingThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningSwS: Self-awareWeakness-driven Problem…SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM ReasoningSWE-RL: Advancing LLMReasoning via…SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution100 Days AfterDeepSeek-R1: A Survey o…100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language ModelsAdaCtrl: TowardsAdaptive and…AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware BudgetingDemystifying LongChain-of-Thought…Demystifying Long Chain-of-Thought Reasoning in LLMs過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。