A Survey of Reinforcement Learning for Large Reasoning Models

In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontier of LLM capabilities, particularly in addressing complex logical tasks such as mathematics and coding. As a result, RL has emerged as a foundational methodology for transforming LLMs into LRMs. With the rapid progress of the field, further scaling of RL for LRMs now faces foundational challenges not only in computational resources but also in algorithm design, training data, and infrastructure. To this end, it is timely to revisit the development of this domain, reassess its trajectory, and explore strategies to enhance the scalability of RL toward Artificial SuperIntelligence (ASI). In particular, we examine research applying RL to LLMs and LRMs for reasoning abilities, especially since the release of DeepSeek-R1, including foundational components, core problems, training resources, and downstream applications, to identify future opportunities and directions for this rapidly evolving area. We hope this review will promote future research on RL for broader reasoning models. Github: https://github.com/TsinghuaC3I/Awesome-RL-for-LRMs

DeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleLogic-RL: Unleashing LLMReasoning with…Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement LearningSailing AI by the Stars:A Survey of Learning…Sailing AI by the Stars: A Survey of Learning from Rewards in Post-Training and Test-Time Scaling of Large Language ModelsLearning to Reason underOff-Policy GuidanceLearning to Reason under Off-Policy GuidanceSimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningMiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionGroup-in-Group PolicyOptimization for LLM…Group-in-Group Policy Optimization for LLM Agent TrainingCritique-GRPO: AdvancingLLM Reasoning with…Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical FeedbackSpurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVRVLM-R1: A Stable andGeneralizable R1-style…VLM-R1: A Stable and Generalizable R1-style Large Vision-Language ModelPass@k Training forAdaptively Balancing…Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning ModelsA Comprehensive Surveyof LLM Alignment…A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and MoreReinforcement LearningMeets Large Language…Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM LifecycleAdvancing MultimodalReasoning: From…Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement LearningSpotlight on TokenPerception for…Spotlight on Token Perception for Multimodal Reinforcement LearningAgent0: UnleashingSelf-Evolving Agents…Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated ReasoningDiffThinker: TowardsGenerative Multimodal…DiffThinker: Towards Generative Multimodal Reasoning with Diffusion ModelsRewardMap: TacklingSparse Rewards in…RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningMemory in the Age of AIAgentsMemory in the Age of AI AgentsAgentic Reasoning forLarge Language ModelsAgentic Reasoning for Large Language ModelsGenerate, Filter,Control, Replay: A…Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement LearningBeyond Mode Elicitation:Diversity-Preserving…Beyond Mode Elicitation: Diversity-Preserving Reinforcement Learning via Latent Diffusion ReasonerCog-DRIFT: Explorationon Adaptively…Cog-DRIFT: Exploration on Adaptively Reformulated Instances Enables Learning from Hard Reasoning ProblemsA Survey ofReinforcement Learning…A Survey of Reinforcement Learning for Large Reasoning Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。