Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated notable success in enhancing the reasoning performance of large language models (LLMs), particularly on mathematics and programming tasks. Similar to how traditional RL helps agents explore and learn new strategies, RLVR is believed to enable LLMs to continuously self-improve, thus acquiring novel reasoning abilities beyond those of the corresponding base models. In this study we critically examine the current state of RLVR by systematically probing the reasoning capability boundaries of RLVR-trained LLMs across various model families, RL algorithms, and math, coding, and visual reasoning benchmarks, using pass@k at large k values as the evaluation metric. Surprisingly, we find that the current training setup does not elicit fundamentally new reasoning patterns. While RLVR-trained models outperform their base models at small k (e.g., k = 1), the base models achieve a higher pass@k score when k is large. Coverage and perplexity analyses show that the observed reasoning abilities originate from and are bounded by the base model. Treating the base model as an upper bound, our quantitative analysis shows that six popular RLVR algorithms perform similarly and remain far from optimal in leveraging the potential of the base model. By contrast, we find that distillation can introduce new reasoning patterns from the teacher and genuinely expand the model's reasoning capabilities. Overall, our findings suggest that current RLVR methods have not yet realized the potential of RL to elicit truly novel reasoning abilities in LLMs. This highlights the need for improved RL paradigms, such as continual scaling and multi-turn agent-environment interaction, to unlock this potential.

DeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveEcho Chamber: RLPost-training Amplifies…Echo Chamber: RL Post-training Amplifies Behaviors Learned in PretrainingDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningREINFORCE++: A Simpleand Efficient Approach…REINFORCE++: A Simple and Efficient Approach for Aligning Large Language ModelsAceReason-Nemotron:Advancing Math and Code…AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement LearningHybridFlow: A Flexibleand Efficient RLHF…HybridFlow: A Flexible and Efficient RLHF FrameworkQwen3 Technical ReportQwen3 Technical ReportKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsStepHint: Multi-levelStepwise Hints Enhance…StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to ReasonThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningAbsolute Zero:Reinforced Self-play…Absolute Zero: Reinforced Self-play Reasoning with Zero DataSpurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVRReinforcement Learningvs. Distillation…Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM ReasoningOn the Mechanism ofReasoning Pattern…On the Mechanism of Reasoning Pattern Selection in Reinforcement Learning for Language ModelsSwS: Self-awareWeakness-driven Problem…SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM ReasoningRisk-Sensitive RL forAlleviating Exploration…Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language ModelsCurriculum ReinforcementLearning from Easy to…Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningAgent0: UnleashingSelf-Evolving Agents…Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated ReasoningGoedel-Prover-V2:Scaling Formal Theorem…Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-CorrectionRESTRAIN: From SpuriousVotes to Signals -…RESTRAIN: From Spurious Votes to Signals - Self-Driven RL with Self-PenalizationScaling Reasoning Tokensvia RL and Parallel…Scaling Reasoning Tokens via RL and Parallel Thinking: Evidence From Competitive ProgrammingDoes ReinforcementLearning Really…Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Earlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.