Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Recent advancements in long chain-of-thought (CoT) reasoning, particularly through the Group Relative Policy Optimization algorithm used by DeepSeek-R1, have led to significant interest in the potential of Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs). While RLVR promises to improve reasoning by allowing models to learn from free exploration, there remains debate over whether it truly enhances reasoning abilities or simply boosts sampling efficiency. This paper systematically investigates the impact of RLVR on LLM reasoning. We revisit Pass@K experiments and demonstrate that RLVR can extend the reasoning boundary for both mathematical and coding tasks. This is supported by our introduction of a novel evaluation metric, CoT-Pass@K, which captures reasoning success by accounting for both the final answer and intermediate reasoning steps. Furthermore, we present a theoretical framework explaining RLVR's incentive mechanism, demonstrating how it can encourage correct reasoning even when rewards are based solely on answer correctness. Our analysis of RLVR's training dynamics reveals that it incentivizes correct reasoning early in the process, with substantial improvements in reasoning quality confirmed through extensive evaluations. These findings provide strong evidence of RLVR's potential to enhance LLM reasoning, offering valuable insights into its mechanisms and performance improvements.

Does ReinforcementLearning Really…Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Reinforcement Learningfor Reasoning in Large…Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveBridging SupervisedLearning and…Bridging Supervised Learning and Reinforcement Learning in Math ReasoningAceReason-Nemotron:Advancing Math and Code…AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement LearningDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleOpen-Reasoner-Zero: AnOpen Source Approach to…Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelProRL: ProlongedReinforcement Learning…ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsRight Question isAlready Half the Answer…Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationRLVER: ReinforcementLearning with Verifiabl…RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic AgentsThe Debate on RLVRReasoning Capability…The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic ViewRL Squeezes, SFTExpands: A Comparative…RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMsStabilizing Knowledge,Promoting Reasoning…Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVRLaSeR: ReinforcementLearning with Last-Toke…LaSeR: Reinforcement Learning with Last-Token Self-RewardingQuestA: ExpandingReasoning Capacity in…QuestA: Expanding Reasoning Capacity in LLMs via Question AugmentationA Survey ofReinforcement Learning…A Survey of Reinforcement Learning for Large Reasoning ModelsUnveiling ImplicitAdvantage Symmetry: Why…Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty AdaptationPlacing Puzzle PiecesWhere They Matter: A…Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement LearningDoes Your ReasoningModel Implicitly Know…Does Your Reasoning Model Implicitly Know When to Stop Thinking?Training LLMs forDivide-and-Conquer…Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time ScalabilityTo Mix or To Merge:Toward Multi-Domain…To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language ModelsReinforcement Learningwith Verifiable Rewards…Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.