Process Reinforcement through Implicit Rewards

Dense process rewards have proven a more effective alternative to the sparse outcome-level rewards in the inference-time scaling of large language models (LLMs), particularly in tasks requiring complex multi-step reasoning. While dense rewards also offer an appealing choice for the reinforcement learning (RL) of LLMs since their fine-grained rewards have the potential to address some inherent issues of outcome rewards, such as training efficiency and credit assignment, this potential remains largely unrealized. This can be primarily attributed to the challenges of training process reward models (PRMs) online, where collecting high-quality process labels is prohibitively expensive, making them particularly vulnerable to reward hacking. To address these challenges, we propose PRIME (Process Reinforcement through IMplicit rEwards), which enables online PRM updates using only policy rollouts and outcome labels through implict process rewards. PRIME combines well with various advantage functions and forgoes the dedicated reward model training phrase that existing approaches require, substantially reducing the development overhead. We demonstrate PRIME's effectiveness on competitional math and coding. Starting from Qwen2.5-Math-7B-Base, PRIME achieves a 15.1% average improvement across several key reasoning benchmarks over the SFT model. Notably, our resulting model, Eurus-2-7B-PRIME, surpasses Qwen2.5-Math-7B-Instruct on seven reasoning benchmarks with 10% of its training data.

Solving math wordproblems with process-…Solving math word problems with process- and outcome-based feedbackTACO: Topics inAlgorithmic COde…TACO: Topics in Algorithmic COde generation datasetQwen2.5-Math TechnicalReport: Toward…Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-ImprovementOlympiadBench: AChallenging Benchmark…OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsBack to Basics:Revisiting REINFORCE…Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsVinePPO: Unlocking RLPotential For LLM…VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit AssignmentRewarding Progress:Scaling Automated…Rewarding Progress: Scaling Automated Process Verifiers for LLM ReasoningFree Process Rewardswithout Process LabelsFree Process Rewards without Process LabelsKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsAceMath: AdvancingFrontier Math Reasoning…AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward ModelingDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningHybridFlow: A Flexibleand Efficient RLHF…HybridFlow: A Flexible and Efficient RLHF FrameworkLearning to Reason underOff-Policy GuidanceLearning to Reason under Off-Policy GuidanceDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleAbsolute Zero:Reinforced Self-play…Absolute Zero: Reinforced Self-play Reasoning with Zero DataS2R: Teaching LLMs toSelf-verify and…S2R: Teaching LLMs to Self-verify and Self-correct via Reinforcement LearningReasonFlux-PRM:Trajectory-Aware PRMs…ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMsSwS: Self-awareWeakness-driven Problem…SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM ReasoningFirst Return,Entropy-Eliciting…First Return, Entropy-Eliciting ExploreSimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningEfficient ReinforcementFinetuning via Adaptive…Efficient Reinforcement Finetuning via Adaptive Curriculum LearningRL-PLUS: CounteringCapability Boundary…RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy OptimizationCan Prompt Difficulty beOnline Predicted for…Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?Not All Rollouts areUseful: Down-Sampling…Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement LearningProcess Reinforcementthrough Implicit RewardsProcess Reinforcement through Implicit RewardsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.