Spurious Rewards: Rethinking Training Signals in RLVR

We show that reinforcement learning with verifiable rewards (RLVR) can elicit strong mathematical reasoning in certain language models even with spurious rewards that have little, no, or even negative correlation with the correct answer. For example, RLVR training with GRPO improves MATH-500 performance for Qwen2.5-Math-7B by 21.4 percentage points using randomly assigned rewards, nearly matching the 29.1-point gain from ground-truth rewards. To explain this counterintuitive observation, we show that GRPO exhibits a clipping bias from the clip term, which can amplify high-prior behaviors learned during pretraining even without informative rewards. As a case study, we identify one such behavior in Qwen2.5-Math models, which we call code reasoning -- reasoning in code without actual code execution; code-reasoning frequency increases from 65 percent to over 90 percent with spurious rewards. However, the presence of such amplifiable behaviors is highly model-dependent. In practice, spurious rewards that are effective for Qwen models often fail to produce gains for other model families, such as Llama3 or OLMo2. Our results highlight the importance of validating RL methods across diverse models rather than relying on a single de facto choice: large gains can arise on Qwen models even from random rewards that do not reflect genuine capability improvements.

DeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsLearning to Reasonwithout External RewardsLearning to Reason without External RewardsTTRL: Test-TimeReinforcement LearningTTRL: Test-Time Reinforcement LearningEcho Chamber: RLPost-training Amplifies…Echo Chamber: RL Post-training Amplifies Behaviors Learned in PretrainingDoes ReinforcementLearning Really…Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Beyond the 80/20 Rule:High-Entropy Minority…Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildRight Question isAlready Half the Answer…Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationOpen-Reasoner-Zero: AnOpen Source Approach to…Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveDeepSeek-R1 incentivizesreasoning in LLMs…DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learningCognitive Behaviors thatEnable Self-Improving…Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsTTRL: Test-TimeReinforcement LearningTTRL: Test-Time Reinforcement LearningCan Large ReasoningModels Self-Train?Can Large Reasoning Models Self-Train?Self-QuestioningLanguage ModelsSelf-Questioning Language ModelsReasoning with Sampling:Your Base Model is…Reasoning with Sampling: Your Base Model is Smarter Than You ThinkNudging the Boundariesof LLM ReasoningNudging the Boundaries of LLM ReasoningReinforcement Learningvs. Distillation…Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM ReasoningA Model Can Help Itself:Reward-Free…A Model Can Help Itself: Reward-Free Self-Training for LLM ReasoningOn the Mechanism ofReasoning Pattern…On the Mechanism of Reasoning Pattern Selection in Reinforcement Learning for Language ModelsRL Squeezes, SFTExpands: A Comparative…RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMsExploration vsExploitation: Rethinkin…Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardMitigating Think-AnswerMismatch in LLM…Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage ReweightingHow Far Can UnsupervisedRLVR Scale LLM Training?How Far Can Unsupervised RLVR Scale LLM Training?Spurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVREarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.