Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Test-time inference has emerged as a powerful paradigm for enabling language models to ``think'' longer and more carefully about complex challenges, much like skilled human experts. While reinforcement learning (RL) can drive self-improvement in language models on verifiable tasks, some models exhibit substantial gains while others quickly plateau. For instance, we find that Qwen-2.5-3B far exceeds Llama-3.2-3B under identical RL training for the game of Countdown. This discrepancy raises a critical question: what intrinsic properties enable effective self-improvement? We introduce a framework to investigate this question by analyzing four key cognitive behaviors -- verification, backtracking, subgoal setting, and backward chaining -- that both expert human problem solvers and successful language models employ. Our study reveals that Qwen naturally exhibits these reasoning behaviors, whereas Llama initially lacks them. In systematic experimentation with controlled behavioral datasets, we find that priming Llama with examples containing these reasoning behaviors enables substantial improvements during RL, matching or exceeding Qwen's performance. Importantly, the presence of reasoning behaviors, rather than correctness of answers, proves to be the critical factor -- models primed with incorrect solutions containing proper reasoning patterns achieve comparable performance to those trained on correct solutions. Finally, leveraging continued pretraining with OpenWebMath data, filtered to amplify reasoning behaviors, enables the Llama model to match Qwen's self-improvement trajectory. Our findings establish a fundamental relationship between initial reasoning behaviors and the capacity for improvement, explaining why some language models effectively utilize additional computation while others plateau.

Fixing Weight DecayRegularization in AdamFixing Weight Decay Regularization in AdamLarge Language Monkeys:Scaling Inference…Large Language Monkeys: Scaling Inference Compute with Repeated SamplingOpenAI o1 System CardOpenAI o1 System CardTowards System 2Reasoning in LLMs…Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-ThoughtLLMs Can Easily Learn toReason from…LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!VinePPO: Unlocking RLPotential For LLM…VinePPO: Unlocking RL Potential For LLM Reasoning Through Refined Credit AssignmentDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningDemystifying LongChain-of-Thought…Demystifying Long Chain-of-Thought Reasoning in LLMsOn the Emergence ofThinking in LLMs I…On the Emergence of Thinking in LLMs I: Searching for the Right IntuitionTraining Language Modelsto Self-Correct via…Training Language Models to Self-Correct via Reinforcement LearningREINFORCE++: A Simpleand Efficient Approach…REINFORCE++: A Simple and Efficient Approach for Aligning Large Language ModelsProcess Reinforcementthrough Implicit RewardsProcess Reinforcement through Implicit RewardsThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningCrosslingual Reasoningthrough Test-Time…Crosslingual Reasoning through Test-Time ScalingDepth-Breadth Synergy inRLVR: Unlocking LLM…Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive ExplorationReMA: Learning toMeta-Think for LLMs wit…ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement LearningSpurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVRHow Much Backtracking isEnough? Exploring the…How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM ReasoningBehavior Injection:Preparing Language…Behavior Injection: Preparing Language Models for Reinforcement LearningCan Large ReasoningModels Self-Train?Can Large Reasoning Models Self-Train?Emerging Properties inUnified Multimodal…Emerging Properties in Unified Multimodal PretrainingCurriculum ReinforcementLearning from Easy to…Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningVCRL: Variance-basedCurriculum Reinforcemen…VCRL: Variance-based Curriculum Reinforcement Learning for Large Language ModelsThinkPrune: Pruning LongChain-of-Thought of LLM…ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement LearningCognitive Behaviors thatEnable Self-Improving…Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.