The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

This paper aims to overcome a major obstacle in scaling RL for reasoning with LLMs, namely the collapse of policy entropy. Such phenomenon is consistently observed across vast RL runs without entropy intervention, where the policy entropy dropped sharply at the early training stage, this diminished exploratory ability is always accompanied with the saturation of policy performance. In practice, we establish a transformation equation R=-a*e^H+b between entropy H and downstream performance R. This empirical law strongly indicates that, the policy performance is traded from policy entropy, thus bottlenecked by its exhaustion, and the ceiling is fully predictable H=0, R=-a+b. Our finding necessitates entropy management for continuous exploration toward scaling compute for RL. To this end, we investigate entropy dynamics both theoretically and empirically. Our derivation highlights that, the change in policy entropy is driven by the covariance between action probability and the change in logits, which is proportional to its advantage when using Policy Gradient-like algorithms. Empirical study shows that, the values of covariance term and entropy differences matched exactly, supporting the theoretical conclusion. Moreover, the covariance term stays mostly positive throughout training, further explaining why policy entropy would decrease monotonically. Through understanding the mechanism behind entropy dynamics, we motivate to control entropy by restricting the update of high-covariance tokens. Specifically, we propose two simple yet effective techniques, namely Clip-Cov and KL-Cov, which clip and apply KL penalty to tokens with high covariances respectively. Experiments show that these methods encourage exploration, thus helping policy escape entropy collapse and achieve better downstream performance.

Gemini: A Family ofHighly Capable…Gemini: A Family of Highly Capable Multimodal ModelsDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsLearning to Reason underOff-Policy GuidanceLearning to Reason under Off-Policy GuidanceTTRL: Test-TimeReinforcement LearningTTRL: Test-Time Reinforcement LearningRight Question isAlready Half the Answer…Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationHybridFlow: A Flexibleand Efficient RLHF…HybridFlow: A Flexible and Efficient RLHF FrameworkKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsACECODER: Acing Coder RLvia Automated Test-Case…ACECODER: Acing Coder RL via Automated Test-Case SynthesisMiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionStabilizing Knowledge,Promoting Reasoning…Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVRSimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningOn Entropy Control inLLM-RL AlgorithmsOn Entropy Control in LLM-RL AlgorithmsEvolving Language Modelswithout Labels: Majorit…Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes VariationSingle-stream PolicyOptimizationSingle-stream Policy OptimizationRandom Policy Valuationis Enough for LLM…Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsETTRL: BalancingExploration and…ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy MechanismFrom System 1 to System2: A Survey of Reasonin…From System 1 to System 2: A Survey of Reasoning Large Language ModelsRisk-Sensitive RL forAlleviating Exploration…Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language ModelsLow-probability TokensSustain Exploration in…Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable RewardHeterogeneous AdaptivePolicy Optimization…Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's NatureThe Entropy Mechanism ofReinforcement Learning…The Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.