Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), while its mechanisms are not yet well understood. In this work, we undertake a pioneering exploration of RLVR through the novel perspective of token entropy patterns, comprehensively analyzing how different tokens influence reasoning performance. By examining token entropy patterns in Chain-of-Thought (CoT) reasoning, we observe that only a small fraction of tokens exhibit high entropy, and these tokens act as critical forks that steer the model toward diverse reasoning pathways. Furthermore, studying how entropy patterns evolve during RLVR training reveals that RLVR largely adheres to the base model's entropy patterns, primarily adjusting the entropy of high-entropy tokens. These findings highlight the significance of high-entropy tokens (i.e., forking tokens) to RLVR. We ultimately improve RLVR by restricting policy gradient updates to forking tokens and uncover a finding even beyond the 80/20 rule: utilizing only 20% of the tokens while maintaining performance comparable to full-gradient updates on the Qwen3-8B base model and significantly surpassing full-gradient updates on the Qwen3-32B (+11.04 on AIME'25 and +7.71 on AIME'24) and Qwen3-14B (+4.79 on AIME'25 and +5.21 on AIME'24) base models, highlighting a strong scaling trend. In contrast, training exclusively on the 80% lowest-entropy tokens leads to a marked decline in performance. These findings indicate that the efficacy of RLVR primarily arises from optimizing the high-entropy tokens that decide reasoning directions. Collectively, our results highlight the potential to understand RLVR through a token-entropy perspective and optimize RLVR by leveraging high-entropy minority tokens to further improve LLM reasoning.

OpenRLHF: AnEasy-to-use, Scalable…OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF FrameworkOpen-Reasoner-Zero: AnOpen Source Approach to…Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildVAPO: Efficient andReliable Reinforcement…VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning TasksDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleReinforcement Learningfor Reasoning in Large…Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleDoes ReinforcementLearning Really…Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Cognitive Behaviors thatEnable Self-Improving…Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsAbsolute Zero:Reinforced Self-play…Absolute Zero: Reinforced Self-play Reasoning with Zero DataDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningSFT Memorizes, RLGeneralizes: A…SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingQwen3 Technical ReportQwen3 Technical ReportMiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionSpurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVRAbsolute Zero:Reinforced Self-play…Absolute Zero: Reinforced Self-play Reasoning with Zero DataFirst Return,Entropy-Eliciting…First Return, Entropy-Eliciting ExploreDCPO: Dynamic ClippingPolicy OptimizationDCPO: Dynamic Clipping Policy OptimizationAgentic Entropy-BalancedPolicy OptimizationAgentic Entropy-Balanced Policy OptimizationOn the Mechanism ofReasoning Pattern…On the Mechanism of Reasoning Pattern Selection in Reinforcement Learning for Language ModelsGLM-4.5: Agentic,Reasoning, and Coding…GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation ModelsRandom Policy Valuationis Enough for LLM…Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsETTRL: BalancingExploration and…ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy MechanismLow-probability TokensSustain Exploration in…Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable RewardHeterogeneous AdaptivePolicy Optimization…Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's NatureBeyond the 80/20 Rule:High-Entropy Minority…Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.