Reasoning with Exploration: An Entropy Perspective

Balancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing large language model (LLM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LLMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LLMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting longer and deeper reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LLM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LLM reasoning.

Beyond the 80/20 Rule:High-Entropy Minority…Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningSEED-GRPO: SemanticEntropy Enhanced GRPO…SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy OptimizationThe Entropy Mechanism ofReinforcement Learning…The Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsSpurious Rewards:Rethinking Training…Spurious Rewards: Rethinking Training Signals in RLVRRight Question isAlready Half the Answer…Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationSkywork Open Reasoner 1Technical ReportSkywork Open Reasoner 1 Technical ReportThe UnreasonableEffectiveness of Entrop…The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningReinforcement Learningfor Reasoning in Large…Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleTTRL: Test-TimeReinforcement LearningTTRL: Test-Time Reinforcement LearningSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveOutcome-basedExploration for LLM…Outcome-based Exploration for LLM ReasoningRandom Policy Valuationis Enough for LLM…Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsStabilizing Knowledge,Promoting Reasoning…Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVREDGE-GRPO:Entropy-Driven GRPO wit…EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage DiversityNudging the Boundariesof LLM ReasoningNudging the Boundaries of LLM ReasoningLearning WhatReinforcement Learning…Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest QuestionsExploration vsExploitation: Rethinkin…Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardHarnessing Uncertainty:Entropy-Modulated Polic…Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM AgentsUnlocking Exploration inRLVR: Uncertainty-aware…Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper ReasoningHeterogeneous AdaptivePolicy Optimization…Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's NatureLook Inward to ExploreOutward: Learning…Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RLF-GRPO: Don't Let YourPolicy Learn the Obviou…F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the RareReasoning withExploration: An Entropy…Reasoning with Exploration: An Entropy Perspective過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。