Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Large language models (LLMs) are increasingly optimized for long reasoning, under the assumption that more reasoning leads to better performance. However, emerging evidence suggests that longer responses can sometimes degrade accuracy rather than improve it. In this paper, we conduct a systematic empirical study of the relationship between reasoning length and answer correctness. We find that LLMs tend to overthink simple problems, generating unnecessarily long outputs, and underthink harder ones, failing to extend their reasoning when it is most needed. This indicates that models might misjudge problem difficulty and fail to calibrate their response length appropriately. Furthermore, we investigate the effects of length reduction with a preference optimization algorithm when simply preferring the shorter responses regardless of answer correctness. Experiments show that the generation length can be significantly reduced while maintaining acceptable accuracy. Our findings highlight generation length as a meaningful signal for reasoning behavior and motivate further exploration into LLMs' self-awareness in reasoning length adaptation.

Do NOT Think That Muchfor 2+3=? On the…Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMsThe Impact of ReasoningStep Length on Large…The Impact of Reasoning Step Length on Large Language ModelsWhen More is Less:Understanding…When More is Less: Understanding Chain-of-Thought Length in LLMsDAST:Difficulty-Adaptive…DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning ModelsCoT-Valve:Length-Compressible…CoT-Valve: Length-Compressible Chain-of-Thought TuningThe Relationship BetweenReasoning and…The Relationship Between Reasoning and Performance in Large Language Models - o3 (mini) Thinks Harder, Not LongerSelf-Training ElicitsConcise Reasoning in…Self-Training Elicits Concise Reasoning in Large Language ModelsReasoning Models Can BeEffective Without…Reasoning Models Can Be Effective Without ThinkingToken-Budget-Aware LLMReasoningToken-Budget-Aware LLM ReasoningL1: Controlling How LongA Reasoning Model Think…L1: Controlling How Long A Reasoning Model Thinks With Reinforcement LearningTowards Thinking-OptimalScaling of Test-Time…Towards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningO1-Pruner:Length-Harmonizing…O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning PruningThink or Not? ExploringThinking Efficiency in…Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic LensShorterBetter: GuidingReasoning Models to Fin…ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient ReasoningThinking Fast and Right:Balancing Accuracy and…Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive RewardsInverse Scaling inTest-Time ComputeInverse Scaling in Test-Time ComputeOverclocking LLMReasoning: Monitoring…Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMsSelf-Aligned Reward:Towards Effective and…Self-Aligned Reward: Towards Effective and Efficient ReasonersTowards Concise andAdaptive Thinking in…Towards Concise and Adaptive Thinking in Large Reasoning Models: A SurveyWhen Reasoning Meets ItsLawsWhen Reasoning Meets Its LawsInternal Bias inReasoning Models leads…Internal Bias in Reasoning Models leads to OverthinkingOptimalThinkingBench:Evaluating Over and…OptimalThinkingBench: Evaluating Over and Underthinking in LLMsMitigating Overthinkingin Large Reasoning…Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation MonitoringEfficient Reasoning withBalanced ThinkingEfficient Reasoning with Balanced ThinkingBetween Underthinkingand Overthinking: An…Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。