Deep Think with Confidence

Large Language Models (LLMs) have shown great potential in reasoning tasks through test-time scaling methods like self-consistency with majority voting. However, this approach often leads to diminishing returns in accuracy and high computational overhead. To address these challenges, we introduce Deep Think with Confidence (DeepConf), a simple yet powerful method that enhances both reasoning efficiency and performance at test time. DeepConf leverages model-internal confidence signals to dynamically filter out low-quality reasoning traces during or after generation. It requires no additional model training or hyperparameter tuning and can be seamlessly integrated into existing serving frameworks. We evaluate DeepConf across a variety of reasoning tasks and the latest open-source models, including Qwen 3 and GPT-OSS series. Notably, on challenging benchmarks such as AIME 2025, DeepConf@512 achieves up to 99.9% accuracy and reduces generated tokens by up to 84.7% compared to full parallel thinking.

Let's Sample Step byStep…Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsEfficiently Serving LLMReasoning Programs with…Efficiently Serving LLM Reasoning Programs with CertaindexAdaptive Inference-TimeCompute: LLMs Can…Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-GenerationDo NOT Think That Muchfor 2+3=? On the…Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMsScalable Best-of-NSelection for Large…Scalable Best-of-N Selection for Large Language Models via Self-CertaintyDynamic Early Exit inReasoning ModelsDynamic Early Exit in Reasoning ModelsKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsLIMO: Less is More forReasoningLIMO: Less is More for Reasonings1: Simple test-timescalings1: Simple test-time scalingLearning to Reasonwithout External RewardsLearning to Reason without External RewardsDon't Overthink it.Preferring Shorter…Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM ReasoningThinkPrune: Pruning LongChain-of-Thought of LLM…ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement LearningThe Majority is notalways right: RL…The Majority is not always right: RL training for solution aggregationCoThink: Token-EfficientReasoning via Instruct…CoThink: Token-Efficient Reasoning via Instruct Models Guiding Reasoning ModelsThinking-Free PolicyInitialization Makes…Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient ReasonersZero-OverheadIntrospection for…Zero-Overhead Introspection for Adaptive Test-Time ComputeLost at the Beginning ofReasoningLost at the Beginning of ReasoningOn the Self-awareness ofLarge Reasoning Models'…On the Self-awareness of Large Reasoning Models' Capability BoundariesMatryoshkaThinking:Recursive Test-Time…MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient ReasoningDeepPrune: ParallelScaling without…DeepPrune: Parallel Scaling without Inter-trace RedundancyCoRefine:Confidence-Guided…CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time ComputeOne-Token Verificationfor Reasoning…One-Token Verification for Reasoning Correctness EstimationThink Deep, Not JustLong: Measuring LLM…Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking TokensSeer Self-Consistency:Advance Budget…Seer Self-Consistency: Advance Budget Estimation for Adaptive Test-Time ScalingDeep Think withConfidenceDeep Think with Confidence過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。