Efficient Reasoning with Hidden Thinking

Chain-of-Thought (CoT) reasoning has become a powerful framework for improving complex problem-solving capabilities in Multimodal Large Language Models (MLLMs). However, the verbose nature of textual reasoning introduces significant inefficiencies. In this work, we propose Heima (as hidden llama), an effective CoT compression framework that condenses lengthy CoTs into a small set of abstract thinking tokens, preserving essential reasoning while removing redundancy. We then conduct a theoretical analysis from an information-theoretic perspective, quantifying the information gap induced by compression, showing that reasoning capability is preserved when non-trivial mutual information is retained. To further explore and quantify this information gap, we design the adaptive interpreter that maps thinking tokens back to variable-length textual sequences, thereby reconstructing the reasoning process. Experiments across diverse reasoning benchmarks demonstrate that Heima improves reasoning efficiency, while maintaining or even achieving better zero-shot accuracy. Moreover, the interpreter reconstructs coherent reasoning progresses from compressed thinking tokens, revealing that the information gap is minimal and validating the effectiveness of the proposed framework. This work paves the way for scalable latent reasoning models and advances our understanding of efficient reasoning processes in large models. Code: https://github.com/shawnricecake/Heima

BERTScore: EvaluatingText Generation with…BERTScore: Evaluating Text Generation with BERTLoRA: Low-RankAdaptation of Large…LoRA: Low-Rank Adaptation of Large Language ModelsGPT-4 Technical ReportGPT-4 Technical ReportTraining Large LanguageModels to Reason in a…Training Large Language Models to Reason in a Continuous Latent SpaceCompressed Chain ofThought: Efficient…Compressed Chain of Thought: Efficient Reasoning Through Dense RepresentationsFrom Explicit CoT toImplicit CoT: Learning…From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by StepDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportMath-Shepherd: Verifyand Reinforce LLMs…Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsDo Large Language ModelsLatently Perform…Do Large Language Models Latently Perform Multi-Hop Reasoning?Dualformer: ControllableFast and Slow Thinking…Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning TracesLLaVA-CoT: Let VisionLanguage Models Reason…LLaVA-CoT: Let Vision Language Models Reason Step-by-StepSoftCoT: SoftChain-of-Thought for…SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMsImplicit Reasoning inLarge Language Models…Implicit Reasoning in Large Language Models: A Comprehensive SurveyHybrid Latent Reasoningvia Reinforcement…Hybrid Latent Reasoning via Reinforcement LearningStop Overthinking: ASurvey on Efficient…Stop Overthinking: A Survey on Efficient Reasoning for Large Language ModelsSoftCoT++: Test-TimeScaling with Soft…SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought ReasoningReasoning BeyondLanguage: A…Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought ReasoningEfficient Inference forLarge Reasoning Models…Efficient Inference for Large Reasoning Models: A SurveyEfficient ReasoningModels: A SurveyEfficient Reasoning Models: A SurveyVeriThinker: Learning toVerify Makes Reasoning…VeriThinker: Learning to Verify Makes Reasoning Model EfficientThink How to Think:Mitigating Overthinking…Think How to Think: Mitigating Overthinking with Autonomous Difficulty Cognition in Large Reasoning ModelsOThink-R1: IntrinsicFast/Slow Thinking Mode…OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning MitigationThink Before Recommend:Unleashing the Latent…Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential RecommendationEfficient Reasoning withHidden ThinkingEfficient Reasoning with Hidden Thinking過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。