DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance in long-context scenarios. (2) Scalable Reinforcement Learning Framework: By implementing a robust reinforcement learning protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par with Gemini-3.0-Pro, achieving gold-medal performance in both the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI). (3) Large-Scale Agentic Task Synthesis Pipeline: To integrate reasoning into tool-use scenarios, we developed a novel synthesis pipeline that systematically generates training data at scale. This methodology facilitates scalable agentic post-training, yielding substantial improvements in generalization and instruction-following robustness within complex, interactive environments.

Fast TransformerDecoding: One Write-Hea…Fast Transformer Decoding: One Write-Head is All You NeedGPQA: A Graduate-LevelGoogle-Proof Q&A…GPQA: A Graduate-Level Google-Proof Q&A BenchmarkDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeek-V2: A Strong,Economical, and…DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelMMLU-Pro: A More Robustand Challenging…MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding BenchmarkGLM-4.5: Agentic,Reasoning, and Coding…GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation ModelsQwen3 Technical ReportQwen3 Technical ReportBrowseComp: A Simple YetChallenging Benchmark…BrowseComp: A Simple Yet Challenging Benchmark for Browsing AgentsGemini 2.5: Pushing theFrontier with Advanced…Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesHumanity's Last ExamHumanity's Last ExamSWE-smith: Scaling Datafor Software Engineerin…SWE-smith: Scaling Data for Software Engineering AgentsThe Tool Decathlon:Benchmarking Language…The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task ExecutionSWE-EVO: BenchmarkingCoding Agents in…SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution ScenariosOne RL to See Them All:Visual Triple Unified…One RL to See Them All: Visual Triple Unified Reinforcement LearningMiMo-V2-Flash TechnicalReportMiMo-V2-Flash Technical ReportDIVE: Scaling Diversityin Agentic Task…DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool UseCL-bench: A Benchmarkfor Context LearningCL-bench: A Benchmark for Context LearningReasoning Shift: HowContext Silently…Reasoning Shift: How Context Silently Shortens LLM ReasoningSkillCraft: Can LLMAgents Learn to Use…SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?LongCoT: BenchmarkingLong-Horizon…LongCoT: Benchmarking Long-Horizon Chain-of-Thought ReasoningDeepSeek-OCR 2: VisualCausal FlowDeepSeek-OCR 2: Visual Causal FlowRethinking the TrustRegion in LLM…Rethinking the Trust Region in LLM Reinforcement LearningLarge-Scale TerminalAgentic Trajectory…Large-Scale Terminal Agentic Trajectory Generation from Dockerized EnvironmentsSame Claim, DifferentJudgment: Benchmarking…Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation DetectionDeepSeek-V3.2: Pushingthe Frontier of Open…DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.