DeepSeek-V3 Technical Report

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models. Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks. The model checkpoints are available at https://github.com/deepseek-ai/DeepSeek-V3.

Measuring MassiveMultitask Language…Measuring Massive Multitask Language UnderstandingMistral 7BMistral 7BLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsQwen Technical ReportQwen Technical ReportLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsC-Eval: A Multi-LevelMulti-Discipline Chines…C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation ModelsInstruction-FollowingEvaluation for Large…Instruction-Following Evaluation for Large Language ModelsDeepSeek LLM: ScalingOpen-Source Language…DeepSeek LLM: Scaling Open-Source Language Models with LongtermismDeepSeekMoE: TowardsUltimate Expert…DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsMMLU-Pro: A More Robustand Challenging…MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding BenchmarkCMMLU: Measuring massivemultitask language…CMMLU: Measuring massive multitask language understanding in ChineseParallel Scaling Law forLanguage ModelsParallel Scaling Law for Language ModelsUnderstandingDifferential Transforme…Understanding Differential Transformer Unchains Pretrained Self-AttentionsRecipes for Pre-trainingLLMs with MXFP8Recipes for Pre-training LLMs with MXFP8IFEvalCode: ControlledCode GenerationIFEvalCode: Controlled Code GenerationLLaVA-PruMerge: AdaptiveToken Reduction for…LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsSearch Arena: AnalyzingSearch-Augmented LLMsSearch Arena: Analyzing Search-Augmented LLMsA Survey on (M)LLM-BasedGUI AgentsA Survey on (M)LLM-Based GUI AgentsToward large reasoningmodels: A survey of…Toward large reasoning models: A survey of reinforced reasoning with large language modelsChartPoint: GuidingMLLMs with Grounding…ChartPoint: Guiding MLLMs with Grounding Reflection for Chart ReasoningNear-Policy:Accelerating On-Policy…Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective PackingUniPool: A GloballyShared Expert Pool for…UniPool: A Globally Shared Expert Pool for Mixture-of-ExpertsConsRoute:Consistency-AwareAdaptive Query Routing…ConsRoute:Consistency-Aware Adaptive Query Routing for Cloud-Edge-Device Large Language ModelsDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.