Kimi K2: Open Agentic Intelligence

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.

C-Eval: A Multi-LevelMulti-Discipline Chines…C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation ModelsDeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportDeepSeek-V2: A Strong,Economical, and…DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelMMLU-Pro: A More Robustand Challenging…MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding BenchmarkMiniCPM: Unveiling thePotential of Small…MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training StrategiesAgentInstruct: TowardGenerative Teaching wit…AgentInstruct: Toward Generative Teaching with Agentic Flowsτ-bench: A Benchmark forTool-Agent-User…τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsMuon is Scalable for LLMTrainingMuon is Scalable for LLM TrainingDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningHumanity's Last ExamHumanity's Last ExamFrom Crowdsourced Datato High-Quality…From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder PipelineSuperGPQA: Scaling LLMEvaluation across 285…SuperGPQA: Scaling LLM Evaluation across 285 Graduate DisciplinesGLM-4.5: Agentic,Reasoning, and Coding…GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation ModelsAutoForge: AutomatedEnvironment Synthesis…AutoForge: Automated Environment Synthesis for Agentic Reinforcement LearningTongyi DeepResearchTechnical ReportTongyi DeepResearch Technical ReportToolForge: A DataSynthesis Pipeline for…ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIsScaling Laws Meet ModelArchitecture: Toward…Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMsHermes 4 TechnicalReportHermes 4 Technical ReportToolMind TechnicalReport: A Large-Scale…ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use DatasetDiversity or Precision?A Deep Dive into Next…Diversity or Precision? A Deep Dive into Next Token PredictionMiMo-V2-Flash TechnicalReportMiMo-V2-Flash Technical ReportLLMRouterBench: AMassive Benchmark and…LLMRouterBench: A Massive Benchmark and Unified Framework for LLM RoutingA.X K1 Technical ReportA.X K1 Technical ReportCL-bench: A Benchmarkfor Context LearningCL-bench: A Benchmark for Context LearningKimi K2: Open AgenticIntelligenceKimi K2: Open Agentic Intelligence過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。