Kimi k1.5: Scaling Reinforcement Learning with LLMs

Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learning (RL) unlocks a new axis for the continued improvement of artificial intelligence, with the promise that large language models (LLMs) can scale their training data by learning to explore with rewards. However, prior published work has not produced competitive results. In light of this, we report on the training practice of Kimi k1.5, our latest multi-modal LLM trained with RL, including its RL training techniques, multi-modal data recipes, and infrastructure optimization. Long context scaling and improved policy optimization methods are key ingredients of our approach, which establishes a simplistic, effective RL framework without relying on more complex techniques such as Monte Carlo tree search, value functions, and process reward models. Notably, our system achieves state-of-the-art reasoning performance across multiple benchmarks and modalities -- e.g., 77.5 on AIME, 96.2 on MATH 500, 94-th percentile on Codeforces, 74.9 on MathVista -- matching OpenAI's o1. Moreover, we present effective long2short methods that use long-CoT techniques to improve short-CoT models, yielding state-of-the-art short-CoT reasoning results -- e.g., 60.8 on AIME, 94.6 on MATH500, 47.3 on LiveCodeBench -- outperforming existing short-CoT models such as GPT-4o and Claude Sonnet 3.5 by a large margin (up to +550%).

Measuring MassiveMultitask Language…Measuring Massive Multitask Language UnderstandingGemini: A Family ofHighly Capable…Gemini: A Family of Highly Capable Multimodal ModelsStarCoder: may thesource be with you!StarCoder: may the source be with you!Scaling Data-ConstrainedLanguage ModelsScaling Data-Constrained Language ModelsOpenAI o1 System CardOpenAI o1 System CardScaling LLM Test-TimeCompute Optimally can b…Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model ParametersDo NOT Think That Muchfor 2+3=? On the…Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMsLet's Verify Step byStepLet's Verify Step by StepThe Llama 3 Herd ofModelsThe Llama 3 Herd of ModelsMathVista: EvaluatingMath Reasoning in Visua…MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal ModelsBack to Basics:Revisiting REINFORCE…Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsLiveCodeBench: Holisticand Contamination Free…LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleParallel Scaling Law forLanguage ModelsParallel Scaling Law for Language ModelsLearning to Reason underOff-Policy GuidanceLearning to Reason under Off-Policy GuidanceBig-Math: A Large-Scale,High-Quality Math…Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language ModelsMM-PRM: EnhancingMultimodal Mathematical…MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level SupervisionThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningSherlock:Self-Correcting…Sherlock: Self-Correcting Reasoning in Vision-Language ModelsSafeChain: Safety ofLanguage Models with…SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning CapabilitiesSimpleVLA-RL: ScalingVLA Training via…SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningDoes Math ReasoningImprove General LLM…Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningAgentGym-RL: TrainingLLM Agents for…AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement LearningReason-RFT:Reinforcement…Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language ModelsKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMs過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。