Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.

Training Deep Nets withSublinear Memory CostTraining Deep Nets with Sublinear Memory CostGPT-4 Technical ReportGPT-4 Technical ReportWhat's Behind PPO'sCollapse in Long-CoT?…What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the SecretHybridFlow: A Flexibleand Efficient RLHF…HybridFlow: A Flexible and Efficient RLHF FrameworkA Survey ofReinforcement Learning…A Survey of Reinforcement Learning for Large Reasoning ModelsrStar2-Agent: AgenticReasoning Technical…rStar2-Agent: Agentic Reasoning Technical ReportMiMo: Unlocking theReasoning Potential of…MiMo: Unlocking the Reasoning Potential of Language Model - From Pretraining to PosttrainingMiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionAdaCoT: Pareto-OptimalAdaptive…AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement LearningPAG: Multi-TurnReinforced LLM…PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative VerifierTowards Concise andAdaptive Thinking in…Towards Concise and Adaptive Thinking in Large Reasoning Models: A SurveyInternBootcamp TechnicalReport: Boosting LLM…InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task ScalingReinforcement Learningfrom Human FeedbackReinforcement Learning from Human FeedbackEnigmata: ScalingLogical Reasoning in…Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable PuzzlesReinforcement LearningOptimization for…Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling LibrarySynLogic: SynthesizingVerifiable Reasoning…SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and BeyondSeed1.5-Thinking:Advancing Superb…Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement LearningEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.