Skywork Open Reasoner 1 Technical Report

The success of DeepSeek-R1 underscores the significant role of reinforcement learning (RL) in enhancing the reasoning capabilities of large language models (LLMs). In this work, we present Skywork-OR1, an effective and scalable RL implementation for long Chain-of-Thought (CoT) models. Building on the DeepSeek-R1-Distill model series, our RL approach achieves notable performance gains, increasing average accuracy across AIME24, AIME25, and LiveCodeBench from 57.8% to 72.8% (+15.0%) for the 32B model and from 43.6% to 57.5% (+13.9%) for the 7B model. Our Skywork-OR1-32B model surpasses both DeepSeek-R1 and Qwen3-32B on the AIME24 and AIME25 benchmarks, while achieving comparable results on LiveCodeBench. The Skywork-OR1-7B and Skywork-OR1-Math-7B models demonstrate competitive reasoning capabilities among models of similar size. We perform comprehensive ablation studies on the core components of our training pipeline to validate their effectiveness. Additionally, we thoroughly investigate the phenomenon of entropy collapse, identify key factors affecting entropy dynamics, and demonstrate that mitigating premature entropy collapse is critical for improved test performance. To support community research, we fully open-source our model weights, training code, and training datasets.

TACO: Topics inAlgorithmic COde…TACO: Topics in Algorithmic COde generation datasetDiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)DeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsThe Llama 3 Herd ofModelsThe Llama 3 Herd of ModelsLight-R1: CurriculumSFT, DPO and RL for Lon…Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and BeyondVAPO: Efficient andReliable Reinforcement…VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning TasksDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleTTRL: Test-TimeReinforcement LearningTTRL: Test-Time Reinforcement LearningDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningLeetCodeDataset: ATemporal Dataset for…LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMsKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsQwen3 Technical ReportQwen3 Technical ReportStabilizing Knowledge,Promoting Reasoning…Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVRJointly ReinforcingDiversity and Quality i…Jointly Reinforcing Diversity and Quality in Language Model GenerationsSqueeze the SoakedSponge: Efficient…Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language ModelASPO: AsymmetricImportance Sampling…ASPO: Asymmetric Importance Sampling Policy OptimizationAgentGym-RL: TrainingLLM Agents for…AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement LearningRandom Policy Valuationis Enough for LLM…Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsThe Choice ofDivergence: A Neglected…The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable RewardThinking-Free PolicyInitialization Makes…Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient ReasonersRethinking EntropyInterventions in RLVR…Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveLook Inward to ExploreOutward: Learning…Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RLLow-probability TokensSustain Exploration in…Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable RewardFIPO: Eliciting DeepReasoning with Future-K…FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy OptimizationSkywork Open Reasoner 1Technical ReportSkywork Open Reasoner 1 Technical ReportEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.