DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

BioXP-0.5B is a 🤗 Medical-AI model trained using our two-stage fine-tuning approach: Supervised Fine-Tuning (SFT): The model was initially fine-tuned on labeled data(MedMCQA) to achieve strong baseline accuracy on multiple-choice medical QA tasks. Group Relative Policy Optimization (GRPO): In the second stage, GRPO was applied to further align the model with human-like reasoning patterns. This reinforcement learning technique enhances the model’s ability to generate coherent, high-quality explanations and improve answer reliability.

Training Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsMeasuring MathematicalProblem Solving With th…Measuring Mathematical Problem Solving With the MATH DatasetQwen Technical ReportQwen Technical ReportLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsGemini: A Family ofHighly Capable…Gemini: A Family of Highly Capable Multimodal ModelsMistral 7BMistral 7BDeepSeek LLM: ScalingOpen-Source Language…DeepSeek LLM: Scaling Open-Source Language Models with LongtermismLet's Verify Step byStepLet's Verify Step by StepMetaMath: Bootstrap YourOwn Mathematical…MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLlemma: An Open LanguageModel For MathematicsLlemma: An Open Language Model For MathematicsDiagnosingNon-Intermittent…Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)WizardMath: EmpoweringMathematical Reasoning…WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructTyphoon 2: A Family ofOpen Text and Multimoda…Typhoon 2: A Family of Open Text and Multimodal Thai Large Language ModelsSearch-R1: Training LLMsto Reason and Leverage…Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningEvaluating Judges asEvaluators: The JETTS…Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsCUDA-L1: Improving CUDAOptimization via…CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement LearningToward large reasoningmodels: A survey of…Toward large reasoning models: A survey of reinforced reasoning with large language modelsTrajectory Balance withAsynchrony: Decoupling…Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-TrainingRank-R1: EnhancingReasoning in LLM-based…Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement LearningContinuousSelf-Improvement of…Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample SelectionLook Back to ReasonForward: Revisitable…Look Back to Reason Forward: Revisitable Memory for Long-Context LLM AgentsHow to Train a Leader:Hierarchical Reasoning…How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMsReasoning over SemanticIDs Enhances Generative…Reasoning over Semantic IDs Enhances Generative RecommendationDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.