s1: Simple test-time scaling

Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts. We seek the simplest approach to achieve test-time scaling and strong reasoning performance. First, we curate a small dataset s1K of 1,000 questions paired with reasoning traces relying on three criteria we validate through ablations: difficulty, diversity, and quality. Second, we develop budget forcing to control test-time compute by forcefully terminating the model's thinking process or lengthening it by appending "Wait" multiple times to the model's generation when it tries to end. This can lead the model to double-check its answer, often fixing incorrect reasoning steps. After supervised finetuning the Qwen2.5-32B-Instruct language model on s1K and equipping it with budget forcing, our model s1-32B exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24). Further, scaling s1-32B with budget forcing allows extrapolating beyond its performance without test-time intervention: from 50% to 57% on AIME24. Our model, data, and code are open-source at https://github.com/simplescaling/s1

Measuring MathematicalProblem Solving With th…Measuring Mathematical Problem Solving With the MATH DatasetGPQA: A Graduate-LevelGoogle-Proof Q&A…GPQA: A Graduate-Level Google-Proof Q&A BenchmarkOlympiadBench: AChallenging Benchmark…OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsMath-Shepherd: Verifyand Reinforce LLMs…Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsLarge Language Monkeys:Scaling Inference…Large Language Monkeys: Scaling Inference Compute with Repeated SamplingLet's Verify Step byStepLet's Verify Step by StepLIMO: Less is More forReasoningLIMO: Less is More for ReasoningKimi k1.5: ScalingReinforcement Learning…Kimi k1.5: Scaling Reinforcement Learning with LLMsTowards System 2Reasoning in LLMs…Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-ThoughtDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningOmni-MATH: A UniversalOlympiad Level…Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language ModelsWizardMath: EmpoweringMathematical Reasoning…WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructSafety Tax: SafetyAlignment Makes Your…Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less ReasonableFrom System 1 to System2: A Survey of Reasonin…From System 1 to System 2: A Survey of Reasoning Large Language ModelsSafeChain: Safety ofLanguage Models with…SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning CapabilitiesOn the Emergence ofThinking in LLMs I…On the Emergence of Thinking in LLMs I: Searching for the Right IntuitionDoes Math ReasoningImprove General LLM…Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningBig-Math: A Large-Scale,High-Quality Math…Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language ModelsReinforcement Learningvs. Distillation…Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM ReasoningLearning to Reason underOff-Policy GuidanceLearning to Reason under Off-Policy GuidanceTwo Heads are BetterThan One: Test-time…Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative ReasoningBoT: Breaking LongThought Processes of…BoT: Breaking Long Thought Processes of o1-like Large Language Models through Backdoor AttackMulti-Step Reasoningwith Large Language…Multi-Step Reasoning with Large Language Models, a SurveyThe First Few Tokens AreAll You Need: An…The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Modelss1: Simple test-timescalings1: Simple test-time scaling過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。