Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Scaling the amount of compute used to train language models has dramatically improved their capabilities. However, when it comes to inference, we often limit models to making only one attempt at a problem. Here, we explore inference compute as another axis for scaling, using the simple technique of repeatedly sampling candidate solutions from a model. Across multiple tasks and models, we observe that coverage -- the fraction of problems that are solved by any generated sample -- scales with the number of samples over four orders of magnitude. Interestingly, the relationship between coverage and the number of samples is often log-linear and can be modelled with an exponentiated power law, suggesting the existence of inference-time scaling laws. In domains like coding and formal proofs, where answers can be automatically verified, these increases in coverage directly translate into improved performance. When we apply repeated sampling to SWE-bench Lite, the fraction of issues solved with DeepSeek-Coder-V2-Instruct increases from 15.9% with one sample to 56% with 250 samples, outperforming the single-sample state-of-the-art of 43%. In domains without automatic verifiers, we find that common methods for picking from a sample collection (majority voting and reward models) plateau beyond several hundred samples and fail to fully scale with the sample budget.

Language Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersEvaluating LargeLanguage Models Trained…Evaluating Large Language Models Trained on CodeTraining Compute-OptimalLarge Language ModelsTraining Compute-Optimal Large Language ModelsLLM-Blender: EnsemblingLarge Language Models…LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionTree of Thoughts:Deliberate Problem…Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsGPT-4 Technical ReportGPT-4 Technical ReportPythia: A Suite forAnalyzing Large Languag…Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingSelf-Refine: IterativeRefinement with…Self-Refine: Iterative Refinement with Self-FeedbackMath-Shepherd: Verifyand Reinforce LLMs…Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsRouteLLM: Learning toRoute LLMs with…RouteLLM: Learning to Route LLMs with Preference DataAre More LLM Calls AllYou Need? Towards…Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference SystemsMixture-of-AgentsEnhances Large Language…Mixture-of-Agents Enhances Large Language Model CapabilitiesParallel Scaling Law forLanguage ModelsParallel Scaling Law for Language ModelsDiverse Inference andVerification for…Diverse Inference and Verification for Advanced ReasoningInference-AwareFine-Tuning for…Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language ModelsIs Best-of-N the Best ofThem? Coverage, Scaling…Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time AlignmentReward Reasoning ModelReward Reasoning ModelEvaluating Judges asEvaluators: The JETTS…Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsHarnessing the ReasoningEconomy: A Survey of…Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language ModelsTwo Heads are BetterThan One: Test-time…Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative ReasoningKernelBench: Can LLMsWrite Efficient GPU…KernelBench: Can LLMs Write Efficient GPU Kernels?Kinetics: RethinkingTest-Time Scaling LawsKinetics: Rethinking Test-Time Scaling LawsEfficiently Learning atTest-Time: Active…Efficiently Learning at Test-Time: Active Fine-Tuning of LLMsScalingInference-Efficient…Scaling Inference-Efficient Language ModelsLarge Language Monkeys:Scaling Inference…Large Language Monkeys: Scaling Inference Compute with Repeated Sampling過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。