Training Verifiers to Solve Math Word Problems

State-of-the-art language models can match human performance on many tasks, but they still struggle to robustly perform multi-step mathematical reasoning. To diagnose the failures of current models and support research, we introduce GSM8K, a dataset of 8.5K high quality linguistically diverse grade school math word problems. We find that even the largest transformer models fail to achieve high test performance, despite the conceptual simplicity of this problem distribution. To increase performance, we propose training verifiers to judge the correctness of model completions. At test time, we generate many candidate solutions and select the one ranked highest by the verifier. We demonstrate that verification significantly improves performance on GSM8K, and we provide strong empirical evidence that verification scales more effectively with increased data than a finetuning baseline.

How well do ComputersSolve Math Word…How well do Computers Solve Math Word Problems? Large-Scale Dataset Construction and EvaluationDeep Neural Solver forMath Word ProblemsDeep Neural Solver for Math Word ProblemsProgram Induction byRationale Generation…Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word ProblemsNeural Math Word ProblemSolver with…Neural Math Word Problem Solver with Reinforcement LearningMathQA: TowardsInterpretable Math Word…MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based FormalismsCommonsenseQA: AQuestion Answering…CommonsenseQA: A Question Answering Challenge Targeting Commonsense KnowledgeA Goal-DrivenTree-Structured Neural…A Goal-Driven Tree-Structured Neural Model for Math Word ProblemsA Diverse Corpus forEvaluating and…A Diverse Corpus for Evaluating and Developing English Math Word Problem SolversLanguage Models areFew-Shot LearnersLanguage Models are Few-Shot LearnersScaling Laws for NeuralLanguage ModelsScaling Laws for Neural Language ModelsApe210K: A Large-Scaleand Template-Rich…Ape210K: A Large-Scale and Template-Rich Dataset of Math Word ProblemsMeasuring MathematicalProblem Solving With th…Measuring Mathematical Problem Solving With the MATH DatasetPlanBench: An ExtensibleBenchmark for Evaluatin…PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about ChangeHave LLMs AdvancedEnough? A Challenging…Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsTrainingChain-of-Thought via…Training Chain-of-Thought via Latent-Variable InferenceMathVista: EvaluatingMath Reasoning in Visua…MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal ModelsChain of ThoughtEmpowers Transformers t…Chain of Thought Empowers Transformers to Solve Inherently Serial ProblemsSALAD-Bench: AHierarchical and…SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language ModelsCan MLLMs Reason inMultimodality? EMMA: An…Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning BenchmarkDifferential TransformerDifferential TransformerEvaluating Judges asEvaluators: The JETTS…Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsdLLM-Cache: AcceleratingDiffusion Large Languag…dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingSafeChain: Safety ofLanguage Models with…SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning CapabilitiesThe SurprisingEffectiveness of…The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningTraining Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.