Self-Consistency Improves Chain of Thought Reasoning in Language Models

Chain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks. In this paper, we propose a new decoding strategy, self-consistency, to replace the naive greedy decoding used in chain-of-thought prompting. It first samples a diverse set of reasoning paths instead of only taking the greedy one, and then selects the most consistent answer by marginalizing out the sampled reasoning paths. Self-consistency leverages the intuition that a complex reasoning problem typically admits multiple different ways of thinking leading to its unique correct answer. Our extensive empirical evaluation shows that self-consistency boosts the performance of chain-of-thought prompting with a striking margin on a range of popular arithmetic and commonsense reasoning benchmarks, including GSM8K (+17.9%), SVAMP (+11.0%), AQuA (+12.2%), StrategyQA (+6.4%) and ARC-challenge (+3.9%).

Learning to SolveArithmetic Word Problem…Learning to Solve Arithmetic Word Problems with Verb CategorizationMAWPS: A Math WordProblem RepositoryMAWPS: A Math Word Problem RepositoryProgram Induction byRationale Generation…Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word ProblemsHotpotQA: A Dataset forDiverse, Explainable…HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question AnsweringHierarchical NeuralStory GenerationHierarchical Neural Story GenerationNumNet: Machine ReadingComprehension with…NumNet: Machine Reading Comprehension with Numerical ReasoningEvaluating LargeLanguage Models Trained…Evaluating Large Language Models Trained on CodeScaling Language Models:Methods, Analysis &…Scaling Language Models: Methods, Analysis & Insights from Training GopherMaking Pre-trainedLanguage Models Better…Making Pre-trained Language Models Better Few-shot LearnersLaMDA: Language Modelsfor Dialog ApplicationsLaMDA: Language Models for Dialog ApplicationsPaLM: Scaling LanguageModeling with PathwaysPaLM: Scaling Language Modeling with PathwaysUL2: Unifying LanguageLearning ParadigmsUL2: Unifying Language Learning ParadigmsDeductive Verificationof Chain-of-Thought…Deductive Verification of Chain-of-Thought ReasoningExchange-of-Thought:Enhancing Large Languag…Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model CommunicationPlan, Verify and Switch:Integrated Reasoning…Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsDetGPT: Detect What YouNeed via ReasoningDetGPT: Detect What You Need via ReasoningSolving Math WordProblems via Cooperativ…Solving Math Word Problems via Cooperative Reasoning induced Language ModelsTrainingChain-of-Thought via…Training Chain-of-Thought via Latent-Variable InferenceMathVista: EvaluatingMath Reasoning in Visua…MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal ModelsSelf-contradictoryHallucinations of Large…Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and MitigationCoarse-to-FineHighlighting: Reducing…Coarse-to-Fine Highlighting: Reducing Knowledge Hallucination in Large Language ModelsMixture-of-AgentsEnhances Large Language…Mixture-of-Agents Enhances Large Language Model CapabilitiesEvaluating Judges asEvaluators: The JETTS…Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling EvaluatorsSurvey and analysis ofhallucinations in large…Survey and analysis of hallucinations in large language models: attribution to prompting strategies or model behaviorSelf-ConsistencyImproves Chain of…Self-Consistency Improves Chain of Thought Reasoning in Language ModelsEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.