CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Recent developments in large language models (LLMs) have been impressive. However, these models sometimes show inconsistencies and problematic behavior, such as hallucinating facts, generating flawed code, or creating offensive and toxic content. Unlike these models, humans typically utilize external tools to cross-check and refine their initial content, like using a search engine for fact-checking, or a code interpreter for debugging. Inspired by this observation, we introduce a framework called CRITIC that allows LLMs, which are essentially "black boxes" to validate and progressively amend their own outputs in a manner similar to human interaction with tools. More specifically, starting with an initial output, CRITIC interacts with appropriate tools to evaluate certain aspects of the text, and then revises the output based on the feedback obtained during this validation process. Comprehensive evaluations involving free-form question answering, mathematical program synthesis, and toxicity reduction demonstrate that CRITIC consistently enhances the performance of LLMs. Meanwhile, our research highlights the crucial importance of external feedback in promoting the ongoing self-improvement of LLMs.

WebGPT: Browser-assistedquestion-answering with…WebGPT: Browser-assisted question-answering with human feedbackTALM: Tool AugmentedLanguage ModelsTALM: Tool Augmented Language ModelsSelf-Refine: IterativeRefinement with…Self-Refine: Iterative Refinement with Self-FeedbackART: Automaticmulti-step reasoning an…ART: Automatic multi-step reasoning and tool-use for large language modelsReflexion: an autonomousagent with dynamic…Reflexion: an autonomous agent with dynamic memory and self-reflectionLanguage Models canSolve Computer TasksLanguage Models can Solve Computer TasksCheck Your Facts and TryAgain: Improving Large…Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated FeedbackProgram of ThoughtsPrompting: Disentanglin…Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning TasksLEVER: Learning toVerify Language-to-Code…LEVER: Learning to Verify Language-to-Code Generation with ExecutionToolformer: LanguageModels Can Teach…Toolformer: Language Models Can Teach Themselves to Use ToolsTeaching Large LanguageModels to Self-DebugTeaching Large Language Models to Self-DebugLarge Language ModelsCannot Self-Correct…Large Language Models Cannot Self-Correct Reasoning YetTPTU: Task Planning andTool Usage of Large…TPTU: Task Planning and Tool Usage of Large Language Model-based AI AgentsE&V: Prompting LargeLanguage Models to…E&V: Prompting Large Language Models to Perform Static Analysis by Pseudo-code Execution and VerificationExplore, Select, Derive,and Recall: Augmenting…Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task AutomationMINT: Evaluating LLMs inMulti-turn Interaction…MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackOctoPack: InstructionTuning Code Large…OctoPack: Instruction Tuning Code Large Language ModelsGaining Wisdom fromSetbacks: Aligning Larg…Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisMAmmoTH: Building MathGeneralist Models…MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningGenerative AI Agents forKnowledge Work…Generative AI Agents for Knowledge Work Augmentation in FinanceOn the Self-VerificationLimitations of Large…On the Self-Verification Limitations of Large Language Models on Reasoning and Planning TasksA Survey of ContextEngineering for Large…A Survey of Context Engineering for Large Language ModelsBuilding Math Agentswith Multi-Turn…Building Math Agents with Multi-Turn Iterative Preference LearningSailing AI by the Stars:A Survey of Learning…Sailing AI by the Stars: A Survey of Learning from Rewards in Post-Training and Test-Time Scaling of Large Language ModelsCRITIC: Large LanguageModels Can Self-Correct…CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.