LLMs Get Lost In Multi-Turn Conversation

Large Language Models (LLMs) are conversational interfaces. As such, LLMs have the potential to assist their users not only when they can fully specify the task at hand, but also to help them define, explore, and refine what they need through multi-turn conversational exchange. Although analysis of LLM conversation logs has confirmed that underspecification occurs frequently in user instructions, LLM evaluation has predominantly focused on the single-turn, fully-specified instruction setting. In this work, we perform large-scale simulation experiments to compare LLM performance in single- and multi-turn settings. Our experiments confirm that all the top open- and closed-weight LLMs we test exhibit significantly lower performance in multi-turn conversations than single-turn, with an average drop of 39% across six generation tasks. Analysis of 200,000+ simulated conversations decomposes the performance degradation into two components: a minor loss in aptitude and a significant increase in unreliability. We find that LLMs often make assumptions in early turns and prematurely attempt to generate final solutions, on which they overly rely. In simpler terms, we discover that *when LLMs take a wrong turn in a conversation, they get lost and do not recover*.

BART: DenoisingSequence-to-Sequence…BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionAre You Sure?Challenging LLMs Leads…Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop ExperimentGemini: A Family ofHighly Capable…Gemini: A Family of Highly Capable Multimodal ModelsThe Llama 3 Herd ofModelsThe Llama 3 Herd of Models2 OLMo 2 Furious2 OLMo 2 FuriousSummary of a Haystack: AChallenge to…Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhi-4 Technical ReportPhi-4 Technical ReportDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningMultiChallenge: ARealistic Multi-Turn…MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMsHumanity's Last ExamHumanity's Last ExamCommand A: AnEnterprise-Ready Large…Command A: An Enterprise-Ready Large Language ModelLiveCodeBench: Holisticand Contamination Free…LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeDoctorAgent-RL: AMulti-Agent…DoctorAgent-RL: A Multi-Agent Collaborative Reinforcement Learning System for Multi-Turn Clinical DialogueDecoding Human-LLMCollaboration in Coding…Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the WildUser Feedback inHuman-LLM Dialogues: A…User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning SignalEchoing: IdentityFailures when LLM Agent…Echoing: Identity Failures when LLM Agents Talk to Each OtherTRAIL: Trace Reasoningand Agentic Issue…TRAIL: Trace Reasoning and Agentic Issue LocalizationBED-LLM: IntelligentInformation Gathering…BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental DesignSeDT:Sentence-Transformer…SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation ReliabilityWhat Prompts Don't Say:Understanding and…What Prompts Don't Say: Understanding and Managing Underspecification in LLM PromptsMAIGO: MitigatingLost-in-Conversation…MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-DistillationIntent Mismatch CausesLLMs to Get Lost in…Intent Mismatch Causes LLMs to Get Lost in Multi-Turn ConversationEvaluating TemporalConsistency in…Evaluating Temporal Consistency in Multi-Turn Language ModelsConsistency of LargeReasoning Models Under…Consistency of Large Reasoning Models Under Multi-Turn AttacksLLMs Get Lost InMulti-Turn ConversationLLMs Get Lost In Multi-Turn ConversationEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.