Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

This paper explores the impact of extending input lengths on the capabilities of Large Language Models (LLMs). Despite LLMs advancements in recent times, their performance consistency across different input lengths is not well understood. We investigate this aspect by introducing a novel QA reasoning framework, specifically designed to assess the impact of input length. We isolate the effect of input length using multiple versions of the same sample, each being extended with padding of different lengths, types and locations. Our findings show a notable degradation in LLMs’ reasoning performance at much shorter input lengths than their technical maximum. We show that the degradation trend appears in every version of our dataset, although at different intensities.Additionally, our study reveals that the traditional metric of next word prediction correlates negatively with performance of LLMs’ on our reasoning dataset. We analyse our results and identify failure modes that can serve as useful guides for future research, potentially informing strategies to address the limitations observed in LLMs.

Training Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsLanguage models showhuman-like content…Language models show human-like content effects on reasoningTree of Thoughts:Deliberate Problem…Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsZeroSCROLLS: A Zero-ShotBenchmark for Long Text…ZeroSCROLLS: A Zero-Shot Benchmark for Long Text UnderstandingLarge Language ModelsAre Human-Level Prompt…Large Language Models Are Human-Level Prompt EngineersTowards Reasoning inLarge Language Models…Towards Reasoning in Large Language Models: A SurveyTraining Trajectories ofLanguage Models Across…Training Trajectories of Language Models Across ScalesScaling Laws vs ModelArchitectures: How does…Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?Lost in the Middle: HowLanguage Models Use Lon…Lost in the Middle: How Language Models Use Long ContextsThe Impact of ReasoningStep Length on Large…The Impact of Reasoning Step Length on Large Language ModelsGenerative Models as aComplex Systems Science…Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?One Thousand and OnePairs: A "novel"…One Thousand and One Pairs: A "novel" challenge for long-context language modelsRetrieval AugmentedGeneration or…Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid ApproachL-CiteEval: DoLong-Context Models…L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding?Stress-TestingLong-Context Language…Stress-Testing Long-Context Language Models with Lifelong ICL and Task HaystackLIFBench: Evaluating theInstruction Following…LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context ScenariosMathHay: An AutomatedBenchmark for…MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMsCLIPPER: Compressionenables long-context…CLIPPER: Compression enables long-context synthetic data generationGSM-Infinite: How DoYour LLMs Behave over…GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?Holistic Reasoning withLong-Context LMs: A…Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual DataNeedle Threading: CanLLMs Follow Threads…Needle Threading: Can LLMs Follow Threads Through Near-Million-Scale Haystacks?The Relationship BetweenReasoning and…The Relationship Between Reasoning and Performance in Large Language Models - o3 (mini) Thinks Harder, Not LongerDocETL: Agentic QueryRewriting and Evaluatio…DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingSame Task, More Tokens:the Impact of Input…Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。