RULER: What's the Real Context Size of Your Long-Context Language Models?

The needle-in-a-haystack (NIAH) test, which examines the ability to retrieve a piece of information (the "needle") from long distractor texts (the "haystack"), has been widely adopted to evaluate long-context language models (LMs). However, this simple retrieval-based test is indicative of only a superficial form of long-context understanding. To provide a more comprehensive evaluation of long-context LMs, we create a new synthetic benchmark RULER with flexible configurations for customized sequence length and task complexity. RULER expands upon the vanilla NIAH test to encompass variations with diverse types and quantities of needles. Moreover, RULER introduces new task categories multi-hop tracing and aggregation to test behaviors beyond searching from context. We evaluate 17 long-context LMs with 13 representative tasks in RULER. Despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases. While these models all claim context sizes of 32K tokens or greater, only half of them can maintain satisfactory performance at the length of 32K. Our analysis of Yi-34B, which supports context length of 200K, reveals large room for improvement as we increase input length and task complexity. We open source RULER to spur comprehensive evaluation of long-context LMs.

Mamba: Linear-TimeSequence Modeling with…Mamba: Linear-Time Sequence Modeling with Selective State SpacesLongBench: A Bilingual,Multitask Benchmark for…LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingBABILong: Testing theLimits of LLMs with Lon…BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-HaystackData Engineering forScaling Language Models…Data Engineering for Scaling Language Models to 128K ContextLM-Infinite: SimpleOn-the-Fly Length…LM-Infinite: Simple On-the-Fly Length Generalization for Large Language ModelsOne Thousand and OnePairs: A "novel"…One Thousand and One Pairs: A "novel" challenge for long-context language modelsLongRoPE: Extending LLMContext Window Beyond 2…LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens∞Bench: Extending LongContext Evaluation…∞Bench: Extending Long Context Evaluation Beyond 100K TokensFlashAttention-2: FasterAttention with Better…FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningLooGLE: Can Long-ContextLanguage Models…LooGLE: Can Long-Context Language Models Understand Long Contexts?LV-Eval: A BalancedLong-Context Benchmark…LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256KSoaring from 4K to 400K:Extending LLM's Context…Soaring from 4K to 400K: Extending LLM's Context with Activation BeaconLV-Eval: A BalancedLong-Context Benchmark…LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256KSeerAttention: LearningIntrinsic Sparse…SeerAttention: Learning Intrinsic Sparse Attention in Your LLMsLarge Language ModelsCan Self-Improve in…Large Language Models Can Self-Improve in Long-context ReasoningRetrieval AugmentedGeneration or…Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid ApproachLongReason: A SyntheticLong-Context Reasoning…LongReason: A Synthetic Long-Context Reasoning Benchmark via Context ExpansionReasoning on MultipleNeedles In A HaystackReasoning on Multiple Needles In A HaystackWildLong: SynthesizingRealistic Long-Context…WildLong: Synthesizing Realistic Long-Context Instruction Data at ScaleDoes RAG Really PerformBad For Long-Context…Does RAG Really Perform Bad For Long-Context Processing?D2O: DynamicDiscriminative…D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language ModelsSqueezed Attention:Accelerating Long…Squeezed Attention: Accelerating Long Context Length LLM InferenceEvaluating Memory in LLMAgents via Incremental…Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsNemotron-H: A Family ofAccurate and Efficient…Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer ModelsRULER: What's the RealContext Size of Your…RULER: What's the Real Context Size of Your Long-Context Language Models?過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。