YaRN: Efficient Context Window Extension of Large Language Models

Rotary Position Embeddings (RoPE) have been shown to effectively encode positional information in transformer-based language models. However, these models fail to generalize past the sequence length they were trained on. We present YaRN (Yet another RoPE extensioN method), a compute-efficient method to extend the context window of such models, requiring 10x less tokens and 2.5x less training steps than previous methods. Using YaRN, we show that LLaMA models can effectively utilize and extrapolate to context lengths much longer than their original pre-training would allow, while also surpassing previous the state-of-the-art at context window extension. In addition, we demonstrate that YaRN exhibits the capability to extrapolate beyond the limited context of a fine-tuning dataset. Code is available at https://github.com/jquesnelle/yarn

Think you have SolvedQuestion Answering? Try…Think you have Solved Question Answering? Try ARC, the AI2 Reasoning ChallengeRoFormer: EnhancedTransformer with Rotary…RoFormer: Enhanced Transformer with Rotary Position EmbeddingGPT-NeoX-20B: AnOpen-Source…GPT-NeoX-20B: An Open-Source Autoregressive Language ModelExtending Context Windowof Large Language Model…Extending Context Window of Large Language Models via Positional InterpolationLandmark Attention:Random-Access Infinite…Landmark Attention: Random-Access Infinite Context Length for TransformersA Length-ExtrapolatableTransformerA Length-Extrapolatable TransformerLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsPaLM: Scaling LanguageModeling with PathwaysPaLM: Scaling Language Modeling with PathwaysThe Impact of PositionalEncoding on Length…The Impact of Positional Encoding on Length Generalization in TransformersCode Llama: OpenFoundation Models for…Code Llama: Open Foundation Models for CodeLM-Infinite: SimpleOn-the-Fly Length…LM-Infinite: Simple On-the-Fly Length Generalization for Large Language ModelsFlashAttention-2: FasterAttention with Better…FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningCLEX: Continuous LengthExtrapolation for Large…CLEX: Continuous Length Extrapolation for Large Language ModelsLongEmbed: ExtendingEmbedding Models for…LongEmbed: Extending Embedding Models for Long Context RetrievalLongRAG: EnhancingRetrieval-Augmented…LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMsQwen2.5-Coder TechnicalReportQwen2.5-Coder Technical ReportMiniCPM: Unveiling thePotential of Small…MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training StrategiesA Comprehensive Surveyon Long Context Languag…A Comprehensive Survey on Long Context Language ModelingSoaring from 4K to 400K:Extending LLM's Context…Soaring from 4K to 400K: Extending LLM's Context with Activation BeaconLongRoPE2: Near-LosslessLLM Context Window…LongRoPE2: Near-Lossless LLM Context Window ScalingKimi Linear: AnExpressive, Efficient…Kimi Linear: An Expressive, Efficient Attention ArchitectureQuest: Query-centricData Synthesis Approach…Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language ModelLongWriter: Unleashing10, 000+ Word Generatio…LongWriter: Unleashing 10, 000+ Word Generation from Long Context LLMsLongLLaDA: UnlockingLong Context…LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMsYaRN: Efficient ContextWindow Extension of…YaRN: Efficient Context Window Extension of Large Language Models過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。