Differential Transformer

Transformer tends to overallocate attention to irrelevant context. In this work, we introduce Diff Transformer, which amplifies attention to the relevant context while canceling noise. Specifically, the differential attention mechanism calculates attention scores as the difference between two separate softmax attention maps. The subtraction cancels noise, promoting the emergence of sparse attention patterns. Experimental results on language modeling show that Diff Transformer outperforms Transformer in various settings of scaling up model size and training tokens. More intriguingly, it offers notable advantages in practical applications, such as long-context modeling, key information retrieval, hallucination mitigation, in-context learning, and reduction of activation outliers. By being less distracted by irrelevant context, Diff Transformer can mitigate hallucination in question answering and text summarization. For in-context learning, Diff Transformer not only enhances accuracy but is also more robust to order permutation, which was considered as a chronic robustness issue. The results position Diff Transformer as a highly effective and promising architecture to advance large language models.

GLU Variants ImproveTransformerGLU Variants Improve TransformerTraining Verifiers toSolve Math Word ProblemsTraining Verifiers to Solve Math Word ProblemsLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsGemini 1.5: Unlockingmultimodal understandin…Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextIn-Context Learning withLong-Context Models: An…In-Context Learning with Long-Context Models: An In-Depth ExplorationFlashAttention-3: Fastand Accurate Attention…FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionFlashAttention-2: FasterAttention with Better…FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningOpenAI o1 System CardOpenAI o1 System CardLongBench: A Bilingual,Multitask Benchmark for…LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingZoology: Measuring andImproving Recall in…Zoology: Measuring and Improving Recall in Efficient Language ModelsDeepSeek-R1:Incentivizing Reasoning…DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningWorld Model onMillion-Length Video An…World Model on Million-Length Video And Language With Blockwise RingAttentionTrustworthy andEfficient LLMs Meet…Trustworthy and Efficient LLMs Meet DatabasesLarge Language ModelsCan Self-Improve in…Large Language Models Can Self-Improve in Long-context ReasoningSelective Attention:Enhancing Transformer…Selective Attention: Enhancing Transformer through Principled Context ControlGated Attention forLarge Language Models…Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeHybrid Architectures forLanguage Models…Hybrid Architectures for Language Models: Systematic Analysis and Design InsightsDINT TransformerDINT TransformerUnderstandingDifferential Transforme…Understanding Differential Transformer Unchains Pretrained Self-AttentionsScaling Stick-BreakingAttention: An Efficient…Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth StudyScalable-Softmax IsSuperior for AttentionScalable-Softmax Is Superior for AttentionSTAR: Synthesis ofTailored ArchitecturesSTAR: Synthesis of Tailored ArchitecturesHuman-inspired EpisodicMemory for Infinite…Human-inspired Episodic Memory for Infinite Context LLMsTransolver++: AnAccurate Neural Solver…Transolver++: An Accurate Neural Solver for PDEs on Million-Scale GeometriesDifferential TransformerDifferential TransformerEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.