Kimi Linear: An Expressive, Efficient Attention Architecture

We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

HGRN2: Gated Linear RNNswith State ExpansionHGRN2: Gated Linear RNNs with State ExpansionGated Slot Attention forEfficient Linear-Time…Gated Slot Attention for Efficient Linear-Time Sequence ModelingComba: ImprovingBilinear RNNs with…Comba: Improving Bilinear RNNs with Closed-loop ControlJet-Nemotron: EfficientLanguage Model with Pos…Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchMoM: Linear SequenceModeling with…MoM: Linear Sequence Modeling with Mixture-of-MemoriesRWKV-7 "Goose" withExpressive Dynamic Stat…RWKV-7 "Goose" with Expressive Dynamic State EvolutionEfficient AttentionMechanisms for Large…Efficient Attention Mechanisms for Large Language Models: A SurveyMoBA: Mixture of BlockAttention for…MoBA: Mixture of Block Attention for Long-Context LLMsLoLCATs: On Low-RankLinearizing of Large…LoLCATs: On Low-Rank Linearizing of Large Language ModelsGated Attention forLarge Language Models…Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeForgetting Transformer:Softmax Attention with…Forgetting Transformer: Softmax Attention with a Forget GateHymba: A Hybrid-headArchitecture for Small…Hymba: A Hybrid-head Architecture for Small Language ModelsSpikingBrain TechnicalReport: Spiking…SpikingBrain Technical Report: Spiking Brain-inspired Large ModelsEnd-to-End Test-TimeTraining for Long…End-to-End Test-Time Training for Long ContextGated KalmaNet: A FadingMemory Layer Through…Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge RegressionMiMo-V2-Flash TechnicalReportMiMo-V2-Flash Technical ReportHySparse: A HybridSparse Attention…HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache SharingMamba-3: ImprovedSequence Modeling using…Mamba-3: Improved Sequence Modeling using State Space PrinciplesAttention ResidualsAttention ResidualsArcee Trinity LargeTechnical ReportArcee Trinity Large Technical ReportSwitch Attention:Towards Dynamic and…Switch Attention: Towards Dynamic and Fine-grained Hybrid TransformersMDN: ParallelizingStepwise Momentum for…MDN: Parallelizing Stepwise Momentum for Delta Linear AttentionMulti-Mixer Models:Flexible Sequence…Multi-Mixer Models: Flexible Sequence Modeling with Shared RepresentationsSpikingBrain2.0:Brain-Inspired…SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform InferenceKimi Linear: AnExpressive, Efficient…Kimi Linear: An Expressive, Efficient Attention ArchitectureEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.