HGRN2: Gated Linear RNNs with State Expansion

Hierarchically gated linear RNN (HGRN, \citealt{HGRN}) has demonstrated competitive training speed and performance in language modeling while offering efficient inference. However, the recurrent state size of HGRN remains relatively small, limiting its expressiveness. To address this issue, we introduce a simple outer product-based state expansion mechanism, which significantly enlarges the recurrent state size without introducing any additional parameters. This enhancement also provides a linear attention interpretation for HGRN2, enabling hardware-efficient training. Our extensive experiments verify the advantage of HGRN2 over HGRN consistently across different settings and competitive with other recurrent models.

Empirical Evaluation ofGated Recurrent Neural…Empirical Evaluation of Gated Recurrent Neural Networks on Sequence ModelingThe Devil in LinearTransformerThe Devil in Linear TransformerFine-Tuning Pre-trainedTransformers into…Fine-Tuning Pre-trained Transformers into Decaying Fast WeightsScaling TransNormer to175 Billion ParametersScaling TransNormer to 175 Billion ParametersRetentive Network: ASuccessor to Transforme…Retentive Network: A Successor to Transformer for Large Language ModelsLLaMA: Open andEfficient Foundation…LLaMA: Open and Efficient Foundation Language ModelsRWKV: Reinventing RNNsfor the Transformer EraRWKV: Reinventing RNNs for the Transformer EraParallelizing LinearTransformers with the…Parallelizing Linear Transformers with the Delta Rule over Sequence LengthGated Linear AttentionTransformers with…Gated Linear Attention Transformers with Hardware-Efficient TrainingSimple linear attentionlanguage models balance…Simple linear attention language models balance the recall-throughput tradeoffZoology: Measuring andImproving Recall in…Zoology: Measuring and Improving Recall in Efficient Language ModelsScaling Laws for LinearComplexity Language…Scaling Laws for Linear Complexity Language ModelsGated Linear AttentionTransformers with…Gated Linear Attention Transformers with Hardware-Efficient TrainingLinear AttentionSequence ParallelismLinear Attention Sequence ParallelismTransformers are SSMs:Generalized Models and…Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityMoM: Linear SequenceModeling with…MoM: Linear Sequence Modeling with Mixture-of-MemoriesLonghorn: State SpaceModels are Amortized…Longhorn: State Space Models are Amortized Online LearnersComba: ImprovingBilinear RNNs with…Comba: Improving Bilinear RNNs with Closed-loop ControlLiger: Linearizing LargeLanguage Models to Gate…Liger: Linearizing Large Language Models to Gated Recurrent StructuresGated Delta Networks:Improving Mamba2 with…Gated Delta Networks: Improving Mamba2 with Delta RuleLattice: Learning toEfficiently Compress th…Lattice: Learning to Efficiently Compress the MemoryLinear-MoE: LinearSequence Modeling Meets…Linear-MoE: Linear Sequence Modeling Meets Mixture-of-ExpertsKimi Linear: AnExpressive, Efficient…Kimi Linear: An Expressive, Efficient Attention ArchitectureLASP-2: RethinkingSequence Parallelism fo…LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its HybridHGRN2: Gated Linear RNNswith State ExpansionHGRN2: Gated Linear RNNs with State Expansion過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。