MiMo-V2-Flash Technical Report

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-V2-Flash adopts a hybrid attention architecture that interleaves Sliding Window Attention (SWA) with global attention, with a 128-token sliding window under a 5:1 hybrid ratio. The model is pre-trained on 27 trillion tokens with Multi-Token Prediction (MTP), employing a native 32k context length and subsequently extended to 256k. To efficiently scale post-training compute, MiMo-V2-Flash introduces a novel Multi-Teacher On-Policy Distillation (MOPD) paradigm. In this framework, domain-specialized teachers (e.g., trained via large-scale reinforcement learning) provide dense and token-level reward, enabling the student model to perfectly master teacher expertise. MiMo-V2-Flash rivals top-tier open-weight models such as DeepSeek-V3.2 and Kimi-K2, despite using only 1/2 and 1/3 of their total parameters, respectively. During inference, by repurposing MTP as a draft model for speculative decoding, MiMo-V2-Flash achieves up to 3.6 acceptance length and 2.6x decoding speedup with three MTP layers. We open-source both the model weights and the three-layer MTP weights to foster open research and community collaboration.

DeepSeek-V3 TechnicalReportDeepSeek-V3 Technical ReportDeepSeekMath: Pushingthe Limits of…DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsGemma 2: Improving OpenLanguage Models at a…Gemma 2: Improving Open Language Models at a Practical SizeKimi K2: Open AgenticIntelligenceKimi K2: Open Agentic IntelligenceMiMo: Unlocking theReasoning Potential of…MiMo: Unlocking the Reasoning Potential of Language Model - From Pretraining to PosttrainingKimi Linear: AnExpressive, Efficient…Kimi Linear: An Expressive, Efficient Attention ArchitectureGated Attention forLarge Language Models…Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeMiniMax-01: ScalingFoundation Models with…MiniMax-01: Scaling Foundation Models with Lightning AttentionGemma 3 Technical ReportGemma 3 Technical ReportDeepSeek-V3.2: Pushingthe Frontier of Open…DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsWhen Attention SinkEmerges in Language…When Attention Sink Emerges in Language Models: An Empirical ViewAREAL: A Large-ScaleAsynchronous…AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningReinforcement Learningfrom Human FeedbackReinforcement Learning from Human FeedbackRethinking On-PolicyDistillation of Large…Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and RecipeHySparse: A HybridSparse Attention…HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache SharingSelf-Distilled RLVRSelf-Distilled RLVRLearning beyond Teacher:Generalized On-Policy…Learning beyond Teacher: Generalized On-Policy Distillation with Reward ExtrapolationRevisiting On-PolicyDistillation: Empirical…Revisiting On-Policy Distillation: Empirical Failure Modes and Simple FixesArcee Trinity LargeTechnical ReportArcee Trinity Large Technical ReportUni-OPD: UnifyingOn-Policy Distillation…Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective RecipeJoyAI-LLM Flash:Advancing Mid-Scale LLM…JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token EfficiencyLightning OPD: EfficientPost-Training for Large…Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy DistillationNemotron-Cascade 2:Post-Training LLMs with…Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy DistillationAttention Editing: AVersatile Framework for…Attention Editing: A Versatile Framework for Cross-Architecture Attention ConversionMiMo-V2-Flash TechnicalReportMiMo-V2-Flash Technical Report過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。