MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combined with a lightning attention mechanism. The model is developed based on our previous MiniMax-Text-01 model, which contains a total of 456 billion parameters with 45.9 billion parameters activated per token. The M1 model natively supports a context length of 1 million tokens, 8x the context size of DeepSeek R1. Furthermore, the lightning attention mechanism in MiniMax-M1 enables efficient scaling of test-time compute. These properties make M1 particularly suitable for complex tasks that require processing long inputs and thinking extensively. MiniMax-M1 is trained using large-scale reinforcement learning (RL) on diverse problems including sandbox-based, real-world software engineering environments. In addition to M1's inherent efficiency advantage for RL training, we propose CISPO, a novel RL algorithm to further enhance RL efficiency. CISPO clips importance sampling weights rather than token updates, outperforming other competitive RL variants. Combining hybrid-attention and CISPO enables MiniMax-M1's full RL training on 512 H800 GPUs to complete in only three weeks, with a rental cost of just $534,700. We release two versions of MiniMax-M1 models with 40K and 80K thinking budgets respectively, where the 40K model represents an intermediate phase of the 80K training. Experiments on standard benchmarks show that our models are comparable or superior to strong open-weight models such as the original DeepSeek-R1 and Qwen3-235B, with particular strengths in complex software engineering, tool utilization, and long-context tasks. We publicly release MiniMax-M1 at https://github.com/MiniMax-AI/MiniMax-M1.

HGRN2: Gated Linear RNNswith State ExpansionHGRN2: Gated Linear RNNs with State ExpansionDAPO: An Open-Source LLMReinforcement Learning…DAPO: An Open-Source LLM Reinforcement Learning System at ScaleMiniMax-01: ScalingFoundation Models with…MiniMax-01: Scaling Foundation Models with Lightning AttentionBeyond the 80/20 Rule:High-Entropy Minority…Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningSeed1.5-Thinking:Advancing Superb…Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement LearningSynLogic: SynthesizingVerifiable Reasoning…SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and BeyondSimpleRL-Zoo:Investigating and Tamin…SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the WildLinear-MoE: LinearSequence Modeling Meets…Linear-MoE: Linear Sequence Modeling Meets Mixture-of-ExpertsUnderstandingR1-Zero-Like Training…Understanding R1-Zero-Like Training: A Critical PerspectiveMoM: Linear SequenceModeling with…MoM: Linear Sequence Modeling with Mixture-of-MemoriesThe Entropy Mechanism ofReinforcement Learning…The Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsTitans: Learning toMemorize at Test TimeTitans: Learning to Memorize at Test TimeThe Art of ScalingReinforcement Learning…The Art of Scaling Reinforcement Learning Compute for LLMsSimpleTIR: End-to-EndReinforcement Learning…SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningJet-Nemotron: EfficientLanguage Model with Pos…Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchGLM-4.5: Agentic,Reasoning, and Coding…GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation ModelsFirst Return,Entropy-Eliciting…First Return, Entropy-Eliciting ExploreDefeating theTraining-Inference…Defeating the Training-Inference Mismatch via FP16Agentic Entropy-BalancedPolicy OptimizationAgentic Entropy-Balanced Policy OptimizationASPO: AsymmetricImportance Sampling…ASPO: Asymmetric Importance Sampling Policy OptimizationProsperity beforeCollapse: How Far Can…Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?F-GRPO: Don't Let YourPolicy Learn the Obviou…F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the RareRethinking the TrustRegion in LLM…Rethinking the Trust Region in LLM Reinforcement LearningLow-probability TokensSustain Exploration in…Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable RewardMiniMax-M1: ScalingTest-Time Compute…MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning AttentionEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.