MiniMax M1 40k
MiniMax · 2025-06-16 · 456.1B parameters
MiniMax M1 40K is MiniMax's open-weight hybrid-attention reasoning checkpoint with a 40K maximum generation length; MiniMax says it represents an intermediate phase of the 80K variant's training. The 456-billion-parameter MoE activates approximately 45.9 billion parameters per token, places one softmax-attention block after every seven Lightning Attention blocks, is described as natively supporting up to 1 million context tokens, and was trained with large-scale reinforcement learning using CISPO.