The MiniMax M3 model officially launched on June 1, 2026, introducing a multimodal foundation model with a 1 million token context window.

This new release supports text, image, and video inputs for complex enterprise workflows. Meanwhile, developers can access the system through the EvoLink platform using standard OpenAI-compatible endpoints.

Notably, the model utilizes a new architecture called MiniMax Sparse Attention to process ultra-long contexts efficiently. Unlike other compression methods, this system operates on a standard Grouped-Query Attention backbone. Consequently, it uses block-level selection on real, uncompressed Key-Value blocks to reduce computational costs.

Launch of the MiniMax M3 model

The MiniMax M3 model represents a significant shift in how the company handles long-context data processing.

Specifically, the company claims the system achieves 15.6 times faster decoding speeds compared to its previous M2 predecessor. Furthermore, prefill speeds are reportedly 9.7 times faster when handling contexts up to 1 million tokens.

Technical Architecture and Performance

However, the company has not yet published independent accuracy or quality tradeoff data alongside these speed benchmarks. Therefore, third-party validation remains necessary to confirm if these speed improvements affect reasoning quality. In addition, the model features a maximum output limit of 131,072 tokens per request.

Development History and Reasoning Capabilities

During previous development cycles, researchers found that sub-quadratic shortcuts often reduced multi-hop reasoning capabilities. As a result, the team initially absorbed the high computational costs of full quadratic attention for older models. Now, the **MiniMax M3 model** uses a novel sub-quadratic framework designed to maintain reasoning quality.

“At larger scale, linear and windowed attention variants exhibited severe reasoning deficits.”

MiniMax Research Team

Market Availability and Pricing

Regarding corporate background, MiniMax was established in 2022 and went public on the Hong Kong Stock Exchange in January 2026. The company receives financial backing from major technology firms including Tencent, Alibaba, and miHoYo. Currently, the firm serves over 200 million users through its consumer applications.

API pricing for the **MiniMax M3 model** is set at 0.30 dollars per million input tokens. Meanwhile, output tokens cost 1.20 dollars per million. Additionally, the model weights will likely be downloadable from Hugging Face, though commercial use requires written authorization.

Specifically, the model is optimized for document understanding, spreadsheet processing, and enterprise agent workflows. These capabilities allow businesses to analyze large codebases and multi-document datasets. Consequently, this release could make ultra-long-context inference economically viable for a wider range of enterprise customers.

Source: X (@intheworldofai) — researched