Models & Research

Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen …

· August 3, 2026
Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen …

What happened

Alibaba’s Qwen team officially launched Qwen3.8-Max, a massive 2.4 trillion parameter Mixture of Experts (MoE) model. After a preview phase, Qwen3.8-Max is now generally available with published per-token pricing. It supports multi-modal input including text, images, and video, handling an impressive 1 million token context window. The team plans to release the model’s weights openly next week. However, no benchmark comparisons have been made public so far.

Why it matters

Qwen3.8-Max pushes the scale of multi-modal AI models well beyond many existing alternatives. Its 2.4 trillion parameters, combined with MoE architecture, enable more efficient scaling by activating only parts of the model per task. The ability to process text, images, and video together over an extremely long context means it can potentially support complex real-world workflows requiring extensive context and cross-media understanding. Publishing per-token pricing signals readiness for commercial API use and cost transparency. Open weights will invite independent evaluation and experimentation, which could accelerate ecosystem adoption or expose limitations.

However, without benchmarks, it is unclear how it performs relative to other state-of-the-art models in practical scenarios. Since MoE models vary significantly in quality depending on routing and training, builders and investors will want to see real workload numbers before committing.

What to watch next

The imminent release of Qwen3.8-Max’s weights will be key. Model researchers and operators can then test performance, fine-tune, or integrate it into workflows. Watch for independent benchmarks on tasks like multi-modal reasoning, video understanding, or long-context NLP to get a clearer picture of its capabilities and efficiency. Also monitor pricing in relation to available compute and competing APIs, as this will influence adoption among startups and enterprises balancing cost and performance. Finally, Qwen’s multi-modal, super long-context support may pressure other AI providers to accelerate multi-domain context models for complex real-world applications.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.