StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agenti…
What it does
StepFun has unveiled Step 5 Preview, a massive sparse Mixture-of-Experts (MoE) model with a total of 600 billion parameters. It activates 27 billion parameters per token, making it highly efficient in scaling compute without a fully dense parameter load. The model supports an unprecedented 1 million token context window and processes inputs across text, image, and video, targeting complex workflows that require long-horizon reasoning and memory. Step 5 Preview is designed for tasks in software engineering, professional knowledge work, and finance that demand extended context and multimodal understanding.
Why it matters
The 1 million token context window is the standout feature here. Most production language models today handle tens of thousands of tokens at best. Stretching that to a million tokens enables genuinely longer-term reasoning, multi-step agentic workflows, and memory retention that can radically improve automation and assistance in professional domains. The sparse MoE approach also keeps active compute manageable compared to fully dense models of this scale, potentially lowering running costs for sustained, complex tasks. Accepting text, images, and video means it can blend multiple data streams natively for richer context in decision-making applications.
Who it is for
Builders aiming to develop next-level AI assistants or agents that need to track and reason over lengthy documents, codebases, or financial data will find this model especially relevant. Companies focused on automating professional knowledge work or providing AI-driven software engineering support gain stronger models for long-horizon tasks. Investors tracking advances in model sparsity and efficiency will note StepFun’s approach as a competitive alternative to monolithic large dense models, which often demand prohibitively high compute and infrastructure costs.
The catch
API access is priced at $1.00 per 1 million input tokens and $2.70 per 1 million output tokens, which is not cheap but aligns with pricing for models with larger context windows and multimodal capabilities. The open weights are not immediately available; StepFun plans to release them on October 15, 2026, delaying wide academic or community adoption. Long context windows also pose real engineering and latency challenges, which mean practical deployment and application integration will still require significant effort.
What to watch next
Watch for how builders integrate Step 5 Preview into real-world workflows, especially in software engineering and finance where long-term agentic tasks have high value. Monitor how the model’s sparse architecture influences operational costs compared to fully dense competitors. The open weight release in late 2026 will be crucial for researchers and startups seeking to experiment with or build on this type of scale and multimodal capability. Lastly, gauge how end users respond to the practicality of handling million-token contexts in actual applications.
AI Quick Briefs Editorial Desk