GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras
What it does
OpenAI introduced a new inference mode called Ultrafast for GPT-5.6 Sol that increases output speed up to 750 tokens per second. This mode runs on specialized Cerebras hardware, a result of OpenAI’s $10 billion partnership with the chipmaker. Ultrafast squeezes 14 times the speed out of the model compared to previous versions, complementing existing Standard and Fast modes to form a three-tier pricing and performance structure.
Why it matters
Speed has been a bottleneck for real-time applications using large language models. By delivering outputs significantly faster, Ultrafast mode lowers latency, making high-volume, time-sensitive tasks like live chat, automated content generation, and interactive AI tools more practical. It shifts the focus from just model capability to inference speed as a standalone product feature. This puts price and performance control directly in the hands of users, who can now choose the right cost-speed tradeoff for their workloads.
Who it is for
The new tiered model pricing appeals to businesses and developers facing strict performance SLAs or handling massive parallel inference at scale. Services processing live customer interactions or updating frequently changing content stand to benefit most. Investors and system operators can use Ultrafast to accelerate workflows and reduce cloud compute expenses per project by cutting request wait times, working smarter on high-volume tasks.
The catch
Ultrafast mode requires running on Cerebras hardware infrastructure, which is not yet widespread. Adoption depends on OpenAI’s and Cerebras’s ability to scale access and pricing to the bulk market. For now, cheaper Standard and Fast modes remain the mainstream options for users without stringent latency or throughput needs. The price delta between tiers could impact marginal cost calculations for businesses relying on heavy inference workloads.
What to watch next
How quickly OpenAI expands Cerebras-powered availability across data centers will determine Ultrafast’s market reach. Watch for developer adoption patterns that weigh speed gains against incremental pricing. Competitors may respond by optimizing their own model inference speeds or unveiling alternative hardware partnerships. Early user feedback will also clarify if Ultrafast makes sense beyond high-end enterprise customers or AI-first startups.
AI Quick Briefs Editorial Desk