OpenAI's Ultrafast mode speeds up GPT-5.6 Sol to 750 tokens per second
Summarized by AI from reporting by OpenAI Blog, published under our editorial policy.
OpenAI previewed Ultrafast, a new API service tier for GPT-5.6 Sol that delivers up to 14× faster inference, reaching 750 output tokens per second via Cerebras hardware.

Key takeaways
- OpenAI previewed Ultrafast, a new API service tier for GPT-5.6 Sol that delivers up to 14× faster inference.
- Ultrafast mode can generate up to 750 output tokens per second.
- The speed boost is powered by Cerebras's specialized AI hardware.
OpenAI previewed a new Ultrafast mode for GPT-5.6 Sol, a version of their latest AI model that can now generate responses up to 14 times faster. This speed boost is powered by Cerebras, a company specializing in AI hardware, and can produce up to 750 output tokens per second.
How Ultrafast mode works
Ultrafast mode is a new service tier in OpenAI's API that significantly increases the speed of GPT-5.6 Sol. In practical terms, this means that AI-generated text—like responses from chatbots or summaries of long documents—can now be produced much more quickly. For example, a response that might have taken a few seconds could now take just a fraction of that time.
Speed benchmarks: 750 tokens per second
The Ultrafast mode can generate up to 750 output tokens per second. To put that into perspective, a typical sentence might contain around 15-20 tokens, so this mode could generate about 35-50 sentences per second. This is a massive leap from previous speeds, which were much slower. The mode is powered by Cerebras's AI hardware, which is designed to handle complex computations quickly and efficiently.
Why speed matters for real-time AI
This speed boost could change how we interact with AI in our daily lives. Imagine using a chatbot to get quick answers, or an AI assistant to help you draft emails—Ultrafast mode could make these interactions feel almost instant. This could be particularly useful in professional settings, like customer service or content creation, where speed is crucial. It could also make AI tools more accessible and user-friendly for everyone.
How to try Ultrafast mode
If you're a developer or a business using OpenAI's API, you can start experimenting with Ultrafast mode right away. OpenAI has made it available as a new service tier, so you can integrate it into your applications and see the difference in speed. If you use apps that rely on OpenAI's models, keep an eye out for updates—you might start noticing faster responses soon.
Frequently asked
- Is Ultrafast mode available to all users?
- Ultrafast mode is currently available as a new service tier in OpenAI's API, primarily for developers and businesses.
- How does Ultrafast mode compare to previous versions?
- Ultrafast mode is up to 14 times faster than previous versions of GPT-5.6 Sol.
- What kind of hardware powers Ultrafast mode?
- Ultrafast mode is powered by Cerebras's specialized AI hardware.