OpenAI launched Ultrafast on Thursday, a new inference mode that runs its flagship GPT-5.6 Sol model at 14x the speed of standard processing and delivers up to 750 output tokens per second. The preview is powered by a partnership with chipmaker Cerebras and is initially available to a small group of customers. It is OpenAI's most direct attempt yet to break the trade-off between model size and response latency.
The pitch is speed without downgrading. Historically, developers who wanted sub-second responses had to route to a smaller, cheaper, and less capable model — Haiku instead of Opus, mini instead of full. Ultrafast keeps GPT-5.6 Sol, OpenAI's most powerful current model, and accelerates it on Cerebras silicon instead.
In a blog post accompanying the release, OpenAI framed the shift plainly.
Key facts
- 01OpenAI's new Ultrafast mode runs GPT-5.6 Sol at 14x the speed of standard processing.
- 02Peak throughput hits 750 output tokens per second, targeting real-time workflows.
- 03The mode is powered by OpenAI's partnership with chipmaker Cerebras.
- 04Ultrafast is in preview with a small group of customers, expanding as capacity grows.
- 05Target use cases include incident response, customer support, financial analysis, and e-commerce.
The 750-tokens-per-second figure is the specification that matters. For comparison, standard GPU-served frontier models typically land in the 50-to-100 tokens-per-second range for large models, which is why the 14x multiplier is being presented as the headline. At that pace, a 3,000-token response — roughly a detailed multi-paragraph answer — completes in about four seconds instead of a minute.
OpenAI is targeting Ultrafast at workflows where latency is the constraint rather than reasoning depth. The company named incident response, customer service and support, financial market analysis, and e-commerce as the initial deployment categories. All four share a pattern: the model quality has to be high, but the human or system consuming the output cannot wait.
The Cerebras dependency is the interesting piece of the announcement. Cerebras builds wafer-scale AI accelerators designed specifically for high-throughput inference, and its architecture has consistently posted faster token-per-second numbers than Nvidia GPU deployments on comparable models. OpenAI running a flagship model on non-Nvidia silicon — even in preview, even for a subset of customers — is a notable diversification.
Competitors have moved in a similar direction but not at this throughput. Anthropic offers a fast mode for Claude that trades some reasoning latency for speed, but it does not reach the 750-tokens-per-second range OpenAI is claiming here. Google's Gemini 3.7 Flash, which shipped three weeks ago at half the price of 3.6, competes on cost and latency but is a distinct, smaller model rather than an acceleration mode on the flagship.
The preview limitation is the caveat. OpenAI has not disclosed pricing for Ultrafast, has not specified which customers currently have access, and has said only that it will expand availability as capacity grows. Cerebras hardware supply is the likely bottleneck — the company's chips are not manufactured at anything close to Nvidia's volume, and dedicating capacity to serve GPT-5.6 Sol inference at 750 tokens per second is not a trivial allocation.
There are also open questions about what Ultrafast costs to run and whether it will be priced accessibly enough to displace the smaller-model workflows it is designed to replace. Real-time customer service at frontier-model quality is genuinely valuable, but only if the per-token economics do not force enterprises back to distilled alternatives.
The strategic read is that OpenAI is treating inference speed as a product surface, not just an engineering metric. Model releases have been converging on capability — every frontier lab now posts strong benchmark numbers — so the next axis of competition is how the model feels in production. Sub-second latency at flagship quality changes what enterprise agents, coding assistants, and support systems can plausibly be built to do, and pairing that with a non-Nvidia silicon partner gives OpenAI a supply story that its rivals will now have to answer.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



