Skip to main content
Live
Main content

OpenAI launches Ultrafast mode, pushing GPT-5.6 Sol to 750 tokens per second

The new preview mode runs at 14x standard speed, powered by a Cerebras partnership, and targets enterprise workflows where latency is the constraint.

Jaeden Schafer
Editor in Chief · · 4 min read
OpenAI logo

OpenAI launched Ultrafast on Thursday, a new inference mode that runs its flagship GPT-5.6 Sol model at 14x the speed of standard processing and delivers up to 750 output tokens per second. The preview is powered by a partnership with chipmaker Cerebras and is initially available to a small group of customers. It is OpenAI's most direct attempt yet to break the trade-off between model size and response latency.

The pitch is speed without downgrading. Historically, developers who wanted sub-second responses had to route to a smaller, cheaper, and less capable model — Haiku instead of Opus, mini instead of full. Ultrafast keeps GPT-5.6 Sol, OpenAI's most powerful current model, and accelerates it on Cerebras silicon instead.

In a blog post accompanying the release, OpenAI framed the shift plainly.

Key facts

  • 01OpenAI's new Ultrafast mode runs GPT-5.6 Sol at 14x the speed of standard processing.
  • 02Peak throughput hits 750 output tokens per second, targeting real-time workflows.
  • 03The mode is powered by OpenAI's partnership with chipmaker Cerebras.
  • 04Ultrafast is in preview with a small group of customers, expanding as capacity grows.
  • 05Target use cases include incident response, customer support, financial analysis, and e-commerce.

The 750-tokens-per-second figure is the specification that matters. For comparison, standard GPU-served frontier models typically land in the 50-to-100 tokens-per-second range for large models, which is why the 14x multiplier is being presented as the headline. At that pace, a 3,000-token response — roughly a detailed multi-paragraph answer — completes in about four seconds instead of a minute.

OpenAI is targeting Ultrafast at workflows where latency is the constraint rather than reasoning depth. The company named incident response, customer service and support, financial market analysis, and e-commerce as the initial deployment categories. All four share a pattern: the model quality has to be high, but the human or system consuming the output cannot wait.

The Cerebras dependency is the interesting piece of the announcement. Cerebras builds wafer-scale AI accelerators designed specifically for high-throughput inference, and its architecture has consistently posted faster token-per-second numbers than Nvidia GPU deployments on comparable models. OpenAI running a flagship model on non-Nvidia silicon — even in preview, even for a subset of customers — is a notable diversification.

Competitors have moved in a similar direction but not at this throughput. Anthropic offers a fast mode for Claude that trades some reasoning latency for speed, but it does not reach the 750-tokens-per-second range OpenAI is claiming here. Google's Gemini 3.7 Flash, which shipped three weeks ago at half the price of 3.6, competes on cost and latency but is a distinct, smaller model rather than an acceleration mode on the flagship.

The preview limitation is the caveat. OpenAI has not disclosed pricing for Ultrafast, has not specified which customers currently have access, and has said only that it will expand availability as capacity grows. Cerebras hardware supply is the likely bottleneck — the company's chips are not manufactured at anything close to Nvidia's volume, and dedicating capacity to serve GPT-5.6 Sol inference at 750 tokens per second is not a trivial allocation.

Related · from this week
OpenAI lifts text-chat limits for ChatGPT free users, rolls out GPT-5.6
Jaeden Schafer · 4 min read →

There are also open questions about what Ultrafast costs to run and whether it will be priced accessibly enough to displace the smaller-model workflows it is designed to replace. Real-time customer service at frontier-model quality is genuinely valuable, but only if the per-token economics do not force enterprises back to distilled alternatives.

The strategic read is that OpenAI is treating inference speed as a product surface, not just an engineering metric. Model releases have been converging on capability — every frontier lab now posts strong benchmark numbers — so the next axis of competition is how the model feels in production. Sub-second latency at flagship quality changes what enterprise agents, coding assistants, and support systems can plausibly be built to do, and pairing that with a non-Nvidia silicon partner gives OpenAI a supply story that its rivals will now have to answer.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

OpenAI logo
News

OpenAI lifts text-chat limits for ChatGPT free users, rolls out GPT-5.6

Free and Go tiers get unlimited text chats and a Think button next week; GPT-5.6 Luna becomes the default model.

Jaeden Schafer4 min read
OpenAI logo
Models

César de la Fuente's lab uses OpenAI's Codex and ChatGPT to hunt new antibiotics

The University of Pennsylvania team mines living and extinct genomes with GPT-powered tools to surface antimicrobial candidates against drug-resistant infections.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI updates ChatGPT voice mode to interrupt users less often

GPT-Live-1 replaces the older turn-based voice model with full-duplex audio that can listen while it speaks.

Jaeden Schafer4 min read