OpenAI on Thursday added three new voice models to its Realtime API: GPT-Realtime-2, GPT-Realtime-Translate and GPT-Realtime-Whisper. The headline model, GPT-Realtime-2, is built on GPT-5-class reasoning and replaces GPT-Realtime-1.5, which shipped with weaker handling of complex multi-turn requests. The translation model covers more than 70 input languages and 13 output languages, while Whisper handles live speech-to-text as conversations unfold.
The pitch is that voice apps stop being scripted call-and-response and start doing actual work mid-conversation. "Together, the models we are launching move real-time audio from simple call-and-response toward voice interfaces that can actually do work: listen, reason, translate, transcribe, and take action as a conversation unfolds," OpenAI said in announcing the release.
Pricing splits across two meters. GPT-Realtime-Translate and GPT-Realtime-Whisper are billed by the minute, which favors steady transcription and translation workloads where audio duration is the relevant cost driver. GPT-Realtime-2 is billed by token consumption, in line with how OpenAI prices its text models, which makes cost forecasting harder for variable-length voice agents but cheaper for short interactions.
Key facts
- 01OpenAI launched GPT-Realtime-2, GPT-Realtime-Translate and GPT-Realtime-Whisper inside its Realtime API on Thursday.
- 02GPT-Realtime-Translate covers 70+ input languages and 13 output languages for live conversational translation.
- 03GPT-Realtime-2 is built on GPT-5-class reasoning, replacing the prior GPT-Realtime-1.5 voice model.
- 04Translate and Whisper are billed by the minute; GPT-Realtime-2 is billed by token consumption.
- 05OpenAI says conversations 'can be halted if they are detected as violating our harmful content guidelines.'
The reasoning upgrade is the part developers will care about. GPT-Realtime-1.5 was capable enough for fixed-script customer service flows but struggled when callers veered off-topic, asked compound questions, or required the model to hold state across several turns. Pushing GPT-5-class reasoning into the audio path is OpenAI's answer to that, and it lines the Realtime API up against Google's Gemini Live and a growing set of voice-agent startups building on top of open-weights models.
“GPT-Realtime-Translate handles more than 70 input languages and outputs 13, aiming to keep pace with a live conversation rather than batch-translating after the fact.”— Jaeden Schafer
Translation is the second pressure point. A 70-language input set with 13 output languages is wide enough to cover most enterprise contact-center deployments, multilingual events, and education products. The asymmetry — many more languages understood than spoken — reflects the hard part of synthetic voice: producing natural-sounding output in a target language is more expensive than recognizing speech in a source one.
OpenAI named customer service, education, media, events and creator platforms as the target buyers. Those are also the verticals where voice-agent startups have been raising at premium multiples for the last 18 months, often by stitching together a third-party speech-to-text model, an LLM, and a separate text-to-speech voice. A single Realtime API call that does all three threatens that integration layer directly.
Abuse is the obvious risk with a model that can hold a fluent live conversation in dozens of languages. OpenAI says it has built guardrails to stop the new features from being used for spam, fraud or other online abuse, with triggers embedded in the system so that "conversations can be halted if they are detected as violating our harmful content guidelines." The company has not detailed how aggressive the classifier is or how often it cuts off legitimate calls.
The translation feature in particular invites scrutiny. Real-time multilingual voice synthesis is exactly the capability scam call centers have been waiting for, and OpenAI's published guardrails are policy commitments rather than technical guarantees. Enforcement will come down to whether OpenAI can detect abuse patterns at the API layer faster than bad actors can iterate on prompts and routing.
There is also a competitive question OpenAI did not address. ElevenLabs, Deepgram, AssemblyAI and a handful of open-source projects have spent the last two years building specialized voice stacks, often with lower latency than general-purpose models. GPT-Realtime-2's reasoning advantage matters for agentic use cases; for pure transcription or pure dubbing, dedicated providers may still win on price and speed.
OpenAI is consolidating the voice stack the same way it consolidated the text stack — one API, one bill, one safety policy, GPT-5-class reasoning underneath. That is good for developers who want to ship a voice agent in a weekend and bad for the middle layer of the voice-AI market, where startups have been charging a premium to glue components together. The Realtime API now does the gluing.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




