Skip to main content
Live
Main content

OpenAI adds GPT-Realtime-2, Translate and Whisper to its Realtime API

The new voice stack handles 70 input languages and 13 output languages, pushing the API beyond simple call-and-response.

Jaeden Schafer
Editor in Chief · · 4 min read
OpenAI logo

OpenAI on Thursday added three new voice models to its Realtime API: GPT-Realtime-2, GPT-Realtime-Translate and GPT-Realtime-Whisper. The headline model, GPT-Realtime-2, is built on GPT-5-class reasoning and replaces GPT-Realtime-1.5, which shipped with weaker handling of complex multi-turn requests. The translation model covers more than 70 input languages and 13 output languages, while Whisper handles live speech-to-text as conversations unfold.

The pitch is that voice apps stop being scripted call-and-response and start doing actual work mid-conversation. "Together, the models we are launching move real-time audio from simple call-and-response toward voice interfaces that can actually do work: listen, reason, translate, transcribe, and take action as a conversation unfolds," OpenAI said in announcing the release.

Pricing splits across two meters. GPT-Realtime-Translate and GPT-Realtime-Whisper are billed by the minute, which favors steady transcription and translation workloads where audio duration is the relevant cost driver. GPT-Realtime-2 is billed by token consumption, in line with how OpenAI prices its text models, which makes cost forecasting harder for variable-length voice agents but cheaper for short interactions.

Key facts

  • 01OpenAI launched GPT-Realtime-2, GPT-Realtime-Translate and GPT-Realtime-Whisper inside its Realtime API on Thursday.
  • 02GPT-Realtime-Translate covers 70+ input languages and 13 output languages for live conversational translation.
  • 03GPT-Realtime-2 is built on GPT-5-class reasoning, replacing the prior GPT-Realtime-1.5 voice model.
  • 04Translate and Whisper are billed by the minute; GPT-Realtime-2 is billed by token consumption.
  • 05OpenAI says conversations 'can be halted if they are detected as violating our harmful content guidelines.'

The reasoning upgrade is the part developers will care about. GPT-Realtime-1.5 was capable enough for fixed-script customer service flows but struggled when callers veered off-topic, asked compound questions, or required the model to hold state across several turns. Pushing GPT-5-class reasoning into the audio path is OpenAI's answer to that, and it lines the Realtime API up against Google's Gemini Live and a growing set of voice-agent startups building on top of open-weights models.

GPT-Realtime-Translate handles more than 70 input languages and outputs 13, aiming to keep pace with a live conversation rather than batch-translating after the fact.
Jaeden Schafer

Translation is the second pressure point. A 70-language input set with 13 output languages is wide enough to cover most enterprise contact-center deployments, multilingual events, and education products. The asymmetry — many more languages understood than spoken — reflects the hard part of synthetic voice: producing natural-sounding output in a target language is more expensive than recognizing speech in a source one.

OpenAI named customer service, education, media, events and creator platforms as the target buyers. Those are also the verticals where voice-agent startups have been raising at premium multiples for the last 18 months, often by stitching together a third-party speech-to-text model, an LLM, and a separate text-to-speech voice. A single Realtime API call that does all three threatens that integration layer directly.

Abuse is the obvious risk with a model that can hold a fluent live conversation in dozens of languages. OpenAI says it has built guardrails to stop the new features from being used for spam, fraud or other online abuse, with triggers embedded in the system so that "conversations can be halted if they are detected as violating our harmful content guidelines." The company has not detailed how aggressive the classifier is or how often it cuts off legitimate calls.

The translation feature in particular invites scrutiny. Real-time multilingual voice synthesis is exactly the capability scam call centers have been waiting for, and OpenAI's published guardrails are policy commitments rather than technical guarantees. Enforcement will come down to whether OpenAI can detect abuse patterns at the API layer faster than bad actors can iterate on prompts and routing.

Related · from this week
Anthropic upgrades Claude voice mode to run on Opus and Sonnet
Jaeden Schafer · 4 min read →

There is also a competitive question OpenAI did not address. ElevenLabs, Deepgram, AssemblyAI and a handful of open-source projects have spent the last two years building specialized voice stacks, often with lower latency than general-purpose models. GPT-Realtime-2's reasoning advantage matters for agentic use cases; for pure transcription or pure dubbing, dedicated providers may still win on price and speed.

OpenAI is consolidating the voice stack the same way it consolidated the text stack — one API, one bill, one safety policy, GPT-5-class reasoning underneath. That is good for developers who want to ship a voice agent in a weekend and bad for the middle layer of the voice-AI market, where startups have been charging a premium to glue components together. The Realtime API now does the gluing.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Anthropic logo
Models

Anthropic upgrades Claude voice mode to run on Opus and Sonnet

Voice mode now taps Opus, Sonnet, and Haiku, plus Gmail, Slack, and Notion — pushing past ChatGPT's tool-less voice.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI updates ChatGPT voice mode to interrupt users less often

GPT-Live-1 replaces the older turn-based voice model with full-duplex audio that can listen while it speaks.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI ships GPT-Live-1, a full-duplex voice model that replaces Advanced Voice Mode

The new model listens and speaks simultaneously, routes to GPT-5.5 for reasoning, and now serves 150M weekly voice users.

Jaeden Schafer4 min read