Thinking Machines Lab announced a new class of models on May 11, 2026 that it calls interaction models, designed to listen and speak at the same time rather than swap turns with the user. The flagship, TML-Interaction-Small, responds in 0.40 seconds — roughly the cadence of natural human conversation, and faster than comparable systems from OpenAI and Google, according to the company. Founder Mira Murati, the former OpenAI CTO who started Thinking Machines last year, is positioning the work as a rethink of how voice AI is structured at the model level.
The technical term is 'full duplex.' Today's voice assistants run a sequential loop: the user speaks, the model transcribes and processes, then the model generates and plays a response. Thinking Machines says TML-Interaction-Small does both at once, processing input audio while producing output audio in parallel.
The 0.40-second figure is the headline benchmark. Human conversational turn-taking typically lands in the 200-millisecond range, and current voice agents from OpenAI and Google generally sit well above that even with streaming optimizations. Closing the gap to roughly 400 milliseconds, if it holds up outside controlled tests, would put a model close to phone-call latency rather than walkie-talkie latency.
Key facts
- 01Thinking Machines Lab announced 'interaction models' on May 11, 2026, led by founder and former OpenAI CTO Mira Murati.
- 02TML-Interaction-Small responds in 0.40 seconds, which the company says beats comparable models from OpenAI and Google.
- 03The architecture is 'full duplex' — the model processes input and generates output simultaneously, rather than alternating turns.
- 04A limited research preview is coming in the next few months, with wider release later this year.
- 05Thinking Machines was founded last year and has not yet shipped a public product.
The framing matters as much as the speed. Most voice features today are built by stitching speech-to-text, a text model, and text-to-speech together — interruption handling and overlap are bolted on at the orchestration layer. Thinking Machines is arguing that interactivity belongs inside the model itself, not in the wrapper around it.
“TML-Interaction-Small responds in 0.40 seconds — roughly the cadence of a natural human exchange, and faster than comparable models from OpenAI and Google.”— Jaeden Schafer
This is a research preview, not a product. Thinking Machines is not releasing TML-Interaction-Small publicly yet. A limited research preview is planned in the next few months, with wider release set for later this year.
Thinking Machines has been quiet on product since its founding last year, raising attention more for its roster than its shipped work. Murati left OpenAI in 2024 and built the company around a thesis that interactivity and personalization are underexplored axes relative to raw model scaling. Interaction models are the first concrete artifact of that thesis.
The competitive frame is straightforward. OpenAI's Realtime API and Advanced Voice Mode are the closest reference points, along with Google's Gemini Live. Both have shipped low-latency voice over the past 18 months, and both still rely on architectures that fundamentally alternate listening and speaking. A model trained from the ground up for simultaneous I/O would be a structural departure rather than an incremental tuning win.
There are real reasons to wait on judgment. Benchmarks measured by the lab that built the model are not the same as field performance, and 'natural conversation' covers everything from clean studio audio to a noisy car. Interruption handling — knowing when to yield, when to keep going, when to stop mid-sentence — is the part that has historically broken voice agents in production, and a 0.40-second response time does not by itself prove the model gets that right.
Thinking Machines also has not disclosed parameter count, training data, pricing, or deployment targets for TML-Interaction-Small. The 'Small' label implies a larger sibling, but the company has not confirmed one. Without an API, third-party evaluations, or a comparison harness against the OpenAI and Google voice stacks, the 0.40-second claim is the company's own.
If full duplex holds up outside the demo, the implication for the voice AI stack is significant. Customer support, in-car assistants, accessibility tools, and consumer voice apps have all been bounded by the same turn-taking ceiling, and the workarounds — barge-in detection, predictive prefetching, aggressive endpointing — are engineering patches on a structural limit. A model that listens while it talks resets the ceiling, and forces the existing voice players to decide whether to retrofit their stacks or rebuild them. For a company that has not shipped a product yet, that is a sharper opening shot than most expected from Thinking Machines this year.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




