Thinking Machines, the AI lab founded by former OpenAI CTO Mira Murati, said on May 11, 2026 that it is building what it calls 'interaction models' — systems designed to continuously process audio, video, and text in real time rather than waiting for a user to finish a turn. The company plans a limited research preview in the coming months and a wider release later this year. It is the most concrete public direction Thinking Machines has staked out since Murati incorporated the company in February 2025.
The framing is a direct critique of how current chatbots work. 'Today's models experience reality in a single thread,' Thinking Machines wrote in its announcement. 'Until the user finishes typing or speaking, the model waits with no perception of what the user is doing or how the user is doing it.' The company argues that the same freeze happens in reverse: once a model starts generating, it stops taking in new information until it finishes or is interrupted.
Thinking Machines calls this a 'bandwidth bottleneck' and pitches interaction models as the fix. 'We believe we can solve this bandwidth bottleneck by making AI interactive in real time across any modality,' the company said. The stated goal is to let 'AI interfaces meet humans where they are, rather than forcing humans to contort themselves to AI interfaces.'
Key facts
- 01Thinking Machines announced its 'interaction models' research direction on May 11, 2026.
- 02The lab was founded by former OpenAI CTO Mira Murati in February 2025.
- 03A limited research preview is planned for the coming months, with a wider release later this year.
- 04Demos included real-time speech translation, animal-mention detection in a story, and posture correction.
- 05The company has lost staff to Meta and back to OpenAI since its founding.
The demonstrations released alongside the announcement are deliberately mundane. One model listens to a story being read aloud and flags every mention of an animal. Another translates speech in real time. A third watches a user on camera and tells them when they are slouching. None of these are headline capabilities on their own; the claim is that the underlying architecture — continuous perception and continuous output — is the novel piece.
“Today's models experience reality in a single thread. Until the user finishes typing or speaking, the model waits with no perception of what the user is doing or how the user is doing it.”— Jaeden Schafer
The pitch sits in contrast to the prevailing chatbot interface, where a turn-based protocol forces both sides to take strict turns. Voice modes from OpenAI, Google, and others have closed some of that latency gap, but they are still fundamentally turn-based under the hood. Thinking Machines is arguing that the right unit of work is not the message but the ongoing stream.
Murati left OpenAI in 2024 after serving as chief technology officer and briefly as interim CEO during the November 2023 board crisis. She incorporated Thinking Machines in February 2025 and pulled in a roster of senior research talent from OpenAI and elsewhere, though the company has not publicly disclosed funding details, headcount, or a product timeline beyond today's research note.
The lab has had a turbulent first year on the personnel side. Several key members have defected to Meta's superintelligence-focused unit, and at least one prominent researcher has gone back to OpenAI. The interaction-models announcement is in part a signal to the market — and to current and prospective hires — that the company has a concrete technical agenda beyond 'ex-OpenAI talent doing something with frontier models.'
A 'limited research preview' is also a deliberately narrow promise. Thinking Machines is not yet committing to a consumer product, an API, or pricing. Real-time multimodal systems are expensive to serve, and the company has not said whether the preview will run on its own infrastructure, a hyperscaler partner, or a hybrid setup. The wider release later this year is the milestone to watch.
The technical bar is high. Continuous bidirectional perception across audio, video, and text means the model has to handle interruption, partial input, and partial output without losing coherence — problems that have tripped up every voice-first AI interface to date. Slouch detection and animal-name spotting are useful proofs of concept, but the harder question is whether interaction models hold up on tasks where the cost of misinterpreting a stream is high, like coding assistance or medical scribing.
For the broader AI market, the interesting variable is whether 'always-on' models become the default interface or remain a specialized mode. If Thinking Machines can ship something that genuinely feels like talking to a colleague rather than dictating to a chatbot, it changes the surface area on which OpenAI, Anthropic, and Google have to compete — and gives a year-old startup a defensible wedge that does not require beating the incumbents on raw benchmark scores. That is a bet worth making, and it is the first time Murati's company has clearly told the market which bet it is.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




