Mira Murati's Thinking Machines Lab previewed a new class of AI it calls interaction models, systems trained to communicate through a camera and microphone rather than a text box. The former OpenAI CTO told WIRED the goal is to keep humans inside the loop as the industry pushes toward superintelligence. The lab has raised billions of dollars since Murati founded it after leaving OpenAI in 2024, and this is only its second product reveal — the first, Tinker, shipped in October 2025.
Unlike conventional voice interfaces that transcribe speech and pass it to a chatbot, the interaction models process continuous audio and video natively. That means they read pauses, interruptions, and shifts in tone instead of waiting for a clean utterance to parse. Thinking Machines demonstrated the capability in several videos but has not opened the models to outside users.
Alexander Kirillov, a founding team member working on multimodal AI, framed the technical gap directly. "The model constantly perceives what you're doing and is constantly there to be able to reply and give you information or search for information or use other tools," he said. "This is something that none of [today's other] models can actually do. The turns [in a conversation] are determined by a much less intelligent system."
Key facts
- 01Thinking Machines Lab previewed interaction models that process camera and microphone input natively, reading pauses, interruptions, and tone shifts.
- 02The lab has raised billions of dollars since Murati left OpenAI as CTO in 2024.
- 03Tinker, the company's first product, launched in October 2025 as an API for fine-tuning open-source models.
- 04Founding team member Alexander Kirillov says current voice models rely on a 'much less intelligent system' to decide conversational turns.
- 05Thinking Machines positions its approach against OpenAI, Anthropic, and Google, which are building models that write software from a single text prompt.
The pitch is a deliberate counter to where OpenAI, Anthropic, and Google are pointing their largest models. Those labs are pushing toward agents that take a single prompt and produce entire software applications with minimal human input. Murati is arguing for the opposite vector: AI that gets more useful as the human stays present.
“At some point we will have super-intelligent machines. But we think that the best way to actually have many possible futures—good futures—is to keep humans in the loop.”— Jaeden Schafer
"At some point we will have super-intelligent machines," Murati told WIRED. "But we think that the best way to actually have many possible futures—good futures—is to keep humans in the loop." She described the interaction models as a first step toward AI that amplifies user preferences rather than substituting for them.
Tinker, the lab's October 2025 release, sits in the same philosophical bucket. It is an API that lets researchers and engineers fine-tune open-source frontier models on their own data, shifting some of the customization power away from the lab and toward the user. The interaction models extend that posture from training-time control to runtime collaboration.
Thinking Machines is not alone in the human-collaboration framing. Other startups, including Humans&, are pursuing AI systems built around collaboration rather than automation, and several economists have publicly argued the field should orient toward human empowerment instead of labor replacement. The frame is also a marketing wedge: every major lab now has a story about agents replacing knowledge work, and few have a story about working with users at the level of voice and gesture.
"This is showing the first bet on human collaboration," Murati said of the interaction models. "Where this is going is really amplifying people's own preferences and values, with AI actually understanding intent and predicting intent." That language — intent prediction, value amplification — is the durable positioning Thinking Machines appears to be staking out.
The open question is delivery. Thinking Machines has raised at a valuation that implies frontier-lab ambitions, but it has shipped one public product in roughly 18 months and the interaction models remain demoware. The lab has not disclosed which base models the interaction system runs on, what its inference cost looks like, or when external developers will get access. Until those land, the human-in-the-loop pitch is a thesis, not a benchmark.
For the broader market, Thinking Machines is an interesting hedge against the dominant agent-autonomy narrative. If OpenAI, Anthropic, and Google are right, the long-run product is software that runs itself and Murati's lab is building a feature. If the autonomy thesis stalls on reliability — which is where most enterprise agent pilots have stuck — then interfaces that read tone and intent in real time become a defensible category of their own, and the company holding the best multimodal interaction stack has a real business.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



