Skip to main content
Live
Main content

Thinking Machines releases Inkling, its first open-weight AI model

Mira Murati's startup ships a 975B-parameter mixture-of-experts model, betting enterprises will fine-tune their own AI rather than rent frontier chatbots.

Jaeden Schafer
Editor in Chief · · 5 min read
Thinking Machines releases Inkling, its first open-weight AI model

Thinking Machines Lab released Inkling on Wednesday, its first proprietary model and, notably, an open-weight one that developers and enterprises can download and modify directly. The startup, founded by former OpenAI CTO Mira Murati, built Inkling as a mixture-of-experts system with 975 billion total parameters, activating about 41 billion for any given task. It was trained on 45 trillion tokens of text, image, audio, and video, and reasons natively across all three modalities.

The release lands about nine months after Thinking Machines began building — a compressed timeline against the roughly five years OpenAI took and the three years Anthropic took to reach comparable milestones. It's the company's first public proof point after a year and a half of infrastructure work, following a May research preview of "interaction models" designed to listen, speak, and even interrupt rather than wait passively like typical chatbots.

On one internal benchmark, Thinking Machines says Inkling uses a third as many tokens as Nvidia's Nemotron 3 Ultra to hit the same coding performance. The model is designed to give calibrated answers, flagging uncertainty rather than guessing, and lets users dial thinking effort up or down to trade accuracy for speed.

Key facts

  • 01Inkling is a mixture-of-experts model with 975 billion total parameters, activating about 41 billion per task.
  • 02The model was trained on 45 trillion tokens across text, image, audio, and video on Nvidia's GB300 NVL72 systems.
  • 03Thinking Machines says Inkling hits the same coding performance as Nvidia's Nemotron 3 Ultra using one third as many tokens.
  • 04A Bridgewater collaboration scored 84.7% on financial reasoning tests at roughly one-fourteenth the run cost of top proprietary models.
  • 05Thinking Machines reached model release in about nine months, versus roughly five years for OpenAI and three for Anthropic.

Thinking Machines is not claiming a frontier crown. Its briefing materials state the model is not the strongest available today, closed or open. The company is instead marketing Inkling as well-rounded and, critically, as a starting point for enterprises to fine-tune through Tinker, its model-customization platform.

not the strongest model available today, closed or open.
Thinking Machines Lab, briefing materials

That's the bet. OpenAI, Anthropic, and Google built ChatGPT, Claude, and Gemini as general-purpose chatbots first, with agentic features layered on. Thinking Machines is arguing the opposite: that a model organizations can adapt for themselves will outperform the one-size-fits-all systems the biggest labs currently sell. A company blog post last week set up the release by arguing that centrally trained, frozen models underperform ones shaped by the enterprises that hold the domain expertise.

The argument is picking up outside endorsers. Microsoft CEO Satya Nadella — whose company has invested billions in both OpenAI and Anthropic — wrote in a Sunday blog post that enterprises using proprietary AI models effectively pay twice: once in subscription costs, and again by handing over business knowledge embedded in prompts and corrections that get absorbed into future model versions. Hugging Face CEO Clem Delangue offered a parallel forecast last week.

The clearest data point in Thinking Machines' pitch comes from a joint project with Bridgewater Associates, the world's largest hedge fund and not a Thinking Machines investor. Researchers took an existing open-source model and trained it further on Bridgewater's financial expertise. The result scored 84.7% on financial reasoning tests, beating top proprietary models while costing roughly a fourteenth as much to run. The results, published jointly in late June, come from the two companies' own evaluation rather than an independent one.

On training provenance, Thinking Machines says it pretrained Inkling from scratch but used other open-weight models — including Moonshot AI's Kimi K2.5 — to generate some early post-training data before reinforcement learning took over. The company says its next model will use fully self-contained post-training. On compute, it struck a strategic partnership with Nvidia in March to deploy a gigawatt of Vera Rubin capacity, and Inkling itself was trained on Nvidia's GB300 NVL72 systems.

Related · from this week
Mira Murati previews Thinking Machines' interaction models with humans in the loop
Jaeden Schafer · 4 min read →

The economics remain the open question. A reported $50 billion fundraising round was said to be coming together in November before stalling by January, and the company has declined to discuss funding since, though Nvidia has confirmed a significant investment alongside the March partnership. Once open weights ship, nothing obligates a downloader to pay Thinking Machines to run them, unlike the metered API access OpenAI and Anthropic sell. Revenue has to come through Tinker — the training, fine-tuning, and hosting layer built around the model.

The strategic wager here is that Thinking Machines doesn't need to match OpenAI's or Anthropic's spend because it isn't playing the same game. If enterprises really do migrate production workloads to customized open-weight systems while reserving frontier chatbots for experimentation, the winner isn't whoever ships the highest benchmark score — it's whoever owns the fine-tuning and deployment layer. Inkling is the loss leader; Tinker is the business. Whether that split materializes is the question every open-weight vendor is now betting on, and Murati has just placed her chip.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Mira Murati previews Thinking Machines' interaction models with humans in the loop
Models

Mira Murati previews Thinking Machines' interaction models with humans in the loop

The ex-OpenAI CTO's lab shows AI that natively reads pauses, interruptions, and tone through camera and mic — its second product since launch.

Jaeden Schafer4 min read
Mira Murati surfaces at Thinking Machines with a new bet on real-time AI
Business

Mira Murati surfaces at Thinking Machines with a new bet on real-time AI

The former OpenAI CTO previewed 'interaction models' that process audio, text, and video in 200-millisecond intervals — her first major appearance in 18 months.

Jaeden Schafer5 min read
Thinking Machines unveils 'full duplex' AI that listens while it talks
Models

Thinking Machines unveils 'full duplex' AI that listens while it talks

Mira Murati's startup says TML-Interaction-Small responds in 0.40 seconds, matching the cadence of natural human conversation.

Jaeden Schafer4 min read