Skip to main content
Live
Main content

Smallest.ai raises $13M to build small voice models that mimic human turn-taking

The India-founded startup is betting specialized voice models beat LLMs at real-time conversation, with RingCentral and Truecaller already on board.

Jaeden Schafer
Editor in Chief · · 4 min read
Smallest.ai raises $13M to build small voice models that mimic human turn-taking

Smallest.ai raised $13 million in Series A funding to build small, specialized voice models that mimic how humans actually talk — listening, thinking, and speaking at the same time rather than waiting for a full prompt to finish. The round was led by Seligman Ventures with participation from Sierra Ventures and 3one4 Capital, pushing the startup's total funding above $21 million less than two years after it was founded in late 2024. Customers already using its voice stack include RingCentral and Truecaller.

Founder and CEO Sudarshan Kamath's pitch is architectural, not incremental. Instead of squeezing lower latency out of a large language model, Smallest.ai runs a purpose-built voice model as a real-time intelligence layer, then hands off to a larger foundation model only when the conversation moves outside its knowledge base. On those handoffs, the agent briefly places the caller on hold to research the issue — the same beat a human agent would take.

The reason for the split is that LLM latency, which is tolerable in a chat window, breaks the illusion in a live phone call. Kamath frames the fix as matching the physiology of human conversation itself: partial listening, overlapping thought, interruption.

The way an LLM works is you give it an entire prompt, and then it starts thinking
Sudarshan Kamath, Founder and CEO of Smallest.ai

Key facts

  • 01Smallest.ai raised $13M in Series A led by Seligman Ventures, bringing total funding above $21M.
  • 02The startup was founded in late 2024 and builds small voice models rather than large foundation models.
  • 03Existing customers include RingCentral and Truecaller; the model supports dozens of languages.
  • 04Smallest.ai competes directly with ElevenLabs, Cartesia, and India-focused rival Sarvam.
  • 05Architecture pairs a real-time small voice model with an offline LLM handoff for complex queries.

That framing drives what Smallest.ai chooses not to optimize for. Its models don't chase dubbing, podcasting, or general-purpose audio generation. They focus narrowly on real-time conversational agents that handle diverse accents, work across dozens of languages, and hold up in noisy environments — the specific failure modes that make current voice bots instantly identifiable.

The commercial thesis is that customer support companies — including well-funded incumbents like Sierra and Decagon — are potential customers rather than competitors, because building a world-class voice model pulls engineering focus away from the actual support product. Kamath said becoming extremely good at doing voice is a distraction from their core business. Smallest.ai wants to be the layer they license instead.

The competitive landscape is already crowded. ElevenLabs leads the category on brand and breadth, Cartesia is pushing on real-time synthesis, and Indian rival Sarvam is building for local-language coverage. Smallest.ai's wedge is enterprise-grade real-time conversation, not creator tools.

The Turing-test framing is aggressive, and Kamath owns it. He treats indistinguishability from a human agent not as a marketing line but as the product spec.

The counterweight is that voice indistinguishability is a moving target every major lab is also chasing, and the enterprise buyers Smallest.ai is targeting tend to consolidate on whichever provider clears procurement, security review, and integration first — not necessarily whichever model sounds most human in a demo. The startup will also need to prove that its two-model architecture, with its research pauses on hard queries, feels more natural in production than a single fast model that never breaks stride. And $21 million total funding is a modest war chest against ElevenLabs, which has raised in the hundreds of millions.

Related · from this week
Fish Audio raises $52M seed at $21M ARR for AI voice models
Jaeden Schafer · 5 min read →

The bet worth watching is the architectural one. If Smallest.ai is right that voice agents converge on a small real-time model plus an offline LLM, that pattern reshapes the buy-versus-build calculus for every customer support platform now assembling its own voice stack, and it gives specialized voice vendors a durable seat under the agent layer rather than getting absorbed into it. If the market instead concludes that one large multimodal model can handle both roles well enough, the specialist tier compresses fast. The next 18 months of enterprise voice deployments will settle which side is right.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Business

Fish Audio raises $52M seed at $21M ARR for AI voice models
Business

Fish Audio raises $52M seed at $21M ARR for AI voice models

The Palo Alto startup has 8 million users and 15,000 natural-language voice controls, with three of its five models open-sourced.

Jaeden Schafer5 min read
Spotify launches ElevenLabs-powered audiobook creation tool for self-publishers
Business

Spotify launches ElevenLabs-powered audiobook creation tool for self-publishers

The streaming platform will let authors generate audiobooks in-app starting June 2026, with no exclusivity required.

Jaeden Schafer5 min read
Wispr Flow's India growth hits 100% as Hinglish voice push lands
Business

Wispr Flow's India growth hits 100% as Hinglish voice push lands

The Bay Area voice-input startup says India is now its second-largest market, with 14% of 2.5M global downloads and aggressive local pricing.

Jaeden Schafer5 min read