Smallest.ai raised $13 million in Series A funding to build small, specialized voice models that mimic how humans actually talk — listening, thinking, and speaking at the same time rather than waiting for a full prompt to finish. The round was led by Seligman Ventures with participation from Sierra Ventures and 3one4 Capital, pushing the startup's total funding above $21 million less than two years after it was founded in late 2024. Customers already using its voice stack include RingCentral and Truecaller.
Founder and CEO Sudarshan Kamath's pitch is architectural, not incremental. Instead of squeezing lower latency out of a large language model, Smallest.ai runs a purpose-built voice model as a real-time intelligence layer, then hands off to a larger foundation model only when the conversation moves outside its knowledge base. On those handoffs, the agent briefly places the caller on hold to research the issue — the same beat a human agent would take.
The reason for the split is that LLM latency, which is tolerable in a chat window, breaks the illusion in a live phone call. Kamath frames the fix as matching the physiology of human conversation itself: partial listening, overlapping thought, interruption.
“The way an LLM works is you give it an entire prompt, and then it starts thinking”— Sudarshan Kamath, Founder and CEO of Smallest.ai
Key facts
- 01Smallest.ai raised $13M in Series A led by Seligman Ventures, bringing total funding above $21M.
- 02The startup was founded in late 2024 and builds small voice models rather than large foundation models.
- 03Existing customers include RingCentral and Truecaller; the model supports dozens of languages.
- 04Smallest.ai competes directly with ElevenLabs, Cartesia, and India-focused rival Sarvam.
- 05Architecture pairs a real-time small voice model with an offline LLM handoff for complex queries.
That framing drives what Smallest.ai chooses not to optimize for. Its models don't chase dubbing, podcasting, or general-purpose audio generation. They focus narrowly on real-time conversational agents that handle diverse accents, work across dozens of languages, and hold up in noisy environments — the specific failure modes that make current voice bots instantly identifiable.
The commercial thesis is that customer support companies — including well-funded incumbents like Sierra and Decagon — are potential customers rather than competitors, because building a world-class voice model pulls engineering focus away from the actual support product. Kamath said becoming extremely good at doing voice is a distraction from their core business. Smallest.ai wants to be the layer they license instead.
The competitive landscape is already crowded. ElevenLabs leads the category on brand and breadth, Cartesia is pushing on real-time synthesis, and Indian rival Sarvam is building for local-language coverage. Smallest.ai's wedge is enterprise-grade real-time conversation, not creator tools.
The Turing-test framing is aggressive, and Kamath owns it. He treats indistinguishability from a human agent not as a marketing line but as the product spec.
The counterweight is that voice indistinguishability is a moving target every major lab is also chasing, and the enterprise buyers Smallest.ai is targeting tend to consolidate on whichever provider clears procurement, security review, and integration first — not necessarily whichever model sounds most human in a demo. The startup will also need to prove that its two-model architecture, with its research pauses on hard queries, feels more natural in production than a single fast model that never breaks stride. And $21 million total funding is a modest war chest against ElevenLabs, which has raised in the hundreds of millions.
The bet worth watching is the architectural one. If Smallest.ai is right that voice agents converge on a small real-time model plus an offline LLM, that pattern reshapes the buy-versus-build calculus for every customer support platform now assembling its own voice stack, and it gives specialized voice vendors a durable seat under the agent layer rather than getting absorbed into it. If the market instead concludes that one large multimodal model can handle both roles well enough, the specialist tier compresses fast. The next 18 months of enterprise voice deployments will settle which side is right.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




