Skip to main content
Live
Main content

Fish Audio raises $52M seed at $21M ARR for AI voice models

The Palo Alto startup has 8 million users and 15,000 natural-language voice controls, with three of its five models open-sourced.

Jaeden Schafer
Editor in Chief · · 5 min read
Fish Audio raises $52M seed at $21M ARR for AI voice models

Fish Audio, a Palo Alto startup building voice generation models, raised $52 million in a seed round led by Coreline Ventures and Capital Today, the company said Tuesday. The round is unusually large for a seed stage and lands as Fish Audio reports $21 million in annual recurring revenue and more than 8 million users across its open-source and hosted products.

359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0 also participated. Fish Audio's pitch to the market rests on two numbers: a library of more than 15,000 natural-language voice controls, and a Fish Speech GitHub repository with over 31,000 stars, which the company uses as evidence of developer adoption beyond its paid tier.

The company shipped five models in the last year — four speech generation models and one speech-to-text model — and open-sourced three of the speech generation models. Its newest release, the S2.1 Pro model, is available only through the paid API, marking Fish Audio's shift toward monetizing frontier capabilities while keeping earlier releases open.

Key facts

  • 01Fish Audio raised $52M in a seed round led by Coreline Ventures and Capital Today.
  • 02The Palo Alto startup reports $21M in annual recurring revenue and more than 8M users across open-source and hosted models.
  • 03Its Fish Speech GitHub repository has surpassed 31,000 stars since launching last year.
  • 04The library offers 15,000+ natural language voice controls; enterprise customers include HeyGen and Sanas.
  • 05Fish Audio automated its DMCA takedown flow to under three minutes after creators alleged voices were uploaded without consent.

Fish Audio started as a side project by former NVIDIA researcher Shijia Liao, who trained a voice generation model on a single GPU after finding existing synthetic voices flat and non-expressive. He open-sourced the result, and the repository became a magnet for indie developers, video game designers, and creators — the base on which the company later layered paid tiers for individuals, teams, and enterprises.

Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters
Rissa Cao, CEO and co-founder, Fish Audio

CEO and co-founder Rissa Cao said the enterprise pull is heterogeneous. HeyGen uses Fish Audio voices for AI avatars and prioritizes realism; a gaming studio wants expressive voices for characters; LiveKit-style voice agent companies want natural sound and low latency for phone calls. The 15,000-control library is the mechanism for serving those different requirements from one platform. Sanas is also cited as an enterprise customer.

One of Fish Audio's growth levers has been asking users to submit their own voices for model training and compensating them when those voices are used. That community model ran into trouble a few months ago, when some creators alleged their voices had been uploaded without consent. The company had a DMCA takedown process, but response times were slow.

Cao said the takedown flow is now automated: a creator can submit a short voice sample or a contract proving ownership, and the voice is removed in under three minutes. The gap in the system remains structural — nothing stops someone from uploading an artist's voice in the first place, and enforcement is reactive until the artist notices. Osuke Honda, a partner at Coreline Ventures, said durable community platforms require consent, transparency, and attribution built into the product, and floated verified voice ownership and revenue sharing as the direction the industry needs to move.

Cao said Fish Audio was running efficiently as an open-source project with creator plans and did not need outside capital, but chose to raise to build more advanced models and serve enterprise demand that was already pulling on the team. Planned releases this year include an audio understanding model and a speech-to-speech model, which would move Fish Audio beyond one-way text-to-speech into the conversational agent stack.

Related · from this week
Smallest.ai raises $13M to build small voice models that mimic human turn-taking
Jaeden Schafer · 4 min read →

The competitive field is packed. ElevenLabs, WellSaid, Cartesia, Speechify, Async — formerly Podcastle — and Krisp are all pursuing creator and enterprise budgets, with ElevenLabs commanding the highest brand recognition. Rico Mallozzi, a partner at 359 Capital, argued Fish Audio's edge is fine-grained developer controls and cost-efficient training, and pointed to the team's ability to ship state-of-the-art models on far smaller budgets than the incumbents.

The consent controversy is the load-bearing risk. A community-sourced voice library scales quickly and cheaply, but if creators cannot trust that uploads are policed at ingestion rather than at complaint, the model becomes both a legal exposure and a reputational one. A three-minute takedown is fast; a three-minute takedown that only fires after discovery still leaves a window.

Voice is turning into a category where the winner is not the lab with the most parameters but the one with the deepest catalog of licensed, controllable voices and the cleanest enterprise integrations. Fish Audio's $21M ARR at seed is the number that matters — it signals the company already sells, and the $52M is fuel to widen the moat before ElevenLabs and Cartesia push into the same creator-plus-enterprise middle. Whether Fish Audio wins depends less on model quality, which its investors say is already competitive, and more on whether it can turn a community model into a licensing regime creators actually endorse.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Business

Smallest.ai raises $13M to build small voice models that mimic human turn-taking
Business

Smallest.ai raises $13M to build small voice models that mimic human turn-taking

The India-founded startup is betting specialized voice models beat LLMs at real-time conversation, with RingCentral and Truecaller already on board.

Jaeden Schafer4 min read
Spotify launches ElevenLabs-powered audiobook creation tool for self-publishers
Business

Spotify launches ElevenLabs-powered audiobook creation tool for self-publishers

The streaming platform will let authors generate audiobooks in-app starting June 2026, with no exclusivity required.

Jaeden Schafer5 min read
Wispr Flow's India growth hits 100% as Hinglish voice push lands
Business

Wispr Flow's India growth hits 100% as Hinglish voice push lands

The Bay Area voice-input startup says India is now its second-largest market, with 14% of 2.5M global downloads and aggressive local pricing.

Jaeden Schafer5 min read