Fish Audio is moving from open-source traction into a broader commercial push for AI voice. The Palo Alto-based startup said on Tuesday that it raised $50 million in a seed round, giving it more room to build models for creators, developers and enterprises.
The company is entering a market where demand is coming from different directions at once. Creative users want synthetic voices that can carry emotion and character. Enterprises want voice systems that can be directed, controlled and used in areas such as customer support and sales operations.
A fast start for Fish Audio
Fish Audio says it now has more than 8 million people using either open-source or hosted versions of its models. The company also says it generates annual recurring revenue of $21 million.
The funding round was led by Coreline Ventures and Capital Today. It also included participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
The company began as a small project by former NVIDIA researcher Shijia Liao. Frustrated by synthetic voices that were not expressive enough, he trained a voice generation model on a single GPU and released it as open source.
That project became Fish Speech on GitHub, where the repository now has more than 31,000 stars. According to the source article, it is used by indie developers, video game designers, and creators.
How the product is expanding
Fish Audio has launched five models in the last year. Four are speech generation models, and one is a speech-to-text model.
The company has open-sourced three of its speech generation models. Its latest S2.1 Pro model, however, is available only through a paid API.
For creators and teams, Fish Audio sells monthly plans that include a set number of generation minutes and voice cloning features. For larger customers, it offers an enterprise version of its APIs and platform.
The company says organizations including HeyGen, Sanas and Plaud are already using it. These customers point to the broader range of expectations around AI voice: some need realism for AI avatars, others need expressive character voices, and voice agent companies need natural-sounding, low-latency output for calls.
Why control matters in AI voice models
Fish Audio is betting partly on control. The company says it has a library of more than 15,000 natural language controls, which is meant to help users shape the way a generated voice sounds and behaves.
That matters because AI voice is not a single use case. A creator may want a voice that sounds dramatic, subtle or character-driven. A company building customer-facing voice systems may care more about predictability, speed and the ability to guide output for a specific workflow.
The same model family may therefore need to serve very different buyers. Fish Audio’s commercial challenge is to make those controls useful enough for technical teams, while still accessible enough for creators who want fast voice generation and cloning tools.
The consent issue is still central
Fish Audio has built part of its voice library by asking users to submit their own voices for training, with compensation when those voices are used. That approach has also created tension.
A few months ago, some creators alleged that their voices had been uploaded to Fish Audio without their consent. The company already had a DMCA content take-down process, but the source article says the take-downs took a long time.
CEO and co-founder Rissa Cao told TechCrunch that Fish Audio has now automated that process. She said creators can submit a short voice sample or a contract to prove that an uploaded voice belongs to them, and the voice will be removed from the platform in less than 3 minutes.
That fix addresses speed after a complaint is filed. It does not, however, stop someone from uploading an artist’s voice without that artist knowing. Until the artist discovers the upload and requests removal, the voice can continue to be used on the platform.
For a community-driven AI voice platform, that distinction is important. Fast takedowns may reduce harm after discovery, but trust also depends on consent, transparency and attribution being clear before voices are used.
What comes next
Fish Audio plans to release an audio understanding model this year. It is also building a speech-to-speech model.
The company is moving forward in a crowded speech generation market. Competitors named in the source article include ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp.
Investors are framing Fish Audio’s advantage around technical efficiency and fine-grained developer controls. Rico Mallozzi, a partner at 359 Capital, said the company’s work shows technical ability in narrowing the gap between artificial-sounding and human-like voices.
The next phase will test whether Fish Audio can turn open-source attention, paid API access and enterprise demand into a durable platform. Its growth numbers are already substantial, but in AI voice, product quality and creator trust are likely to remain tightly linked.