Fish Audio raises $52M seed for AI voice models

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Fish Audio raised a $52 million seed round led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
- Fish Audio claims more than 8 million users across its open-source and hosted products, $21 million in annual recurring revenue, and a library of over 15,000 natural language controls for steerability.
- Shijia Liao, a former Nvidia researcher frustrated by non-expressive synthetic voices, started Fish Audio by training a voice-generation model on a single GPU and open-sourcing it; the Fish Speech GitHub repo now has over 31,000 stars.
- Fish Audio launched five models in the past year — four speech-generation and one speech-to-text — with its latest S2.1 Pro model available only via paid API; enterprise customers include HeyGen, Sanas, and LiveKit.
- Fish Audio faced creator backlash after some users alleged their voices were uploaded without consent; CEO Rissa Cao told TechCrunch the startup has since automated its DMCA takedown process to under three minutes.
- Fish Audio plans to release an audio understanding model and a speech-to-speech model this year as it competes with ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp.
Why it matters: Fish Audio enters the crowded enterprise voice-AI market with $21M ARR and 8M users already in hand, but its community-submitted voice library has already triggered consent disputes — meaning the $52M must fund not just more advanced models but the consent, attribution, and takedown infrastructure that lead investor Coreline Ventures called essential to durability.




