Fish Audio raises $50M seed for AI voice models

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Fish Audio raised a $50 million seed round led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
- The company has reached 8 million users across its open-source and hosted voice models and now generates $21 million in annual recurring revenue despite only formally taking outside capital now.
- Founder Shijia Liao, a former NVIDIA researcher, started Fish Audio as a side project after training a voice generation model on a single GPU; the open-source Fish Speech repository has since accumulated more than 31,000 GitHub stars.
- Fish Audio shipped five models in the past year—four speech-generation and one speech-to-text—with its latest S2.1 Pro available only via paid API; enterprise customers include HeyGen, Sanas, Plaud, and LiveKit.
- The startup faced complaints that creators' voices were uploaded without consent; CEO Rissa Cao says the company has automated takedowns so voices can be removed in under 3 minutes, though uploads still aren't proactively blocked.
- Fish Audio plans to release an audio understanding model and a speech-to-speech model later this year as it competes against ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp.
Why it matters: Fish Audio built 8 million users and $21 million ARR before raising outside money, so this 'seed' round is effectively a growth-stage bet against well-capitalized incumbents like ElevenLabs. The unresolved consent gap—where any voice can be uploaded first and removed only after a complaint—remains a structural trust risk that Coreline partner Oskue Honda flagged as existential for community-driven voice platforms.




