More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Fish Audio, a Palo Alto startup founded by ex-Nvidia researcher Shijia Liao, has raised a $52 million seed round led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist, Bayhouse, Carya, and HF0. The company has built a library of over 15,000 natural language controls across five models—four text-to-speech and one speech-to-text—and counts 8 million users on its open-source and hosted offerings. Last year’s GitHub release of Fish Speech drew more than 31,000 stars and attracted indie developers, game designers, and content creators; its latest S2.1 Pro voice model, however, is available exclusively via a paid API.
Fish Audio pulls in $21 million in annual recurring revenue from tiered subscription plans for individuals and teams, plus an enterprise API used by companies like HeyGen and Sanas. To train its models, the startup asks users to submit voice samples and compensates them when their recordings are used. After some creators flagged unauthorized uploads, Fish Audio automated its takedown process: anyone can file a claim with a voice sample or contract and have the matching voice removed in under three minutes. Coreline partner Osuke Honda says true community trust hinges on built-in consent, transparency, and fair licensing. Looking ahead, Fish Audio plans to launch an audio understanding model and a speech-to-speech system, aiming to stand out against rivals such as ElevenLabs, WellSaid, and Speechify by offering fine-grained controls and cost-efficient model training.
Questions about this article
No questions yet.