The Palo Alto startup's open source voice models hit 8M users in a year. Some creators say their voices were uploaded to the platform without consent.
Fish Audio raised a $50M seed round led by Coreline Ventures and Capital Today, with 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0 participating, according to TechCrunch. Some creators say their voices were uploaded to the platform without consent, and the company leans on a DMCA take-down process to sort it out. The raise and the consent issue are both consequences of the same growth: an open-source voice model that reached 8M users in a year, and consent scaffolding that did not keep up.
Founder Shijia Liao trained his first voice generation model on a single GPU and open-sourced it. The repository, Fish Speech, now has more than 31,000 GitHub stars. Five models have shipped in the past year: four for speech generation and one for speech-to-text, with three of the speech models released as open weights. The company says its products reach 8M users and pull in $21M in annual recurring revenue, a standard SaaS metric that measures subscription income annualized, TechCrunch reported. That figure has not been independently audited.
The round is paying for that growth. Named enterprise customers include HeyGen, an AI avatar company, Sanas, a real-time voice translation service for call centers, Plaud, a hardware voice recorder, and LiveKit, a voice-agent platform that runs under live customer calls. Indie game designers and video creators use the open weights directly. The product positioning is a library of more than 15,000 natural language controls, the labels and pacing tags a user can stack on a voice to make it read as apologetic, urgent, or amused, that work for both a creator dubbing a character and a sales-team agent trying to sound less like a script.
That growth also made the consent problem visible. Some creators have alleged that their voices were uploaded to Fish Audio and used to train models or sold to enterprise customers without permission. The company's stated workflow, as described by spokesperson Cao, is to ask users to submit voices and to compensate them when those voices are used. When a creator objects, the company leans on a DMCA take-down process. A DMCA notice is a copyright tool, not a consent tool: it requires the person objecting to prove ownership of a specific recording, know to file, and have the time to do it. The available reporting on the consent dispute is truncated, and the number of take-downs filed, granted, or refused is not disclosed.
A $50M seed in 2026 shows that investors are betting voice is becoming infrastructure, in the same way text-to-image models became a layer underneath ad creative and e-commerce product photos. Fish Audio's open distribution is what got it onto indie developers' machines fast enough to win that bet. Open distribution is also what makes the consent problem hard. The voices that show up in the training mix are not always the voices the company knows about.
The company plans to use the funding to push S2.1 Pro, its first model shipped as a paid API only rather than open weights, and to expand enterprise sales. Both moves put Fish Audio closer to the end customer, and further from the open-source community that built the original dataset and tested the model in the wild.
The consent workflow has not scaled as fast as the model. A DMCA notice is a legal backstop, not a governance system. What would actually address the underlying complaint: a provenance layer, a signed manifest of which voices went into a given model, a creator opt-in registry searchable by voice, or a per-use royalty ledger. A clearer take-down form would not.
Fish Audio is the latest open-source AI startup to discover that seed-stage size is also a governance signal. The faster a model grows, the faster it outruns the scaffolding its community built around it. Whether the next round of voice-AI funding treats that as a problem to solve or a cost to externalize will shape who ends up running the voice layer of the AI stack.