Fish Audio Secures $50M to Advance AI Voice Models for Creators and Enterprises

Palo Alto-based startup Fish Audio has successfully raised $50 million in a seed funding round led by Coreline Ventures and Capital Today. Additional investors include 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. This substantial investment aims to enhance Fish Audio’s development of AI-generated voice models tailored for both creative professionals and enterprise applications.

Since its inception last year, Fish Audio has attracted over 8 million users to its open-source and hosted voice models, achieving an annual recurring revenue of $21 million. The company’s offerings include a library of more than 15,000 natural language controls, designed to provide expressive and steerable AI voices suitable for various use cases, from creative projects to customer support automation.

Fish Audio’s journey began as a personal project by former NVIDIA researcher Shijia Liao, who sought to create more expressive synthetic voices. His initial model, trained on a single GPU and open-sourced, has evolved into the Fish Speech repository on GitHub, now boasting over 31,000 stars and serving a diverse user base, including indie developers, video game designers, and content creators.

Over the past year, the company has launched five models: four for speech generation and one for speech-to-text. While three speech generation models are open-source, the latest S2.1 Pro model is accessible exclusively through a paid API. Fish Audio offers subscription plans for creators and teams, providing a set number of generation minutes and voice cloning features. Additionally, an enterprise version of its APIs and platform is available, with organizations like HeyGen, Sanas, and Plaud already utilizing these services.

To expand its voice library, Fish Audio has encouraged users to submit their own voices for model training, offering compensation for contributions. However, this approach led to concerns when some creators alleged their voices were uploaded without consent. In response, Fish Audio has automated its takedown process, allowing creators to submit a short voice sample or contract to verify ownership, ensuring removal from the platform within three minutes.

Oskue Honda, a partner at Coreline Ventures, emphasized the importance of trust in community-driven models, stating that consent, transparency, and attribution must be integral to the product. He advocates for the industry to adopt verified voice ownership, clear licensing terms, and efficient reporting and takedown processes.

Fish Audio’s rapid growth and significant funding highlight the increasing demand for advanced AI voice models. As the company continues to innovate, it faces the challenge of balancing technological advancement with ethical considerations, ensuring that creator rights and user trust remain at the forefront of its development strategy.