Back to News Hub
🟢TechCrunch AI
July 28, 2026
Funding & Investment

Fish Audio raises $50M seed to build AI voice models for creators and enterprises

Overview

Since launching last year, the startup today has more than 8 million people using the open-source or hosted version of its models, and now generates annual recurring revenue of $21 million. The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable.

Key Takeaways

  • Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls.

    Since launching last year, the startup today has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million.

  • The Fish Speech repository on GitHub now has more than 31,000 stars, and is used by indie developers, video game designers, and creators.

    The company has launched five models in the last year: four speech generation models and one speech-to-text model.

  • For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls," Cao said.

    One way the startup has built its library of voices is by simply asking users to submit their own voices for training its models, and compensating them if their voices are used.

  • Creators can easily submit a short voice sample or a contract to prove that an uploaded voice belongs to them, and their voice will be taken off the startup's platform in less than 3 minutes, she said.

    Still, that doesn't prevent anyone from uploading an artist's voice without their knowledge.

  • I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially," he said.

Stats & Key Facts

  • #Since launching last year, the startup today has more than 8 million people using the open-source or hosted version of its models, and now generates annual recurring revenue of $21 million.
  • #Since launching last year, the startup today has more than 8 million people using the open-source or hosted version of its models, and now generates annual recurring revenue of $21 million.
  • #Since launching last year, the startup today has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million.
  • #To continue building on that traction, the startup on Tuesday said it has raised $50 million in a seed round that was led by Coreline Ventures and Capital Today.

Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup today has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million. To continue building on that traction, the startup on Tuesday said it has raised $50 million in a seed round that was led by Coreline Ventures and Capital Today.

The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. Fish Audio started as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice generation model on a single GPU, which he open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars, and is used by indie developers, video game designers, and creators.

The company has launched five models in the last year: four speech generation models and one speech-to-text model. It has open-sourced three of its speech generation models, but its latest S2. 1 Pro model is available only through its paid API.

For more details please read the original article at TechCrunch AI.

Continue Learning

Originally published by TechCrunch AI
Read the original

Comments

Sign in to join the conversation