Skip to main content
Fish Audio logo

Fish Audio

Fast voice cloning and multilingual text-to-speech

audio#text-to-speech#voice-cloning#open-source#tts
Free plan Claimed API Self-hosted
Toolglade’s take

Fish Audio combines expressive, natural TTS and quick voice cloning with an unusual advantage: open-source models you can self-host at no per-generation cost. That flexibility, plus competitive pricing, makes it appealing to both creators and developers. As a young company (founded 2025), it is less battle-tested than incumbents like ElevenLabs, and voice cloning raises the usual consent and misuse concerns, so use it responsibly and verify licensing for commercial work.

About Fish Audio

Fish Audio is a voice AI platform for expressive text-to-speech and fast voice cloning, able to clone a voice from as little as a 15-second sample and support zero-shot cloning across many languages via its S2 model. It offers dozens of emotion and tone tags and low-latency streaming for conversational AI. Uniquely, its core OpenAudio and Fish Speech models are open source, allowing self-hosted, no-per-generation-cost deployment alongside a hosted API and web app. Founded in 2025 and headquartered in Palo Alto, it raised a $52 million seed round in mid-2026.

Fish Audio is a fast-growing voice AI company that produces expressive text-to-speech and voice cloning. Its models can clone a voice from a sample as short as 15 seconds, capturing timbre, pacing, and style, and its S2 model supports zero-shot cloning across a large set of languages. A hallmark is granular expressiveness, with dozens of emotion and tone tags plus low-latency streaming aimed at conversational AI. What sets Fish Audio apart is that its core models, including Fish Speech and the OpenAudio S1 and S2 releases, are published as open-source repositories. That allows developers to self-host and run generation locally without per-use cost, while a hosted platform and API provide a managed option with a public voice library. The company reports strong benchmark rankings for naturalness and expressiveness. Founded in 2025 by former NVIDIA researcher Shijia Liao and headquartered in Palo Alto, Fish Audio raised a $52 million seed round in mid-2026 led by Coreline Ventures and Capital Today, and reports rapid growth to millions of users. It is a compelling choice for creators, developers, and enterprises wanting flexible, affordable, expressive voice generation.

TL;DR

Fish Audio is a voice AI platform for expressive TTS and fast voice cloning from as little as a 15-second sample, supporting many languages including zero-shot cloning via its S2 model. Its core OpenAudio and Fish Speech models are open source, so developers can self-host with no per-generation platform fee, alongside a hosted API and web app. Founded in 2025 in Palo Alto, it raised a $52 million seed round in mid-2026 and reports rapid user growth. It suits creators, developers, and enterprises wanting flexible, affordable, expressive voice.

Company overview

Fish Audio is a voice AI company founded in 2025 by former NVIDIA researcher Shijia Liao and headquartered in Palo Alto. It grew out of the popular open-source Fish Speech project and has expanded into a hosted platform with the OpenAudio model family.

The company raised a $52 million seed round in mid-2026 led by Coreline Ventures and Capital Today, with participation from several venture firms and angels, and reports rapid growth to millions of users and strong annual recurring revenue for its stage.

Product features

Fish Audio provides expressive text-to-speech and voice cloning from short samples, with dozens of emotion and tone tags and low-latency streaming for conversational AI. Its S2 model supports zero-shot cloning across many languages.

The core models are open source, enabling self-hosted deployment without per-generation platform fees, while a hosted API, web app, and public voice library offer a managed path.

Target market

Content creators, developers, and enterprises needing expressive, multilingual voice generation and cloning, including those who want the option to self-host open-source models.

Buyer personas

End users

Creators, narrators, and developers generating or cloning voices for content and applications.

Buyers

Product and content leaders selecting a TTS/voice-cloning provider.

Key influencers

Open-source and AI-voice communities that adopt and benchmark the models.

Ideal customer profile

Creators and developer teams that want expressive, affordable, multilingual voice generation with the flexibility to self-host.

Funding & performance

Fish Audio raised a $52 million seed round in mid-2026 led by Coreline Ventures and Capital Today, with participation from firms including 359 Capital, HF0, and 645 Ventures.

Pros & cons

Pros

  • Voice cloning from a 15-second sample
  • Expressive output with dozens of emotion/tone tags
  • Open-source models enable free self-hosting
  • Broad multilingual and zero-shot cloning support
  • Low-latency streaming for conversational use
  • Strong benchmark rankings for naturalness
  • Competitive subscription pricing plus a free tier

Cons

  • Young company with a shorter track record
  • Voice cloning raises consent and misuse concerns
  • Self-hosting requires technical setup and GPUs
  • Commercial licensing terms need careful checking
  • No mobile app or browser extension
  • Smaller ecosystem than ElevenLabs

Pricing plans

Free
$0 / month
  • Access to voice library
  • Basic TTS and cloning
  • Usage limits apply
Plus
$11 / month
  • More generation quota
  • Voice cloning
  • Higher-quality models
Pro
$75 / month
  • High-volume generation
  • Priority access
  • Advanced features
Max
$749 / month
  • Very high volume
  • Priority support
  • For heavy production use
Enterprise
Custom
  • Custom volume pricing
  • Dedicated support
  • Commercial terms

Key features

API
Self-hosted
Multi-language
Integrations
REST API, Open-source model weights (GitHub), Web app, Streaming API
Input types
text, audio
Output types
audio
Best For
voice cloning from short samples, multilingual TTS, self-hosted voice generation, conversational AI voices

Compare key features

View all alternatives →
Feature
Fish Audio
Cartesia
Deepgram
Pricing
Freemium
Freemium
Freemium
Free plan
Yes
Yes
Yes
Free trial
No
Yes
Yes
API
Yes
Yes
Yes
Self-hosted
Yes
No
Yes
Team support
No
Yes
No

Frequently asked questions

How much audio does Fish Audio need to clone a voice?+

As little as a 15-second sample, from which it captures timbre, pacing, and speaking style. Its S2 model also supports zero-shot cloning across many languages.

Are Fish Audio's models open source?+

Yes. Core models including Fish Speech and OpenAudio S1/S2 are published as open-source repositories on GitHub, allowing self-hosted, local generation without per-use platform fees.

How many languages does Fish Audio support?+

It supports dozens of languages for TTS, with its S2 model enabling zero-shot voice cloning across a broad set of languages.

Is there a free plan?+

Yes. Fish Audio has a free tier, with paid subscriptions (Plus, Pro, Max) and an Enterprise option for higher volume and additional capabilities.

Can I use cloned voices commercially?+

Commercial use depends on the plan and licensing terms, and voice cloning requires proper consent. Review Fish Audio's current terms and ensure you have rights to any voice you clone.

Reviews (0)

Write a review

Pick a rating
Loading reviews…
Compare

Compare Fish Audio with other AI tools

Side-by-side pages for pricing, features, and best-fit use cases.

All comparisons →

Similar tools you may like