Cartesia
Ultra-low-latency, real-time voice AI and text-to-speech built on state space models

Fast voice cloning and multilingual text-to-speech
Fish Audio combines expressive, natural TTS and quick voice cloning with an unusual advantage: open-source models you can self-host at no per-generation cost. That flexibility, plus competitive pricing, makes it appealing to both creators and developers. As a young company (founded 2025), it is less battle-tested than incumbents like ElevenLabs, and voice cloning raises the usual consent and misuse concerns, so use it responsibly and verify licensing for commercial work.
Fish Audio is a voice AI platform for expressive text-to-speech and fast voice cloning, able to clone a voice from as little as a 15-second sample and support zero-shot cloning across many languages via its S2 model. It offers dozens of emotion and tone tags and low-latency streaming for conversational AI. Uniquely, its core OpenAudio and Fish Speech models are open source, allowing self-hosted, no-per-generation-cost deployment alongside a hosted API and web app. Founded in 2025 and headquartered in Palo Alto, it raised a $52 million seed round in mid-2026.
Fish Audio is a fast-growing voice AI company that produces expressive text-to-speech and voice cloning. Its models can clone a voice from a sample as short as 15 seconds, capturing timbre, pacing, and style, and its S2 model supports zero-shot cloning across a large set of languages. A hallmark is granular expressiveness, with dozens of emotion and tone tags plus low-latency streaming aimed at conversational AI. What sets Fish Audio apart is that its core models, including Fish Speech and the OpenAudio S1 and S2 releases, are published as open-source repositories. That allows developers to self-host and run generation locally without per-use cost, while a hosted platform and API provide a managed option with a public voice library. The company reports strong benchmark rankings for naturalness and expressiveness. Founded in 2025 by former NVIDIA researcher Shijia Liao and headquartered in Palo Alto, Fish Audio raised a $52 million seed round in mid-2026 led by Coreline Ventures and Capital Today, and reports rapid growth to millions of users. It is a compelling choice for creators, developers, and enterprises wanting flexible, affordable, expressive voice generation.
Fish Audio is a voice AI platform for expressive TTS and fast voice cloning from as little as a 15-second sample, supporting many languages including zero-shot cloning via its S2 model. Its core OpenAudio and Fish Speech models are open source, so developers can self-host with no per-generation platform fee, alongside a hosted API and web app. Founded in 2025 in Palo Alto, it raised a $52 million seed round in mid-2026 and reports rapid user growth. It suits creators, developers, and enterprises wanting flexible, affordable, expressive voice.
Fish Audio is a voice AI company founded in 2025 by former NVIDIA researcher Shijia Liao and headquartered in Palo Alto. It grew out of the popular open-source Fish Speech project and has expanded into a hosted platform with the OpenAudio model family.
The company raised a $52 million seed round in mid-2026 led by Coreline Ventures and Capital Today, with participation from several venture firms and angels, and reports rapid growth to millions of users and strong annual recurring revenue for its stage.
Fish Audio provides expressive text-to-speech and voice cloning from short samples, with dozens of emotion and tone tags and low-latency streaming for conversational AI. Its S2 model supports zero-shot cloning across many languages.
The core models are open source, enabling self-hosted deployment without per-generation platform fees, while a hosted API, web app, and public voice library offer a managed path.
Content creators, developers, and enterprises needing expressive, multilingual voice generation and cloning, including those who want the option to self-host open-source models.
Creators, narrators, and developers generating or cloning voices for content and applications.
Product and content leaders selecting a TTS/voice-cloning provider.
Open-source and AI-voice communities that adopt and benchmark the models.
Creators and developer teams that want expressive, affordable, multilingual voice generation with the flexibility to self-host.
Fish Audio raised a $52 million seed round in mid-2026 led by Coreline Ventures and Capital Today, with participation from firms including 359 Capital, HF0, and 645 Ventures.
As little as a 15-second sample, from which it captures timbre, pacing, and speaking style. Its S2 model also supports zero-shot cloning across many languages.
Yes. Core models including Fish Speech and OpenAudio S1/S2 are published as open-source repositories on GitHub, allowing self-hosted, local generation without per-use platform fees.
It supports dozens of languages for TTS, with its S2 model enabling zero-shot voice cloning across a broad set of languages.
Yes. Fish Audio has a free tier, with paid subscriptions (Plus, Pro, Max) and an Enterprise option for higher volume and additional capabilities.
Commercial use depends on the plan and licensing terms, and voice cloning requires proper consent. Review Fish Audio's current terms and ensure you have rights to any voice you clone.
Side-by-side pages for pricing, features, and best-fit use cases.
Ultra-low-latency, real-time voice AI and text-to-speech built on state space models
Fast, scalable speech-to-text and voice AI APIs
Ultra low-latency text-to-speech for developers
Emotionally intelligent voice AI with an empathic interface and expression measurement, built for developers.