Deepgram
Fast, scalable speech-to-text and voice AI APIs

Ultra low-latency text-to-speech for developers
LMNT's calling card is latency: if you're building a voice agent or interactive app where responsiveness is critical, its 150-200 ms streaming is a genuine advantage. It is developer-focused, so it won't replace a full creator suite, and its voice catalog and language coverage are narrower than the largest TTS players. Public funding data is thin and inconsistent, so evaluate the company on product fit rather than headline numbers.
LMNT is a developer-focused text-to-speech API built for ultra low-latency, real-time streaming voice, with reported latency of around 150-200 ms. It supports fast voice cloning from a short recording and offers polished voices across roughly 24 languages. Designed to scale, with concurrency/rate-limit removal for enterprise, it is aimed at conversational AI, voice agents, and interactive applications. Pricing is freemium and usage-based, from a free tier and an affordable entry plan to per-character overage and custom enterprise rates.
LMNT provides a text-to-speech API engineered for speed, delivering lifelike streaming audio with latency reported in the 150-200 millisecond range. That responsiveness makes it well-suited to conversational AI, voice agents, and interactive experiences where the gap between text and audio must feel instantaneous. It also supports fast voice cloning from a short recording and offers a catalog of polished, natural-sounding voices across roughly two dozen languages. The product is aimed squarely at developers building real-time voice features, with an API designed to scale and, for enterprise customers, options that remove concurrency and rate limits. Pricing follows a freemium, usage-based model with a free tier, an affordable entry plan, per-character overage rates, and custom enterprise pricing. Founded in the late 2010s and based in San Francisco, LMNT is a smaller, focused player in the TTS space rather than a broad consumer brand. Public funding details are limited and somewhat inconsistent across sources, so treat totals cautiously. LMNT is best for teams that specifically need low-latency streaming voice and are comfortable integrating an API.
LMNT is a developer-focused text-to-speech API optimized for ultra low-latency (~150-200 ms) real-time streaming voice, with fast voice cloning and roughly 24 languages. It targets conversational AI, voice agents, and interactive apps where responsiveness is critical. Pricing is freemium and usage-based, from a free tier and ~$10/month entry plan to per-character overage and custom enterprise rates. It is a smaller, focused player; public funding data is limited and inconsistent.
LMNT is a San Francisco-based text-to-speech company founded in the late 2010s, focused on lifelike, low-latency streaming voice for developers. It positions itself around speed and real-time responsiveness rather than a broad consumer brand.
Public funding information is limited and inconsistent across sources (figures reported range widely), so specific totals should be treated cautiously and verified directly with the company.
LMNT's core is a streaming text-to-speech API with reported latency around 150-200 ms, fast voice cloning from short recordings, and a catalog of polished voices across roughly 24 languages.
The API is designed to scale, and enterprise options can remove concurrency and rate limits for high-throughput real-time applications.
Developers and teams building real-time voice experiences such as conversational agents, interactive games, and multilingual assistants where low latency is essential.
Developers integrating real-time speech into agents, apps, and games.
Engineering and product leaders selecting a low-latency TTS provider.
Voice-AI engineers benchmarking latency and naturalness.
Teams building latency-sensitive, real-time voice applications who value speed and a scalable developer API.
Public funding details for LMNT are limited and inconsistent across sources; no reliable, clearly disclosed total is available. Verify with the company.
LMNT reports streaming latency of roughly 150-200 milliseconds, making it well-suited for real-time voice agents and interactive applications.
Yes. LMNT offers fast voice cloning from a short recording, in addition to its catalog of prebuilt voices.
LMNT supports roughly 24 languages, which is narrower than the largest multilingual TTS providers but sufficient for many applications.
Not really. LMNT is an API aimed at developers building real-time voice features. Non-technical users seeking a full creator suite should consider other tools.
It is freemium and usage-based, with a free tier, an entry plan around $10/month, per-character overage on higher tiers, and custom enterprise pricing. Verify current rates with LMNT.
Side-by-side pages for pricing, features, and best-fit use cases.
Fast, scalable speech-to-text and voice AI APIs
Fast voice cloning and multilingual text-to-speech
Ultra-low-latency, real-time voice AI and text-to-speech built on state space models
Speech AI models and APIs for developers