Skip to main content
Deepgram logo

Deepgram

Fast, scalable speech-to-text and voice AI APIs

audio#speech-to-text#voice-ai#transcription#api
Free plan Free trial Claimed API Self-hosted
Toolglade’s take

Deepgram is one of the strongest developer-first speech APIs on the market, with excellent per-minute economics, fast streaming, and a broad feature set including diarization and keyterm prompting. It is aimed squarely at developers, so non-technical users will find little to use directly. Pricing is transparent and usage-based, and the free $200 in credits makes it easy to evaluate. For production voice apps at scale, it is a top contender.

About Deepgram

Deepgram is a developer-focused voice AI platform providing low-latency speech-to-text, text-to-speech, and a Voice Agent API. Its Nova-3 model transcribes pre-recorded and streaming audio in 40+ languages with speaker diarization, smart formatting, automatic language detection, and keyterm prompting. Pricing is usage-based and among the most affordable in the category, with $200 in free starting credits and volume discounts via prepaid annual plans. Deepgram raised a $130M Series C at a $1.3B valuation in January 2026 and is best suited for teams building voice-enabled products at scale.

Deepgram builds voice AI infrastructure for developers and enterprises, centered on high-accuracy, low-cost speech-to-text. Its Nova model family (Nova-3 is the current flagship) transcribes pre-recorded and streaming audio in more than 40 languages, with features like speaker diarization, smart formatting, automatic language detection, and keyterm prompting to boost recognition of domain-specific vocabulary. Beyond transcription, Deepgram has expanded into text-to-speech and a Voice Agent API that stitches speech recognition, language models, and speech synthesis into real-time conversational applications. The platform is priced on usage, with per-minute rates that are among the lowest in the category, plus free credits to start and volume discounts through prepaid annual commitments. In January 2026 Deepgram raised a $130 million Series C at a $1.3 billion valuation, underscoring its position as a well-funded, developer-first alternative to hyperscaler speech services. It is a strong fit for teams building voice-enabled products at scale who want transcription accuracy, speed, and predictable cost.

TL;DR

Deepgram is a developer-first voice AI platform with low-latency speech-to-text, text-to-speech, and a Voice Agent API. Its Nova-3 model transcribes 40+ languages with diarization, smart formatting, and keyterm prompting at some of the lowest per-minute rates in the market. New accounts get about $200 in free credits, and enterprise volume discounts are available. It raised a $130M Series C at a $1.3B valuation in January 2026 and is ideal for teams building voice products at scale.

Company overview

Deepgram is a voice AI company building speech recognition, synthesis, and agent infrastructure for developers and enterprises. It positions itself as a fast, affordable, developer-first alternative to hyperscaler speech services.

The company has raised a total of roughly $229 million across multiple rounds, most recently a $130 million Series C at a $1.3 billion valuation announced in January 2026, led by AVP with participation from investors including Y Combinator, Madrona, Tiger, Wing, and BlackRock-managed funds.

Product features

Deepgram's core is the Nova family of speech-to-text models, supporting pre-recorded and streaming transcription in 40+ languages with speaker diarization, smart formatting, automatic language detection, and keyterm prompting.

The platform also offers text-to-speech and a Voice Agent API for building real-time conversational applications, plus self-hosted deployment for enterprises with strict data requirements.

Target market

Developers and enterprises building voice-enabled products, including call center analytics, media captioning, voice agents, and voice search, especially where high volume and low per-minute cost matter.

Buyer personas

End users

Developers and ML engineers integrating transcription or voice agents into applications.

Buyers

Engineering leaders and CTOs at companies building voice-enabled products or analytics.

Key influencers

Solutions architects and data engineers evaluating STT accuracy and cost.

Ideal customer profile

High-volume voice application teams that need accurate, low-latency, cost-efficient speech-to-text with enterprise deployment options.

Funding & performance

Deepgram has raised roughly $229 million total, most recently a $130 million Series C at a $1.3 billion valuation announced in January 2026, led by AVP.

Pros & cons

Pros

  • Very competitive per-minute pricing
  • Low-latency streaming and batch transcription
  • Nova-3 supports 40+ languages
  • Rich features: diarization, smart formatting, keyterm prompting
  • Generous $200 free starting credit
  • Self-hosted / on-prem deployment options
  • Well-funded with strong developer tooling and SDKs

Cons

  • Developer-only; no ready-to-use consumer app
  • Accuracy varies by language and audio quality
  • Advanced features and add-ons increase cost
  • Enterprise growth plans require annual commitments
  • No mobile app or browser extension
  • Requires engineering effort to integrate

Pricing plans

Free Credits
$0
  • About $200 in free credits
  • Access to Nova models
  • Pay-as-you-go after credits
Pay As You Go
~$0.0048/min
  • Nova-3 batch from ~$0.0048/min
  • Streaming from ~$0.0077/min
  • Diarization and formatting add-ons
  • No commitment
Growth / Enterprise
Custom (prepaid annual)
  • Discounted usage rates
  • Prepaid annual credits
  • Self-hosted / on-prem options
  • Priority support

Key features

API
Self-hosted
Multi-language
Integrations
REST API, WebSocket streaming, SDKs (Python, Node, Go, .NET), Voice Agent API
Input types
audio, text
Output types
text, audio
Best For
real-time transcription at scale, call center and voice analytics, building voice agents, cost-sensitive STT workloads

Compare key features

View all alternatives →
Feature
Deepgram
AssemblyAI
Speechmatics
Pricing
Freemium
Freemium
Freemium
Free plan
Yes
Yes
Yes
Free trial
Yes
Yes
Yes
API
Yes
Yes
Yes
Self-hosted
Yes
No
Yes

Frequently asked questions

What is Deepgram's flagship model?+

Nova-3 is the current flagship speech-to-text model, offering high accuracy across pre-recorded and streaming audio in 40+ languages with features like diarization and keyterm prompting.

How much does Deepgram cost?+

It is usage-based. As of August 2026, Nova-3 batch transcription is roughly $0.0048/minute and streaming around $0.0077/minute, with new accounts getting about $200 in free credits. Verify current rates with Deepgram.

Can Deepgram be self-hosted?+

Yes. Deepgram offers on-premises and self-hosted deployment options for enterprises with data residency or security requirements, in addition to its cloud API.

Does Deepgram do text-to-speech too?+

Yes. Beyond speech-to-text, Deepgram offers text-to-speech and a Voice Agent API that combines recognition, language models, and synthesis for real-time voice applications.

Is Deepgram suitable for non-developers?+

Not really. Deepgram is an API platform aimed at developers and engineering teams. If you need a finished transcription app, a consumer-facing tool would be a better fit.

Reviews (0)

Write a review

Pick a rating
Loading reviews…
Compare

Compare Deepgram with other AI tools

Side-by-side pages for pricing, features, and best-fit use cases.

All comparisons →

Similar tools you may like