Skip to main content
Course

Voice & Conversational AI

Build voice agents that feel natural — the STT→LLM→TTS pipeline vs. speech-native models, latency and turn-taking, and the cloning-fraud and disclosure law you can't ignore.

Module 1 freeIntermediate ~5h
Start free — How voice AI works: the pipeline

Voice is becoming a primary interface for AI — customer service, assistants, and agents that talk. But building voice AI that feels natural is hard: latency, turn-taking, interruption, and a cascading error stack where one misheard word derails the whole conversation. This course teaches developers and product teams the voice-AI pipeline, the technology components (STT, TTS, speech-to-speech), conversation design for voice, and the real ethical and legal issues — voice-cloning fraud and disclosure law.

It is practical and current for 2026, and honest about what's hard (latency isn't as low as the marketing suggests). Specific products and latency numbers change fast, so the course emphasizes durable concepts.

What you'll be able to do

  • Understand the voice-AI pipeline (STT, LLM, TTS) and speech-native models
  • Design for latency, turn-taking, and interruption
  • Handle the cascading error stack and build reliable voice agents
  • Navigate voice-cloning fraud, consent, and disclosure law (FCC/TCPA)
Voice AIConversational designSpeech-to-text / TTSVoice agentsAI disclosure compliance

Curriculum

Module 1: Foundations

Free

How voice AI works, cascade vs. speech-native, the honest truth about latency, and turn-taking.

Module 2: The Technology Components

Premium

Speech-to-text and Whisper's hallucination issue, TTS and voice cloning, speech-to-speech, and orchestration frameworks.

  • Speech-to-text and the Whisper hallucination problem15 min
  • Text-to-speech and voice cloning12 min
  • Realtime speech-to-speech models12 min
  • Orchestration frameworks12 min

Module 3: Building Good Voice Agents

Premium

Conversation design for voice, the cascading error stack, telephony, and grounding/guardrails.

  • Conversation design for voice12 min
  • The cascading error stack12 min
  • Telephony integration12 min
  • Grounding and guardrails for voice12 min

Module 4: Risks, Ethics & Law

Premium

Voice-cloning fraud, the FCC/TCPA ruling and disclosure law, accessibility, and the road ahead.

  • Voice cloning and deepfake fraud12 min
  • The law: FCC/TCPA, disclosure, and voice rights15 min
  • Accessibility: benefits and considerations12 min
  • The road ahead12 min

Get the full course — one-time $29.

Start Module 1 free today. Buy once for lifetime access to the remaining 3 modules — no subscription.

One-time payment · lifetime access
Hand-written & fact-checked Certificate on completion Module 1 free

Want all 53? Get the All-Access Bundle for $99 — one purchase, every course.

Student reviews

Reviews come from students who own this course. Enroll to share yours.
Loading reviews…