The fastest, most accurate speech-to-text API for developers. Nova-2 sets the bar for real-time transcription.
Deepgram is what Twilio, Spotify, and NASA reach for when they need to transcribe audio at scale. The Nova-2 model delivers 90%+ accuracy on noisy real-world audio where Whisper and Google STT stumble โ and it streams results back in under 300ms, fast enough to power live voice agents.
Who it's for: Developers building voice agents, call center analytics, video captioning, podcast search, or any product that needs to ingest hours of audio.
90.4% WER on the Switchboard benchmark โ beats Whisper Large-v3 by 18 points on noisy audio. Trained on 1M+ hours of labeled speech.
WebSocket-based streaming API. Start receiving words as the speaker says them. Powers voice agents like Vapi, Retell, and Synthflow.
36 languages out of the box, plus code-switching detection for mixed-language audio. Custom models for jargon-heavy domains (medical, legal, finance).
Pay-as-you-go at $0.0043/min for Nova-2. Volume discounts to $0.0025/min at 1M+ minutes. $200 free credit on signup โ that's 46,500 minutes of transcription.
If you're building anything voice-first in 2026 โ voice agents, transcription, call analytics โ Deepgram is the default. It's the fastest, the most accurate on hard audio, and the developer experience is the best in class. The only reason to skip it: you need on-prem deployment, or your volume is so high that the per-minute math hurts (then negotiate or look at self-hosted Whisper.cpp).
OpenAI's open-source speech recognition. Free, runs locally, slower than Deepgram but no API cost.
Best-in-class AI voice generation. Pair with Deepgram for full speech-to-text-to-speech pipelines.
Meeting transcription and notes. Built on top of speech-to-text APIs like Deepgram.
Voice agent platform. Deepgram under the hood for STT, GPT-4 for reasoning, ElevenLabs for TTS.