The gold standard for speech-to-text. Open-source, 99 languages, near-human accuracy.
Whisper is the default speech-to-text model for 2026.
Who it's for: Podcasters, transcribers, video editors, developers.
Trained on 680,000 hours of multilingual audio. Out-of-the-box accuracy in English, Spanish, Mandarin, Hindi, Arabic, and 94 more.
Run Whisper locally on your own hardware. No per-minute API fees. No data leaves your machine.
Translate speech from any of 99 languages directly to English text. Built-in, no separate pipeline.
The latest model improves word error rate 10-20% over V2. Same 1.5B parameters, better training data.
Whisper is the default speech-to-text model for 2026. If you need a one-off transcription, the hosted API is cheap. If you transcribe 100+ hours/month, self-hosting on a single A10G GPU pays back in 2 months. For real-time voice agents, use the hosted API with streaming.