Deepgram offers Voice AI APIs for speech-to-text, text-to-speech, voice agents, and audio intelligence in real-time and batch, cloud and self-hosted configurations.
This is the broad commercial model reported for Deepgram. Product-level terms are listed below when available.
Trust profile
Deepgram has released updated Nova-3 monolingual models, expanding language coverage in speech-to-text. New batch models are available for Kannada and Telugu. New streaming models are available for Greek, Kannada, and Telugu. These models are live for all users and do not require changes to existing API requests.
The Flux speech-to-text model now supports numeral formatting, which converts spoken numbers into numerical format (e.g., "nine hundred" to "900"). This feature is available for the English Flux model and the Flux Multilingual model (covering English, Spanish, French, German, Russian, Portuguese, Italian, and Dutch). Users can enable this by setting the 'numerals=true' query parameter during the initial WebSocket connection.
Deepgram released improved Nova-3 monolingual models for eleven languages, including German, French, Turkish, and Ukrainian, among others. The updates, which apply to both batch and streaming workloads for specific languages, enhance transcription quality and accuracy for all existing users without requiring modifications to API requests.
Deepgram released the July 2026 version of its self-hosted suite (260714). Key features include improved number and percentage formatting, such as maintaining trailing zeros in decimals, and enhanced Malay language numeral conversion. Additionally, the update changes how streaming sessions report no-audio timeouts, providing a specific 'no_audio_timeout' reason instead of generic closing signals to facilitate better error handling for client applications.
Deepgram has updated its Voice Agent features, introducing an 'interrupt' behavior for injected messages that allows messages to immediately override ongoing speech. The company also moved its LatencyReport feature from experimental to fully supported, providing automated, granular reporting on LLM, TTS, and end-to-end latency metrics for each turn in a conversation.