Imagine: a user speaks into the microphone, and text appears on screen with a delay of less than half a second. This is a non-trivial engineering task, but with Deepgram Nova-2 it is solved. The model gives a median latency of about 300 ms and a WER of about 5% on Russian. Such indicators are unavailable to classic batch solutions. In this article, we'll cover how to set up WebSocket transcription on iOS and android, which parameters are critical, and how to avoid typical mistakes.
Typical difficulties include unstable connections, duplication of interim results, and incorrect audio codec configuration. Our team, with 10+ years of experience in ASR implementation, guarantees a stable connection and low latency. On a real project, a client complained about interim duplication: each new word was appended to the previous one. After implementing the buffer replacement pattern, the problem disappeared completely.
Comparison of Deepgram Nova-2 and Whisper
Deepgram Nova-2 provides low latency on streaming: median of about 300 ms from the end of a phrase to text. Whisper cannot do that in principle – it is synchronous. If the task is "user speaks – text appears on screen" with sub-second delay, it's Deepgram.
| Characteristic | Deepgram Nova-2 | Whisper (synchronous) |
|---|---|---|
| Final latency | 300–500 ms | 2–5 seconds |
| Streaming | Asynchronous, streaming | Synchronous, batch |
| Interim results | Yes | No |
| Russian support | Excellent | Good |
| Price (per hour) | Upon request | Free (self-hosted) |
For a mobile scenario, Deepgram wins by 6–10 times in speed. Additionally, Nova-2 achieves WER ≤5% on Russian, while Whisper large-v2 is around 7%.
Configuring the connection protocol
Deepgram works via WebSocket. Endpoint:
wss://api.deepgram.com/v1/listen?model=nova-2&language=ru&encoding=linear16&sample_rate=16000&channels=1&interim_results=true
Parameters are critical: encoding=linear16 means raw PCM 16-bit little-endian. Any other format without explicit codec specification risks a 1008 Policy Violation. interim_results=true enables partial results – they create the real-time feel.
iOS: AVAudioEngine + URLSessionWebSocketTask
class DeepgramStreamer {
private var audioEngine = AVAudioEngine()
private var webSocket: URLSessionWebSocketTask?
func start() throws {
let session = URLSession(configuration: .default)
var request = URLRequest(url: URL(string: "wss://api.deepgram.com/v1/listen?model=nova-2&language=ru&encoding=linear16&sample_rate=16000&channels=1&interim_results=true")!)
request.setValue("Token \(apiKey)", forHTTPHeaderField: "Authorization")
webSocket = session.webSocketTask(with: request)
webSocket?.resume()
receiveLoop()
let inputNode = audioEngine.inputNode
let format = AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 16000, channels: 1, interleaved: false)!
inputNode.installTap(onBus: 0, bufferSize: 4096, format: format) { buffer, _ in
guard let channelData = buffer.int16ChannelData else { return }
let frameLength = Int(buffer.frameLength)
let data = Data(bytes: channelData[0], count: frameLength * 2)
self.webSocket?.send(.data(data)) { _ in }
}
try audioEngine.start()
}
private func receiveLoop() {
webSocket?.receive { [weak self] result in
if case .success(let message) = result, case .string(let text) = message {
// Decode Deepgram JSON response
self?.handleTranscript(text)
}
self?.receiveLoop()
}
}
}
Important detail: AVAudioEngine.inputNode on iOS 16+ requires explicit microphone permission via AVAudioSession.sharedInstance().requestRecordPermission. And обязательно AVAudioSession.setCategory(.record, mode: .measurement) – the .measurement mode disables AEC and AGC, which can distort the signal for transcription.
Android: AudioRecord + OkHttp WebSocket
class DeepgramStreamer(private val apiKey: String) {
private val client = OkHttpClient()
private var webSocket: WebSocket? = null
private var audioRecord: AudioRecord? = null
fun start(onTranscript: (String, Boolean) -> Unit) {
val request = Request.Builder()
.url("wss://api.deepgram.com/v1/listen?model=nova-2&language=ru&encoding=linear16&sample_rate=16000&channels=1&interim_results=true")
.header("Authorization", "Token $apiKey")
.build()
webSocket = client.newWebSocket(request, object : WebSocketListener() {
override fun onMessage(webSocket: WebSocket, text: String) {
val json = JSONObject(text)
val channel = json.getJSONObject("channel")
val alternatives = channel.getJSONArray("alternatives")
val transcript = alternatives.getJSONObject(0).getString("transcript")
val isFinal = json.getBoolean("is_final")
if (transcript.isNotEmpty()) onTranscript(transcript, isFinal)
}
})
val bufferSize = AudioRecord.getMinBufferSize(16000, AudioFormat.CHANNEL_IN_MONO, AudioFormat.ENCODING_PCM_16BIT)
audioRecord = AudioRecord(MediaRecorder.AudioSource.MIC, 16000, AudioFormat.CHANNEL_IN_MONO, AudioFormat.ENCODING_PCM_16BIT, bufferSize)
audioRecord?.startRecording()
Thread {
val buffer = ShortArray(bufferSize / 2)
while (audioRecord?.recordingState == AudioRecord.RECORDSTATE_RECORDING) {
val read = audioRecord!!.read(buffer, 0, buffer.size)
if (read > 0) {
val byteBuffer = ByteBuffer.allocate(read * 2).order(ByteOrder.LITTLE_ENDIAN)
buffer.take(read).forEach { byteBuffer.putShort(it) }
webSocket?.send(byteBuffer.array().toByteString())
}
}
}.start()
}
}
ByteOrder.LITTLE_ENDIAN is mandatory. Deepgram expects LE PCM. Sending BE will work but with noticeably worse quality.
How to avoid typical streaming audio mistakes?
-
Interim duplication: never accumulate all interim as separate lines. Store the current utterance in a buffer and overwrite it with each new interim. When
is_final: truearrives, finalize the buffer. - Connection loss: implement reconnect with exponential backoff (1,2,4,8 sec). Deepgram does not support session resumption, so after reconnect you must start a new stream.
- Incorrect sample rate: strictly use 16000 Hz. Higher rates increase traffic without quality gain, lower rates degrade recognition.
Handling interim results
Deepgram returns two types of messages: with is_final: false (interim) and is_final: true (final). Correct UI pattern:
- Display interim in gray or italics – the user sees recognition in progress
- When
is_final: trueis received, replace all previous interim of that utterance with the final text -
speech_final: trueindicates the end of a pause – a good moment to start processing the phrase
Nova-2 parameters that affect quality
| Parameter | Value | Description |
|---|---|---|
| model | nova-2 | Recognition model |
| encoding | linear16 | Audio encoding |
| sample_rate | 16000 | Sample rate |
| interim_results | true | Enable partial results |
| utterance_end_ms | 1000 | Finalization on pause |
| diarize | false | Speaker separation |
| punctuate | true | Auto-punctuation |
| smart_format | true | Formatting numbers and dates |
-
utterance_end_ms: 1000– Deepgram automatically finalizes utterance after 1 second of silence. Useful for dictation without explicit "stop" commands. -
diarize: true– speaker separation, addsspeakerto each word. -
punctuate: true– auto-punctuation. Without it, text lacks periods and commas. -
smart_format: true– formats numbers, dates, phones. "twenty-fifth of March" → "25 March".
What's included in the work
- Setting up WebSocket connection with authorization
- Audio capture via AVAudioEngine / AudioRecord in correct format
- Processing interim and final results without duplication
- Reconnect on network drop with exponential backoff
- Testing on real devices (iOS/Android)
- Documentation and team training
Work process
- Analysis – we study the application architecture and transcription requirements
- Design – we choose the model, protocol, and parameters
- Implementation – code integration of WebSocket, audio capture, UI
- Testing – load testing, verification on different devices and networks
- Deployment – publishing in App Store and Google Play
Timeline
Basic integration of WebSocket + AudioRecord/AVAudioEngine + text output – 4–7 days. Adding diarization, network switching handling (reconnect), background mode, and result export – 8–14 days.
Savings on developing your own ASR engine can reach 70%. Transcription cost is low – significantly cheaper than alternatives with batch processing.
Get a consultation on integrating Deepgram for your mobile project. We will analyze the architecture, choose the optimal configuration, and implement low-latency transcription turnkey. Order a demo version with low latency today. Contact us to discuss your project.
Integrating Deepgram allowed us to reduce transcription latency from 5 seconds to 300 ms – client feedback from the financial sector.







