Mobile App Transcription with Deepgram Nova-2

Imagine: a user speaks into the microphone, and text appears on screen with a delay of less than half a second. This is a non-trivial engineering task, but with **Deepgram Nova-2** it is solved. The model gives a median latency of about 300 ms and a WER of about 5% on Russian. Such indicators are un

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Mobile App Transcription with Deepgram Nova-2
Medium
~3-5 days

Our competencies:

Frequently Asked Questions

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    896
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    782
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1216
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1079
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    1003
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    597

Imagine: a user speaks into the microphone, and text appears on screen with a delay of less than half a second. This is a non-trivial engineering task, but with Deepgram Nova-2 it is solved. The model gives a median latency of about 300 ms and a WER of about 5% on Russian. Such indicators are unavailable to classic batch solutions. In this article, we'll cover how to set up WebSocket transcription on iOS and android, which parameters are critical, and how to avoid typical mistakes.

Typical difficulties include unstable connections, duplication of interim results, and incorrect audio codec configuration. Our team, with 10+ years of experience in ASR implementation, guarantees a stable connection and low latency. On a real project, a client complained about interim duplication: each new word was appended to the previous one. After implementing the buffer replacement pattern, the problem disappeared completely.

Comparison of Deepgram Nova-2 and Whisper

Deepgram Nova-2 provides low latency on streaming: median of about 300 ms from the end of a phrase to text. Whisper cannot do that in principle – it is synchronous. If the task is "user speaks – text appears on screen" with sub-second delay, it's Deepgram.

Characteristic Deepgram Nova-2 Whisper (synchronous)
Final latency 300–500 ms 2–5 seconds
Streaming Asynchronous, streaming Synchronous, batch
Interim results Yes No
Russian support Excellent Good
Price (per hour) Upon request Free (self-hosted)

For a mobile scenario, Deepgram wins by 6–10 times in speed. Additionally, Nova-2 achieves WER ≤5% on Russian, while Whisper large-v2 is around 7%.

Configuring the connection protocol

Deepgram works via WebSocket. Endpoint:

wss://api.deepgram.com/v1/listen?model=nova-2&language=ru&encoding=linear16&sample_rate=16000&channels=1&interim_results=true 

Parameters are critical: encoding=linear16 means raw PCM 16-bit little-endian. Any other format without explicit codec specification risks a 1008 Policy Violation. interim_results=true enables partial results – they create the real-time feel.

iOS: AVAudioEngine + URLSessionWebSocketTask

class DeepgramStreamer { private var audioEngine = AVAudioEngine() private var webSocket: URLSessionWebSocketTask? func start() throws { let session = URLSession(configuration: .default) var request = URLRequest(url: URL(string: "wss://api.deepgram.com/v1/listen?model=nova-2&language=ru&encoding=linear16&sample_rate=16000&channels=1&interim_results=true")!) request.setValue("Token \(apiKey)", forHTTPHeaderField: "Authorization") webSocket = session.webSocketTask(with: request) webSocket?.resume() receiveLoop() let inputNode = audioEngine.inputNode let format = AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 16000, channels: 1, interleaved: false)! inputNode.installTap(onBus: 0, bufferSize: 4096, format: format) { buffer, _ in guard let channelData = buffer.int16ChannelData else { return } let frameLength = Int(buffer.frameLength) let data = Data(bytes: channelData[0], count: frameLength * 2) self.webSocket?.send(.data(data)) { _ in } } try audioEngine.start() } private func receiveLoop() { webSocket?.receive { [weak self] result in if case .success(let message) = result, case .string(let text) = message { // Decode Deepgram JSON response self?.handleTranscript(text) } self?.receiveLoop() } } } 

Important detail: AVAudioEngine.inputNode on iOS 16+ requires explicit microphone permission via AVAudioSession.sharedInstance().requestRecordPermission. And обязательно AVAudioSession.setCategory(.record, mode: .measurement) – the .measurement mode disables AEC and AGC, which can distort the signal for transcription.

Android: AudioRecord + OkHttp WebSocket

class DeepgramStreamer(private val apiKey: String) { private val client = OkHttpClient() private var webSocket: WebSocket? = null private var audioRecord: AudioRecord? = null fun start(onTranscript: (String, Boolean) -> Unit) { val request = Request.Builder() .url("wss://api.deepgram.com/v1/listen?model=nova-2&language=ru&encoding=linear16&sample_rate=16000&channels=1&interim_results=true") .header("Authorization", "Token $apiKey") .build() webSocket = client.newWebSocket(request, object : WebSocketListener() { override fun onMessage(webSocket: WebSocket, text: String) { val json = JSONObject(text) val channel = json.getJSONObject("channel") val alternatives = channel.getJSONArray("alternatives") val transcript = alternatives.getJSONObject(0).getString("transcript") val isFinal = json.getBoolean("is_final") if (transcript.isNotEmpty()) onTranscript(transcript, isFinal) } }) val bufferSize = AudioRecord.getMinBufferSize(16000, AudioFormat.CHANNEL_IN_MONO, AudioFormat.ENCODING_PCM_16BIT) audioRecord = AudioRecord(MediaRecorder.AudioSource.MIC, 16000, AudioFormat.CHANNEL_IN_MONO, AudioFormat.ENCODING_PCM_16BIT, bufferSize) audioRecord?.startRecording() Thread { val buffer = ShortArray(bufferSize / 2) while (audioRecord?.recordingState == AudioRecord.RECORDSTATE_RECORDING) { val read = audioRecord!!.read(buffer, 0, buffer.size) if (read > 0) { val byteBuffer = ByteBuffer.allocate(read * 2).order(ByteOrder.LITTLE_ENDIAN) buffer.take(read).forEach { byteBuffer.putShort(it) } webSocket?.send(byteBuffer.array().toByteString()) } } }.start() } } 

ByteOrder.LITTLE_ENDIAN is mandatory. Deepgram expects LE PCM. Sending BE will work but with noticeably worse quality.

How to avoid typical streaming audio mistakes?

  • Interim duplication: never accumulate all interim as separate lines. Store the current utterance in a buffer and overwrite it with each new interim. When is_final: true arrives, finalize the buffer.
  • Connection loss: implement reconnect with exponential backoff (1,2,4,8 sec). Deepgram does not support session resumption, so after reconnect you must start a new stream.
  • Incorrect sample rate: strictly use 16000 Hz. Higher rates increase traffic without quality gain, lower rates degrade recognition.

Handling interim results

Deepgram returns two types of messages: with is_final: false (interim) and is_final: true (final). Correct UI pattern:

  • Display interim in gray or italics – the user sees recognition in progress
  • When is_final: true is received, replace all previous interim of that utterance with the final text
  • speech_final: true indicates the end of a pause – a good moment to start processing the phrase

Nova-2 parameters that affect quality

Parameter Value Description
model nova-2 Recognition model
encoding linear16 Audio encoding
sample_rate 16000 Sample rate
interim_results true Enable partial results
utterance_end_ms 1000 Finalization on pause
diarize false Speaker separation
punctuate true Auto-punctuation
smart_format true Formatting numbers and dates
  • utterance_end_ms: 1000 – Deepgram automatically finalizes utterance after 1 second of silence. Useful for dictation without explicit "stop" commands.
  • diarize: true – speaker separation, adds speaker to each word.
  • punctuate: true – auto-punctuation. Without it, text lacks periods and commas.
  • smart_format: true – formats numbers, dates, phones. "twenty-fifth of March" → "25 March".

What's included in the work

  • Setting up WebSocket connection with authorization
  • Audio capture via AVAudioEngine / AudioRecord in correct format
  • Processing interim and final results without duplication
  • Reconnect on network drop with exponential backoff
  • Testing on real devices (iOS/Android)
  • Documentation and team training

Work process

  1. Analysis – we study the application architecture and transcription requirements
  2. Design – we choose the model, protocol, and parameters
  3. Implementation – code integration of WebSocket, audio capture, UI
  4. Testing – load testing, verification on different devices and networks
  5. Deployment – publishing in App Store and Google Play

Timeline

Basic integration of WebSocket + AudioRecord/AVAudioEngine + text output – 4–7 days. Adding diarization, network switching handling (reconnect), background mode, and result export – 8–14 days.

Savings on developing your own ASR engine can reach 70%. Transcription cost is low – significantly cheaper than alternatives with batch processing.

Get a consultation on integrating Deepgram for your mobile project. We will analyze the architecture, choose the optimal configuration, and implement low-latency transcription turnkey. Order a demo version with low latency today. Contact us to discuss your project.

Integrating Deepgram allowed us to reduce transcription latency from 5 seconds to 300 ms – client feedback from the financial sector.