Voice messages are the worst format for quick information retrieval. Especially in corporate chat: a minute-long voice note instead of one line of text. We solve this by embedding transcription directly into your mobile app. The user speaks—the app converts speech to text with up to 95% accuracy. The text is synchronized with audio: tap any word to hear it. Technically, this means integrating with Whisper API or local solutions, implementing audio capture, converting to the optimal format (16 kHz mono MP3 32 kbps), sending to the server, and receiving a transcript with timestamps every 200–300 ms. Then post-processing, noise notation filtering, and displaying in an interactive UI. The entire cycle from button press to visible text takes 0.5 to 3 seconds depending on recording length. One of our projects in fintech showed a 60% reduction in meeting protocol processing time—saving approximately $2,000 per month.
How Audio Capture Works in the App
In practice, a mobile app works with two paths:
Recording inside the app. The user records directly in your app—native capture, full control over the format. We use AVAudioRecorder on iOS and MediaRecorder on Android. The audio sample rate of 16kHz is chosen based on the Nyquist theorem for speech (max 8kHz frequency content), and the MP3 codec at 32kbps ensures compression without significant loss of phonetic information.
Importing an external file. Get a WAV/MP3/OGG from a messenger via share sheet. On iOS—UTType.audio in UIDocumentPickerViewController. On Android—ACTION_GET_CONTENT with "audio/*". File format matters. OGG Opus (Telegram format) Whisper understands natively. AMR (old Android messengers)—needs conversion. On the server, ffmpeg handles conversion of any format:
import subprocess
def convert_to_mp3(input_path: str, output_path: str) -> None:
subprocess.run([
"ffmpeg", "-i", input_path,
"-ar", "16000", # 16kHz is enough for speech
"-ac", "1", # mono
"-b:a", "32k", # 32kbps for speech
output_path
], check=True)
16kHz mono MP3 32kbps—optimal for Whisper: quality doesn't drop, file size is minimal.
Why Whisper API Is Not the Only Option
Whisper API: A 10-second message processes in 0.5–1.5 s. A 1-minute message in 3–8 s. This includes processing time on OpenAI servers plus network. Acceptable for the user if progress is shown. According to OpenAI Whisper GitHub repository, the model achieves state-of-the-art accuracy.
Deepgram Nova-2—real-time streaming transcription, latency <300 ms on short fragments. More expensive than Whisper, but faster. Deepgram Nova-2 offers latency under 300ms, which is up to 10x faster than Whisper API for short fragments.
Local Whisper (self-hosted). faster-whisper on GPU (RTX 3090) processes 1 minute of audio in 2–4 seconds. On CPU—15–30 seconds. If data cannot be sent to the cloud—the only option.
Client-side transcription on iOS. SFSpeechRecognizer—native Apple Speech framework, works on-device (since iOS 16), free, no data sent. But: supports only a limited set of languages, quality lower than Whisper, limit of 1 minute per request.
// iOS — local transcription via SFSpeechRecognizer
let recognizer = SFSpeechRecognizer(locale: Locale(identifier: "ru-RU"))
let request = SFSpeechURLRecognitionRequest(url: audioURL)
request.shouldReportPartialResults = true
recognizer?.recognitionTask(with: request) { result, error in
guard let result else { return }
DispatchQueue.main.async {
self.transcriptText = result.bestTranscription.formattedString
}
}
For short personal notes, SFSpeechRecognizer is a good option without server costs. For corporate meeting recordings—Whisper or Deepgram.
Comparison of Transcription Methods
| Method | Latency | Quality | Cost | Privacy |
|---|---|---|---|---|
| Whisper API | 0.5–8 s | Excellent | $0.006/min | Data sent to server |
| Deepgram Nova-2 | <300 ms | Excellent | Higher | Data sent to server |
| Local Whisper (GPU) | 2–4 s per minute | Excellent | Hardware only | Fully local |
| SFSpeechRecognizer (iOS) | Instant | Medium | Free | Fully local |
How to Display Transcript with Timestamps
Simple transcription—just text. Good transcription on mobile:
- Interactive text with timestamps: tap a word → audio jumps to that moment
- Punctuation (Whisper restores it well, but not perfectly—sometimes post-processing needed)
- Paragraphs by pauses (Whisper segments audio—use
segmentsfor splitting) - Copy all text button
- Search within transcript
For messenger-style functionality: transcript appears streaming—don't wait for full completion, show segments as they become ready.
Transcript Post-Processing
Whisper sometimes inserts [Music], [Applause] in Whisper notation, transcribes background noise. We filter them:
import re
def clean_transcript(text: str) -> str:
# Remove Whisper notations like [Music], [Noise]
text = re.sub(r'\[.*?\]', '', text)
# Remove extra spaces
text = re.sub(r'\s+', ' ', text).strip()
return text
For business scenarios, LLM post-processing is useful: fix proper names, terms, add punctuation where Whisper made mistakes. This server-side transcription for mobile ensures high quality.
What's Included in the Work
- Source code of the transcription module for iOS and Android
- Documentation on architecture and REST API (if server side)
- Access to services (OpenAI, Deepgram) with ready keys
- Team training and consultations during integration
- 24/7 support after launch
How to Implement Transcription: Step-by-Step Plan
- Analysis — discuss use cases, stack, latency and privacy requirements.
- Design — architecture of capture, transcription, and display.
- Implementation — integrate Whisper/Deepgram, code for iOS/Android, server-side conversion.
- Testing — validate on real recordings, optimize for your case.
- Deploy — release to App Store and Google Play, set up monitoring.
- Documentation and training — hand over code, instructions, train your team.
- Support — guarantee 24/7 stability after launch.
Timelines and Cost
| Stage | Duration |
|---|---|
| Audio capture + file import | 3–5 days |
| Server-side transcription (Whisper) + progress | 5–7 days |
| Post-processing and formatting | 2–3 days |
| Mobile UI with interactive transcript | 5–7 days |
| Optional: streaming, local SFSpeechRecognizer | +3–5 days |
Basic transcription via Whisper with plain text display — 1–2 weeks. Full tool with interactive text, timestamps, and post-processing — 3–4 weeks. Integration cost is determined individually.
With over 5 years of experience in mobile speech integration and 30+ successful implementations for fintech and healthcare clients, our team ensures reliable delivery. Contact us for a free assessment of your project. Request a demo version to test on real data.







