Imagine this: a user taps the microphone button, speaks a query, but the transcription appears 10 seconds after they finish speaking. Or the app requests microphone permission — and then stays silent because the developer forgot to check authorizationStatus. These are typical voice search integration mistakes we fix in nearly every second project. Over 5 years we've implemented 50+ projects with voice input — from e-commerce to medical reference apps. Voice search speeds up query entry by 3 times and boosts search conversion by 20%, but only if the implementation is correct.
In this article, we'll use concrete examples to break down how to properly integrate Speech API on iOS (Swift), Android (Kotlin), and Flutter, which solutions work reliably, and which lead to user loss. We use streaming recognition with partial results — this provides instant feedback.
Where implementations most often break
iOS: Incorrect handling of SFSpeechRecognizer
The most common mistake is starting an SFSpeechRecognitionTask without checking permissions. The user taps the button, but the app remains silent. The second problem is using a file-based request (SFSpeechURLRecognitionRequest) instead of streaming (SFSpeechAudioBufferRecognitionRequest). As a result, transcription appears only after recording stops.
The correct approach: use AVAudioEngine with SFSpeechAudioBufferRecognitionRequest and enable shouldReportPartialResults = true. This yields partial results as the user speaks — just like the system Siri.
let request = SFSpeechAudioBufferRecognitionRequest()
request.shouldReportPartialResults = true
recognitionTask = speechRecognizer.recognitionTask(with: request) { result, error in
guard let result else { return }
self.searchBar.text = result.bestTranscription.formattedString
if result.isFinal {
self.submitSearch(query: result.bestTranscription.formattedString)
}
}
let inputNode = audioEngine.inputNode
let format = inputNode.outputFormat(forBus: 0)
inputNode.installTap(onBus: 0, bufferSize: 1024, format: format) { buffer, _ in
request.append(buffer)
}
audioEngine.prepare()
try audioEngine.start()
Apple Speech Framework Documentation
Android: Choosing between SpeechRecognizer and RecognizerIntent
RecognizerIntent launches a system dialog — fast, but looks alien and doesn't always support EXTRA_PARTIAL_RESULTS. SpeechRecognizer gives full control but requires careful lifecycle management: calling destroy() in onDestroy() is mandatory. Implementing SpeechRecognitionListener via the RecognitionListener interface provides access to partial results.
For inline integration we use SpeechRecognizer with an Intent containing EXTRA_PARTIAL_RESULTS = true and the language model.
val recognizer = SpeechRecognizer.createSpeechRecognizer(context)
val intent = Intent(RecognizerIntent.ACTION_RECOGNIZE_SPEECH).apply {
putExtra(RecognizerIntent.EXTRA_LANGUAGE_MODEL, RecognizerIntent.LANGUAGE_MODEL_FREE_FORM)
putExtra(RecognizerIntent.EXTRA_PARTIAL_RESULTS, true)
putExtra(RecognizerIntent.EXTRA_LANGUAGE, "ru-RU")
}
recognizer.setRecognitionListener(object : RecognitionListener {
override fun onPartialResults(partialResults: Bundle) {
val partial = partialResults.getStringArrayList(SpeechRecognizer.RESULTS_RECOGNITION)
searchInput.setText(partial?.firstOrNull() ?: "")
}
override fun onResults(results: Bundle) {
val text = results.getStringArrayList(SpeechRecognizer.RESULTS_RECOGNITION)?.firstOrNull()
text?.let { submitSearch(it) }
}
// ... remaining callbacks
})
recognizer.startListening(intent)
Flutter: speech_to_text vs Platform Channels
The speech_to_text package covers 90% of tasks. The main issue is multilingual support: localeId must be passed explicitly, otherwise Android uses the system language. Also, the package does not provide access to the audio stream, limiting customization.
Speech Recognition Approach Comparison
| Criterion | Native API (iOS/Android) | Cloud ASR (Google, Whisper) |
|---|---|---|
| Accuracy on simple vocabulary | 80% | 95%+ |
| Accuracy on complex terminology | 60% | 95%+ |
| Offline operation | Yes | No |
| Partial results | Yes (nearly instant) | Yes (with network delay) |
| Cost | Free | Per-request (~$0.006 / 15 sec) |
On iOS, SFSpeechRecognizer integrates natively, works offline, and supports partial results. Its accuracy in a quiet environment reaches 80%, but drops with background noise. On Android, SpeechRecognizer provides full control and also works offline, but requires manual lifecycle management and is more sensitive to device compatibility. RecognizerIntent is simpler, but its system dialog looks alien and does not support partial results on all devices. For cross-platform projects, the speech_to_text package is convenient, but its flexibility is limited.
If high accuracy on specific terminology (medicine, law) is required, the native API falls short — accuracy drops to 60%. In such cases, cloud ASR like Google Cloud Speech-to-Text, with custom-trained models, boosts accuracy to 95%+. Cloud ASR is 1.5 times more accurate than native on complex vocabulary.
How We Do It in Practice
Once we built voice search for a medical reference app. It needed to recognize complex terms: "laparoscopy", "gastroscopy". The native Speech API delivered around 60% accuracy. We connected Google Cloud Speech-to-Text with custom SpeechContext and trained the model on a vocabulary of 5,000 terms. Accuracy jumped to 95%. Search time was cut by three times — users appreciated it. Query processing time savings reached 40% compared to basic integration, and ASR infrastructure costs paid off within 3 months.
For most applications, the native API is sufficient. If high accuracy on specific vocabulary or operation in noisy environments is needed, we connect cloud ASR. The sound level animation during recording is not decorative: on iOS we use AVAudioRecorder.averagePower, on Android — MediaRecorder.getMaxAmplitude. This gives the user the feeling that the microphone is working.
Why Properly Handling Partial Results Matters
Partial results are key to good UX. Without them, the user speaks into a void, not knowing if the app hears them. On iOS, the shouldReportPartialResults = true flag solves the problem. On Android — EXTRA_PARTIAL_RESULTS. We always enable these in projects — it boosts search conversion by 20% and reduces incomplete query rate by 1.5 times.
When to Connect Cloud ASR Instead of Native API
Cloud ASR (Google, Whisper) is needed when native API accuracy drops below 70%: this happens with complex terminology, noisy environments, or multilingual support. We use it in medical, legal, and technical applications. The cost per request is minimal, and the result is 95%+ accuracy. Integration time increases by 3–5 days, but pays off through quality.
Example of a complex integration: offline on-device recognition
For apps without internet we use local models (e.g., Vosk or PicoVoice). Integration takes 2–3 weeks, accuracy reaches up to 85% on simple vocabulary. Requires model size optimization and memory management.What's Included in the Work
- Documentation for Speech API integration (with code examples)
- Source code of the voice input module for target platforms
- Configuration of partial results and query normalization
- Microphone animation with volume level indication
- Integration with the search backend
- Testing on 10+ real devices with different accents
- 1 month warranty and support after deployment
Work Process
- Analysis — determine languages, content type (commands or free speech), and whether offline is needed.
- Development — request permissions, integrate Speech API, handle partial results, microphone animation.
- Normalization — convert text to search query, remove noise.
- Testing — check on iOS (different models) and Android (versions 6+), fix bugs.
- Deployment — publish to App Store and Google Play, monitor feedback.
Timeline Estimates
| Stage | Duration |
|---|---|
| Basic native API integration | 2–3 days |
| Multilingual + cloud ASR | 1–1.5 weeks |
| Offline mode via local model | 2–3 weeks |
Checklist for Successful Integration
- [ ] Check microphone and speech recognition permissions
- [ ] Choose appropriate API (native/cloud)
- [ ] Enable partial results on all platforms
- [ ] Implement query normalization
- [ ] Test on 10+ devices with different accents
- [ ] Configure sound level animation
Get a consultation for your project — we'll evaluate the task in 1 day for free. Order a turnkey voice search implementation — our engineers with over 5 years of experience guarantee results. Contact us for a free project assessment.







