A user dictates a command, waits for a response — and after 3 seconds gets the wrong thing. Sound familiar? In a voice assistant, speed and recognition accuracy are paramount. We solve this by building a pipeline from VAD, STT, NLU, and TTS with a total latency from the end of the phrase to the response ≤1.5 seconds. This is a technical constraint we overcome by optimizing each stage: from voice detection to speech synthesis. We handle the entire cycle: design, development, integration with your CRM or ERP, and app store publication. Over 7+ years, we have deployed voice assistants in 15+ projects for iOS and Android. Evaluate our approach — contact us for a consultation.
How the Voice Assistant Pipeline Works
Microphone → VAD → STT → NLU → Logic → TTS → Speaker. Each component contributes to latency.
VAD — Voice Activity Detection, cuts silence. We use WebRTCVAD or SileroVAD (ONNX/TFLite, ~1 MB). This reduces empty STT requests and saves traffic.
Example VAD configuration on Android
```kotlin val vad = SileroVAD.create(context) vad.start { frame -> if (vad.isVoice(frame)) { // send audio to STT } } ```STT — Speech-to-Text. Options: native SFSpeechRecognizer (iOS) or Android Speech for simple scenarios; for high accuracy Russian — Yandex SpeechKit or OpenAI Whisper API. Apple Developer Documentation recommends using SFSpeechRecognizer for basic commands. Cloud recognition costs ~$0.006 per audio minute for Whisper API; on-device is free. A typical request lasts 3–5 seconds, so costs are minimal.
Why Latency ≤1.5s Is Critical
Users expect instant response. Delay over 2 seconds feels like a hang. We achieve this through parallel requests, local processing, and caching frequent intents. Each component has its typical delay:
| Component | Typical Latency | Cost (per audio minute) |
|---|---|---|
| VAD (Silero) | 30–50 ms | $0 (on-device) |
| STT (Whisper API) | 200–400 ms | ~$0.006 |
| NLU (Rasa) | 200–400 ms | $0 (self-hosted) |
| TTS (Yandex SpeechKit) | 200–500 ms | ~$0.002 |
Total — up to 1.5 seconds. Replacing NLU with an LLM (GPT-4) can increase latency to 4 seconds, unacceptable for real-time.
Intent Recognition: What Actually Works
For a limited domain (smart home, internet banking) — Rasa NLU or Dialogflow with 50–200 training examples per intent. For an open domain — LLM with function calling. Rasa NLU is better than Dialogflow for confidential data since it runs on your server and does not send speech to the cloud.
| Feature | Rasa NLU | Dialogflow | LLM (GPT-4) |
|---|---|---|---|
| Privacy | Full | Google Cloud | Cloud (prompt) |
| Domain Accuracy | 90%+ | 85%+ | 95%+ (but slower) |
| Setup Difficulty | Medium | Low | High (prompt engineering) |
| Latency | 200–400 ms | 500–800 ms | 1–4 sec |
Case Study from Our Practice
Corporate assistant for field employees: voice task creation in CRM without unlocking the phone. Stack: SileroVAD on-device -> Yandex SpeechKit streaming -> Rasa NLU (self-hosted, 23 intents) -> CRM REST API -> Yandex SpeechKit TTS. Latency: median 1.1 s, p95 2.3 s. Rasa NLU provided full data control. The client estimated time savings of ~25% for employees.
How to Implement a Voice Assistant: Step-by-Step Plan
- Scenario Analysis — define command list and contexts (up to 3 days).
- Component Selection — STT, NLU, TTS considering language, privacy, and budget.
- Pipeline Integration — connect modules, tune VAD parameters and timeouts.
- Testing on Real Data — record dialogues, A/B tests, optimization.
- App Store Release — prepare metadata, test with TestFlight/Internal Track.
When On-Device STT Is Needed?
If the app must work offline or requires minimal latency — choose on-device. Accuracy is lower (80–90%), but latency is 300–500 ms and no API call costs. For Russian, on-device still lags behind cloud solutions but works for a limited set of phrases.
What’s Included in the Work
- Architecture documentation and API specifications
- Configured CI/CD for build and deployment
- Source code and repository access
- Training your team on the voice pipeline
- Post-release support for 2 weeks
Estimated Timelines
- Basic pipeline STT + NLU + TTS: 2–3 weeks
- With wake word and context: 4–6 weeks
- With integration into existing infrastructure: individually defined
Cost is determined after analyzing your requirements. Get a consultation for your scenario — contact us. We will evaluate your project in 1–2 days.







