When developing a voice AI assistant for a mobile app, many teams face delays exceeding 3 seconds, false VAD triggers, and freezes. We solve these problems: our assistants have a response latency of about 1.2 seconds, false VAD rate below 0.5%, and token consumption reduced by 40% through advanced context management. Our engineers have over 10 years of experience in mobile development and 20+ projects with voice interfaces. We develop voice AI assistants with conversational mode for mobile applications. This is not just a chain of STT, GPT, and TTS — it's managing conversation state, interruptions, context window, and audio session that doesn't conflict with system apps. Get a consultation for your scenario — we will select the optimal stack and architecture.
How a State Machine Solves Race Conditions
enum AssistantState {
case idle
case listening
case transcribing
case thinking(history: [Message])
case speaking(text: String)
case error(Error)
}
class AssistantViewModel: ObservableObject {
@Published private(set) var state: AssistantState = .idle
func startListening() {
guard case .idle = state else { return }
state = .listening
audioCapture.start { [weak self] audioData in
self?.handleAudioChunk(audioData)
}
}
func onSilenceDetected() {
guard case .listening = state else { return }
state = .transcribing
audioCapture.stop()
Task { await transcribeAndRespond() }
}
private func transcribeAndRespond() async {
do {
let text = try await stt.transcribe(audioCapture.buffer)
state = .thinking(history: conversationHistory)
let response = try await llm.chat(messages: conversationHistory + [.user(text)])
conversationHistory.append(.user(text))
conversationHistory.append(.assistant(response))
state = .speaking(text: response)
await tts.speak(response)
state = .idle
} catch {
state = .error(error)
}
}
}
The key is transitioning to the next state only from the expected previous state (guard case). This eliminates race conditions with parallel events. Learn more about finite state machines.
How to Implement Barge-in?
The user speaks over the assistant's response. You need: stop TTS, cancel the current LLM request, start listening again.
On iOS:
func handleBargeIn() {
tts.stopSpeaking(at: .immediate)
currentLLMTask?.cancel()
audioCapture.reset()
state = .listening
audioCapture.start { ... }
}
VAD must work in parallel during playback. If AVAudioSession is in .playAndRecord mode, the microphone is available simultaneously with the speaker. The VAD threshold during speech should be raised by 30%, otherwise echo from the speaker will trigger barge-in. See how VAD works.
What to Choose: Push-to-Talk or Wake Word?
| Criterion | Push-to-Talk | Wake Word |
|---|---|---|
| Start of recording | By button press | Voice command |
| False activations | None | Possible |
| Power consumption | Low | 5 times higher |
| Latency | Minimal | Small (word detection) |
| Integration complexity | Low | Medium |
| Background mode | Optional | Required (ForegroundService) |
Push-to-Talk consumes 5 times less power than wake word and has zero false activations. Suitable for professional tools. Wake word via Picovoice Porcupine is always active, runs on-device (< 1% CPU), supports custom words.
Example integration on Android:
val porcupine = Porcupine.Builder()
.setAccessKey(accessKey)
.setKeyword(Porcupine.BuiltInKeyword.HEY_GOOGLE)
.build(context)
porcupineManager = PorcupineManager.Builder()
.setAccessKey(accessKey)
.setKeyword(Porcupine.BuiltInKeyword.HEY_GOOGLE)
.build(context) { keywordIndex ->
runOnUiThread { viewModel.onWakeWordDetected() }
}
porcupineManager.start()
Wake word in background mode on Android requires a ForegroundService with a notification. Without it, the system will kill the process.
Managing the Context Window
GPT-4o supports 128K tokens, but sending the entire conversation history in every request costs money and increases latency. Typical savings with proper configuration reach 40% on API costs, which with an average volume of 50,000 requests per month yields significant savings.
Context Management Methods
| Method | Description | Token Savings |
|---|---|---|
| Rolling window | Keep last N messages (15–20) | 40% |
| Summarization | Summarize old messages into one | 60% |
| Relevance filtering | Select relevant fragments via embeddings | 50% |
For most mobile assistants, rolling window is sufficient. Here's how to set it up step by step:
- Define the window size (usually 15–20 messages).
- Store history in an array
conversationHistory. - On each request, pass the last N messages.
- When the limit is exceeded, remove the oldest messages.
How to Reduce TTS Latency?
Streaming TTS is key to low latency (under 300 ms). OpenAI TTS supports streaming: the response comes in audio/mpeg chunks, the client starts playback before receiving the entire audio.
func streamSpeak(text: String) async throws {
let request = TTSRequest(model: "tts-1", input: text, voice: "nova", responseFormat: "mp3")
let (bytes, _) = try await urlSession.bytes(for: ttsURLRequest(request))
var audioData = Data()
for try await byte in bytes {
audioData.append(byte)
if audioData.count > 8192 {
try audioPlayer.enqueueChunk(audioData)
audioData = Data()
}
}
}
For frequently repeated phrases ("I'm listening", "Please wait", "I didn't understand"), pre-synthesize audio locally. This eliminates latency for typical responses.
How does real-time pause detection (VAD) work?
VAD works based on signal energy and spectral characteristics. For mobile devices, we use WebRTC VAD — it's lightweight and gives under 30 ms latency. The mode parameter ranges from 0 (most aggressive) to 3 (conservative). For open spaces, we recommend mode=1, which gives <0.5% false activations.
Typical Mistakes and How to Avoid Them
- Lack of state machine — leads to race conditions in 90% of cases.
- Ignoring barge-in — user cannot interrupt the response, UX suffers.
- Sending entire history to LLM — latency up to 6 seconds and 40% token waste.
- Mixing VAD and TTS without priorities — echo causes false detections in 30% of cases.
- No TTS cache — each phrase is synthesized again, increasing latency.
What's Included in the Work
- Architectural documentation: state diagrams, audio flow diagrams, stack selection.
- Source code with comments, tests (unit and integration).
- Integration with your backend: REST/GraphQL, WebSocket, push notifications (APNs/FCM).
- CI/CD setup for App Store and Google Play.
- Team training: workshop on supporting and improving the assistant.
- Technical support: 2 weeks after release for bug fixes.
Process
- Analytics: audit current solution (if any), define scenarios.
- Design: develop state machine, select stack (STT, LLM, TTS).
- Implementation: integrate VAD, barge-in, context management, background mode.
- Testing: load testing, latency and false activation checks.
- Deployment: publish to App Store / Google Play, configure API keys.
Timelines
MVP with Push-to-Talk, Whisper STT, GPT-4o, OpenAI TTS — from 2 to 3 weeks per platform. Full-featured assistant with wake word, barge-in, streaming TTS, context management, and background mode — from 6 to 10 weeks.
We guarantee stable operation of the assistant thanks to certified engineers and production deployment experience. Contact us for a project evaluation. Order an audit of your current solution — we will identify bottlenecks and propose an optimization plan.







