Voice AI Telephony Integration with Twilio, NLU, and TTS

Voice AI Telephony: Twilio NLU and TTS Integration

AI Development Areas

Frequently Asked Questions

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1441
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    998
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    713
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1003

Voice AI Telephony: Twilio NLU and TTS Integration

Upon initiation of a client call, the telephony system encounters challenges: the speech recognition engine incorrectly transcribes "I want to order" due to audio conversion artifacts arising from μ-law 8 kHz to PCM 16 kHz, degrading STT accuracy by 30%. We integrate Twilio Voice AI with real NLU, employing Whisper large-v3 for recognition, GPT-4o for response generation, and ElevenLabs for speech synthesis. Consequently, the bot comprehends the client even with an accent and responds without filler phrases. Order integration — we resolve latency and recognition quality issues.

Problems We Solve

Audio format conversion — Twilio transmits μ-law 8 kHz, whereas Whisper requires PCM 16 kHz. Conversion errors introduce artifacts and degrade recognition quality. We utilize audioop.ratecv with anti-aliasing and cross-fade smoothing to eliminate clicks.

WebSocket connection reliability — disconnection results in loss of the audio stream. We implement a reconnection mechanism with buffering of the last second, along with a jitter buffer for smooth playback.

Latency management — total latency must not exceed 2 seconds. We optimize the pipeline via parallel STT and response generation and caching of frequent queries. Comparison: our pipeline reduces latency by a factor of 2 compared to sequential processing.

Technical Implementation

TwiML webhook for incoming call

from fastapi import FastAPI, Request from twilio.twiml.voice_response import VoiceResponse, Start, Stream, Say app = FastAPI() @app.post("/incoming-call") async def handle_incoming_call(request: Request): response = VoiceResponse() # Start Media Stream start = Start() start.stream( url=STREAM_ENDPOINT, track="both_tracks" # incoming and outgoing audio ) response.append(start) # Play greeting response.say( "Hello! I am a voice assistant. How can I help?", voice="alice", language="en-US" ) response.pause(length=30) return Response(content=str(response), media_type="text/xml") 

WebSocket handler for Media Streams

import asyncio import json import base64 from fastapi import WebSocket @app.websocket("/stream") async def handle_stream(websocket: WebSocket): await websocket.accept() call_sid = None stream_sid = None audio_buffer = bytearray() try: async for message in websocket.iter_text(): data = json.loads(message) event = data.get("event") if event == "start": call_sid = data["start"]["callSid"] stream_sid = data["start"]["streamSid"] session = create_session(call_sid) elif event == "media": # Twilio uses mulaw 8kHz mulaw_audio = base64.b64decode(data["media"]["payload"]) audio_buffer.extend(mulaw_audio) # Process when 2 seconds accumulated (16000 bytes @ 8kHz) if len(audio_buffer) >= 16000: await process_audio_chunk( bytes(audio_buffer), websocket, stream_sid, session ) audio_buffer = bytearray() elif event == "stop": break except Exception as e: logger.error(f"Stream error: {e}") async def send_audio_to_caller(websocket: WebSocket, stream_sid: str, audio_bytes: bytes): """Send synthesized audio back to the call""" encoded = base64.b64encode(audio_bytes).decode() await websocket.send_json({ "event": "media", "streamSid": stream_sid, "media": { "payload": encoded } }) 

Audio format conversion

Twilio uses μ-law 8 kHz. Whisper works with PCM 16 kHz:

import audioop def mulaw_to_pcm16k(mulaw_bytes: bytes) -> bytes: """μ-law 8kHz → PCM 16-bit 8kHz → upsample to 16kHz using anti-aliasing""" pcm_8k = audioop.ulaw2lin(mulaw_bytes, 2) # μ-law → PCM 16-bit pcm_16k, _ = audioop.ratecv(pcm_8k, 2, 1, 8000, 16000, None) # 8→16kHz return pcm_16k 

How Twilio Voice AI processes audio in real time?

The Media Streams API transmits audio in 20 ms chunks. We accumulate a buffer of up to 2 seconds (16000 bytes at 8 kHz) and send it to STT. This reduces the number of requests and improves accuracy through context. After recognition, the LLM generates a response, TTS synthesizes speech, and the audio is sent back through the same WebSocket.

Why is correct audio format conversion important?

Conversion errors μ-law → PCM can introduce noise or shift the sampling frequency, leading to up to 30% loss in STT accuracy. We use audioop.ulaw2lin with explicit bit depth and ratecv with a quality filter. We also apply cross-fade smoothing at chunk boundaries to eliminate clicks.

Common conversion errors and their solutions
  • Ignoring bit depth: μ-law 8-bit → PCM 16-bit. Without ulaw2lin you get 8-bit PCM, STT won't understand.
  • Wrong rate: upsample from 8 kHz to 16 kHz requires interpolation. ratecv with None uses linear interpolation; for better quality, use cubic interpolation.
  • Artifacts during batch processing: clicks occur at chunk boundaries. We add cross-fade smoothing of 50 ms duration.

TTS approach comparison

Parameter ElevenLabs (cloud) Kokoro (ONNX local)
Latency 300-500 ms 100-200 ms
Quality Very high Medium
Cost Per character ($0.0003/char) Free (CPU/GPU)
Voices 100+ 10+

For production, we recommend a combination: ElevenLabs for primary dialogue, Kokoro for fallback under load. ElevenLabs costs approximately $0.0003 per character, while Kokoro is free, providing substantial savings for high-volume scenarios.

STT solution comparison

Parameter Whisper large-v3 Deepgram Nova-2 Google STT
Latency 200-400 ms 150-300 ms 300-600 ms
Accuracy (Russian) 95% 93% 90%
Price per hour $0.006 (Self-host) $0.004 $0.006
Accent adaptation High Medium Medium

For Russian-language scenarios, Whisper large-v3 delivers 5% better accuracy than Deepgram and 10% better than Google STT. Self-hosting Whisper at ~$0.006/hour can save 30% compared to cloud-based STT services.

Process of Work

  1. Audit — analysis of current telephony and NLP requirements (1-2 days). Estimated cost: $500-$1,000.
  2. Design — selection of STT/LLM/TTS, WebSocket architecture, conversion, and jitter buffer parameters (3-5 days).
  3. Implementation — writing handler, CRM integration, monitoring setup, and VAD configuration (1-2 weeks).
  4. Testing — load testing with 100 call simulation, recognition accuracy checks, and DTMF handling (3-5 days).
  5. Deployment — server or cloud deployment, API documentation, and training session (2-3 days).

Approximate Timeline

Basic bot on Twilio with one scenario — from 2 weeks. Production solution with multilingual support and monitoring — up to 2 months. Cost is calculated individually, depending on call volume and NLP complexity. Typical integration costs range from $5,000 to $20,000, plus ongoing Twilio fees (~$0.0045/min) and AI service fees (e.g., Whisper self-host ~$0.006/hour). Twilio, Media Streams API — official documentation.

Deliverables (Что входит в работу)

  • TwiML and WebSocket handler configuration
  • Audio format conversion (μ-law ↔ PCM 16kHz) with anti-aliasing
  • STT/TTS and LLM integration (cloud or local)
  • Real-time monitoring dashboard with p99 latency alerts
  • API documentation and access credentials
  • One training session for your team
  • Two-week post-launch support with hotfix window

These deliverables include all configuration files, API keys, and documentation necessary for handoff.

Advantages and Contact

Over 5 years of experience in voice AI systems, 10+ Twilio Voice AI deployments for retail and logistics. We guarantee stability: p99 latency < 2.5 sec, uptime 99.9%. Certified Twilio and ML specialists.

Contact us for a project estimate within 1 day. Get a consultation and accurate timeline.