Voice AI Telephony: Twilio NLU and TTS Integration
Upon initiation of a client call, the telephony system encounters challenges: the speech recognition engine incorrectly transcribes "I want to order" due to audio conversion artifacts arising from μ-law 8 kHz to PCM 16 kHz, degrading STT accuracy by 30%. We integrate Twilio Voice AI with real NLU, employing Whisper large-v3 for recognition, GPT-4o for response generation, and ElevenLabs for speech synthesis. Consequently, the bot comprehends the client even with an accent and responds without filler phrases. Order integration — we resolve latency and recognition quality issues.
Problems We Solve
Audio format conversion — Twilio transmits μ-law 8 kHz, whereas Whisper requires PCM 16 kHz. Conversion errors introduce artifacts and degrade recognition quality. We utilize audioop.ratecv with anti-aliasing and cross-fade smoothing to eliminate clicks.
WebSocket connection reliability — disconnection results in loss of the audio stream. We implement a reconnection mechanism with buffering of the last second, along with a jitter buffer for smooth playback.
Latency management — total latency must not exceed 2 seconds. We optimize the pipeline via parallel STT and response generation and caching of frequent queries. Comparison: our pipeline reduces latency by a factor of 2 compared to sequential processing.
Technical Implementation
TwiML webhook for incoming call
from fastapi import FastAPI, Request
from twilio.twiml.voice_response import VoiceResponse, Start, Stream, Say
app = FastAPI()
@app.post("/incoming-call")
async def handle_incoming_call(request: Request):
response = VoiceResponse()
# Start Media Stream
start = Start()
start.stream(
url=STREAM_ENDPOINT,
track="both_tracks" # incoming and outgoing audio
)
response.append(start)
# Play greeting
response.say(
"Hello! I am a voice assistant. How can I help?",
voice="alice",
language="en-US"
)
response.pause(length=30)
return Response(content=str(response), media_type="text/xml")
WebSocket handler for Media Streams
import asyncio
import json
import base64
from fastapi import WebSocket
@app.websocket("/stream")
async def handle_stream(websocket: WebSocket):
await websocket.accept()
call_sid = None
stream_sid = None
audio_buffer = bytearray()
try:
async for message in websocket.iter_text():
data = json.loads(message)
event = data.get("event")
if event == "start":
call_sid = data["start"]["callSid"]
stream_sid = data["start"]["streamSid"]
session = create_session(call_sid)
elif event == "media":
# Twilio uses mulaw 8kHz
mulaw_audio = base64.b64decode(data["media"]["payload"])
audio_buffer.extend(mulaw_audio)
# Process when 2 seconds accumulated (16000 bytes @ 8kHz)
if len(audio_buffer) >= 16000:
await process_audio_chunk(
bytes(audio_buffer), websocket, stream_sid, session
)
audio_buffer = bytearray()
elif event == "stop":
break
except Exception as e:
logger.error(f"Stream error: {e}")
async def send_audio_to_caller(websocket: WebSocket, stream_sid: str, audio_bytes: bytes):
"""Send synthesized audio back to the call"""
encoded = base64.b64encode(audio_bytes).decode()
await websocket.send_json({
"event": "media",
"streamSid": stream_sid,
"media": {
"payload": encoded
}
})
Audio format conversion
Twilio uses μ-law 8 kHz. Whisper works with PCM 16 kHz:
import audioop
def mulaw_to_pcm16k(mulaw_bytes: bytes) -> bytes:
"""μ-law 8kHz → PCM 16-bit 8kHz → upsample to 16kHz using anti-aliasing"""
pcm_8k = audioop.ulaw2lin(mulaw_bytes, 2) # μ-law → PCM 16-bit
pcm_16k, _ = audioop.ratecv(pcm_8k, 2, 1, 8000, 16000, None) # 8→16kHz
return pcm_16k
How Twilio Voice AI processes audio in real time?
The Media Streams API transmits audio in 20 ms chunks. We accumulate a buffer of up to 2 seconds (16000 bytes at 8 kHz) and send it to STT. This reduces the number of requests and improves accuracy through context. After recognition, the LLM generates a response, TTS synthesizes speech, and the audio is sent back through the same WebSocket.
Why is correct audio format conversion important?
Conversion errors μ-law → PCM can introduce noise or shift the sampling frequency, leading to up to 30% loss in STT accuracy. We use audioop.ulaw2lin with explicit bit depth and ratecv with a quality filter. We also apply cross-fade smoothing at chunk boundaries to eliminate clicks.
Common conversion errors and their solutions
- Ignoring bit depth: μ-law 8-bit → PCM 16-bit. Without
ulaw2linyou get 8-bit PCM, STT won't understand. - Wrong rate: upsample from 8 kHz to 16 kHz requires interpolation.
ratecvwithNoneuses linear interpolation; for better quality, use cubic interpolation. - Artifacts during batch processing: clicks occur at chunk boundaries. We add cross-fade smoothing of 50 ms duration.
TTS approach comparison
| Parameter | ElevenLabs (cloud) | Kokoro (ONNX local) |
|---|---|---|
| Latency | 300-500 ms | 100-200 ms |
| Quality | Very high | Medium |
| Cost | Per character ($0.0003/char) | Free (CPU/GPU) |
| Voices | 100+ | 10+ |
For production, we recommend a combination: ElevenLabs for primary dialogue, Kokoro for fallback under load. ElevenLabs costs approximately $0.0003 per character, while Kokoro is free, providing substantial savings for high-volume scenarios.
STT solution comparison
| Parameter | Whisper large-v3 | Deepgram Nova-2 | Google STT |
|---|---|---|---|
| Latency | 200-400 ms | 150-300 ms | 300-600 ms |
| Accuracy (Russian) | 95% | 93% | 90% |
| Price per hour | $0.006 (Self-host) | $0.004 | $0.006 |
| Accent adaptation | High | Medium | Medium |
For Russian-language scenarios, Whisper large-v3 delivers 5% better accuracy than Deepgram and 10% better than Google STT. Self-hosting Whisper at ~$0.006/hour can save 30% compared to cloud-based STT services.
Process of Work
- Audit — analysis of current telephony and NLP requirements (1-2 days). Estimated cost: $500-$1,000.
- Design — selection of STT/LLM/TTS, WebSocket architecture, conversion, and jitter buffer parameters (3-5 days).
- Implementation — writing handler, CRM integration, monitoring setup, and VAD configuration (1-2 weeks).
- Testing — load testing with 100 call simulation, recognition accuracy checks, and DTMF handling (3-5 days).
- Deployment — server or cloud deployment, API documentation, and training session (2-3 days).
Approximate Timeline
Basic bot on Twilio with one scenario — from 2 weeks. Production solution with multilingual support and monitoring — up to 2 months. Cost is calculated individually, depending on call volume and NLP complexity. Typical integration costs range from $5,000 to $20,000, plus ongoing Twilio fees (~$0.0045/min) and AI service fees (e.g., Whisper self-host ~$0.006/hour). Twilio, Media Streams API — official documentation.
Deliverables (Что входит в работу)
- TwiML and WebSocket handler configuration
- Audio format conversion (μ-law ↔ PCM 16kHz) with anti-aliasing
- STT/TTS and LLM integration (cloud or local)
- Real-time monitoring dashboard with p99 latency alerts
- API documentation and access credentials
- One training session for your team
- Two-week post-launch support with hotfix window
These deliverables include all configuration files, API keys, and documentation necessary for handoff.
Advantages and Contact
Over 5 years of experience in voice AI systems, 10+ Twilio Voice AI deployments for retail and logistics. We guarantee stability: p99 latency < 2.5 sec, uptime 99.9%. Certified Twilio and ML specialists.
Contact us for a project estimate within 1 day. Get a consultation and accurate timeline.







