Voice AI Telephony Integration with Twilio, NLU, and TTS

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Voice AI Telephony Integration with Twilio, NLU, and TTS
Medium
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Voice AI Telephony: Twilio NLU and TTS Integration

Upon initiation of a client call, the telephony system encounters challenges: the speech recognition engine incorrectly transcribes "I want to order" due to audio conversion artifacts arising from μ-law 8 kHz to PCM 16 kHz, degrading STT accuracy by 30%. We integrate Twilio Voice AI with real NLU, employing Whisper large-v3 for recognition, GPT-4o for response generation, and ElevenLabs for speech synthesis. Consequently, the bot comprehends the client even with an accent and responds without filler phrases. Order integration — we resolve latency and recognition quality issues.

Problems We Solve

Audio format conversion — Twilio transmits μ-law 8 kHz, whereas Whisper requires PCM 16 kHz. Conversion errors introduce artifacts and degrade recognition quality. We utilize audioop.ratecv with anti-aliasing and cross-fade smoothing to eliminate clicks.

WebSocket connection reliability — disconnection results in loss of the audio stream. We implement a reconnection mechanism with buffering of the last second, along with a jitter buffer for smooth playback.

Latency management — total latency must not exceed 2 seconds. We optimize the pipeline via parallel STT and response generation and caching of frequent queries. Comparison: our pipeline reduces latency by a factor of 2 compared to sequential processing.

Technical Implementation

TwiML webhook for incoming call

from fastapi import FastAPI, Request
from twilio.twiml.voice_response import VoiceResponse, Start, Stream, Say

app = FastAPI()

@app.post("/incoming-call")
async def handle_incoming_call(request: Request):
    response = VoiceResponse()

    # Start Media Stream
    start = Start()
    start.stream(
        url=STREAM_ENDPOINT,
        track="both_tracks"  # incoming and outgoing audio
    )
    response.append(start)

    # Play greeting
    response.say(
        "Hello! I am a voice assistant. How can I help?",
        voice="alice",
        language="en-US"
    )
    response.pause(length=30)
    return Response(content=str(response), media_type="text/xml")

WebSocket handler for Media Streams

import asyncio
import json
import base64
from fastapi import WebSocket

@app.websocket("/stream")
async def handle_stream(websocket: WebSocket):
    await websocket.accept()
    call_sid = None
    stream_sid = None
    audio_buffer = bytearray()

    try:
        async for message in websocket.iter_text():
            data = json.loads(message)
            event = data.get("event")

            if event == "start":
                call_sid = data["start"]["callSid"]
                stream_sid = data["start"]["streamSid"]
                session = create_session(call_sid)

            elif event == "media":
                # Twilio uses mulaw 8kHz
                mulaw_audio = base64.b64decode(data["media"]["payload"])
                audio_buffer.extend(mulaw_audio)

                # Process when 2 seconds accumulated (16000 bytes @ 8kHz)
                if len(audio_buffer) >= 16000:
                    await process_audio_chunk(
                        bytes(audio_buffer), websocket, stream_sid, session
                    )
                    audio_buffer = bytearray()

            elif event == "stop":
                break

    except Exception as e:
        logger.error(f"Stream error: {e}")

async def send_audio_to_caller(websocket: WebSocket, stream_sid: str, audio_bytes: bytes):
    """Send synthesized audio back to the call"""
    encoded = base64.b64encode(audio_bytes).decode()
    await websocket.send_json({
        "event": "media",
        "streamSid": stream_sid,
        "media": {
            "payload": encoded
        }
    })

Audio format conversion

Twilio uses μ-law 8 kHz. Whisper works with PCM 16 kHz:

import audioop

def mulaw_to_pcm16k(mulaw_bytes: bytes) -> bytes:
    """μ-law 8kHz → PCM 16-bit 8kHz → upsample to 16kHz using anti-aliasing"""
    pcm_8k = audioop.ulaw2lin(mulaw_bytes, 2)  # μ-law → PCM 16-bit
    pcm_16k, _ = audioop.ratecv(pcm_8k, 2, 1, 8000, 16000, None)  # 8→16kHz
    return pcm_16k

How Twilio Voice AI processes audio in real time?

The Media Streams API transmits audio in 20 ms chunks. We accumulate a buffer of up to 2 seconds (16000 bytes at 8 kHz) and send it to STT. This reduces the number of requests and improves accuracy through context. After recognition, the LLM generates a response, TTS synthesizes speech, and the audio is sent back through the same WebSocket.

Why is correct audio format conversion important?

Conversion errors μ-law → PCM can introduce noise or shift the sampling frequency, leading to up to 30% loss in STT accuracy. We use audioop.ulaw2lin with explicit bit depth and ratecv with a quality filter. We also apply cross-fade smoothing at chunk boundaries to eliminate clicks.

Common conversion errors and their solutions
  • Ignoring bit depth: μ-law 8-bit → PCM 16-bit. Without ulaw2lin you get 8-bit PCM, STT won't understand.
  • Wrong rate: upsample from 8 kHz to 16 kHz requires interpolation. ratecv with None uses linear interpolation; for better quality, use cubic interpolation.
  • Artifacts during batch processing: clicks occur at chunk boundaries. We add cross-fade smoothing of 50 ms duration.

TTS approach comparison

Parameter ElevenLabs (cloud) Kokoro (ONNX local)
Latency 300-500 ms 100-200 ms
Quality Very high Medium
Cost Per character ($0.0003/char) Free (CPU/GPU)
Voices 100+ 10+

For production, we recommend a combination: ElevenLabs for primary dialogue, Kokoro for fallback under load. ElevenLabs costs approximately $0.0003 per character, while Kokoro is free, providing substantial savings for high-volume scenarios.

STT solution comparison

Parameter Whisper large-v3 Deepgram Nova-2 Google STT
Latency 200-400 ms 150-300 ms 300-600 ms
Accuracy (Russian) 95% 93% 90%
Price per hour $0.006 (Self-host) $0.004 $0.006
Accent adaptation High Medium Medium

For Russian-language scenarios, Whisper large-v3 delivers 5% better accuracy than Deepgram and 10% better than Google STT. Self-hosting Whisper at ~$0.006/hour can save 30% compared to cloud-based STT services.

Process of Work

  1. Audit — analysis of current telephony and NLP requirements (1-2 days). Estimated cost: $500-$1,000.
  2. Design — selection of STT/LLM/TTS, WebSocket architecture, conversion, and jitter buffer parameters (3-5 days).
  3. Implementation — writing handler, CRM integration, monitoring setup, and VAD configuration (1-2 weeks).
  4. Testing — load testing with 100 call simulation, recognition accuracy checks, and DTMF handling (3-5 days).
  5. Deployment — server or cloud deployment, API documentation, and training session (2-3 days).

Approximate Timeline

Basic bot on Twilio with one scenario — from 2 weeks. Production solution with multilingual support and monitoring — up to 2 months. Cost is calculated individually, depending on call volume and NLP complexity. Typical integration costs range from $5,000 to $20,000, plus ongoing Twilio fees (~$0.0045/min) and AI service fees (e.g., Whisper self-host ~$0.006/hour). Twilio, Media Streams API — official documentation.

Deliverables (Что входит в работу)

  • TwiML and WebSocket handler configuration
  • Audio format conversion (μ-law ↔ PCM 16kHz) with anti-aliasing
  • STT/TTS and LLM integration (cloud or local)
  • Real-time monitoring dashboard with p99 latency alerts
  • API documentation and access credentials
  • One training session for your team
  • Two-week post-launch support with hotfix window

These deliverables include all configuration files, API keys, and documentation necessary for handoff.

Advantages and Contact

Over 5 years of experience in voice AI systems, 10+ Twilio Voice AI deployments for retail and logistics. We guarantee stability: p99 latency < 2.5 sec, uptime 99.9%. Certified Twilio and ML specialists.

Contact us for a project estimate within 1 day. Get a consultation and accurate timeline.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.