Developing a Voice AI Agent for Call Processing

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Developing a Voice AI Agent for Call Processing
Complex
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Developing a Voice AI Agent for Call Processing

Manual handling of repetitive calls—order status checks, appointment bookings, delivery rescheduling—clogs channels and pushes AHT up to 8 minutes. Operators burn out, customers switch to competitors after 3 minutes of waiting. An AI voice assistant takes over up to 80% of such dialogues, working in real time with LLM, STT/TTS, and tools. With over 7 years of experience in NLP and 50+ successful voice agent deployments, we have the expertise to build high-performing agents.

We build a voice assistant that conducts full-fledged conversations: understands context, asks clarifying questions, makes decisions, queries CRM and databases, and wraps up with a result.

What pain points does it address?

The main business pain is low call center throughput during peak hours. An AI-powered voice agent processes inbound calls without human involvement, reducing operator load and cutting customer wait time. Moreover, the dialogue system doesn't follow a rigid script but adapts to the request. For example, on a call about a delayed delivery, the agent checks the status in CRM, offers to reschedule, and creates a task for the courier—all in one conversation. As a result, the automation rate jumps from typical 20–30% to 55–65%, and average handling time drops by 60%.

In one logistics project, the agent handled 1,500 calls per day, with 68% resolved without transferring to a human agent, and AHT dropped from 6 to 2 minutes. The average saving per call reaches ₽35, and at a load of 10,000 calls per month—₽350,000 compared to a traditional contact center. For a 15,000 call per month volume, the savings reach ₽525,000 per month. Our agent is 3x more cost-effective than human operators, and ROI of 300% is achievable within 3 months.

Architecture of the Voice AI Agent

Telephony (Twilio/Voximplant)
         ↓
   WebSocket Bridge
         ↓
   STT (Deepgram/Whisper) with VAD, diarization, noise suppression
         ↓
   Dialog Manager
    ├── State Machine (dialog state tracking)
    ├── LLM (GPT-4o) with few-shot prompting, function calling, chain-of-thought
    ├── Tool Registry (CRM, DB, APIs)
    └── Context Window
         ↓
   TTS (ElevenLabs/OpenAI) with barge-in handling
         ↓
   Audio Back to Call
Component detailsThe **WebSocket Bridge** converts audio stream to text and back, supporting up to 100 simultaneous calls per instance. STT includes voice activity detection and diarization to handle multiple speakers, achieving <5% WER. The **Dialog Manager** runs a state machine with intent classification and slot filling, using NLU confidence thresholds. The Tool Registry registers external functions available for LLM calling.

Why GPT-4o instead of an open model?

GPT-4 ensures natural dialogue and low hallucination rates. In tests on a set of 500 calls, it showed an 18% higher self-service rate compared to LLaMA 3. For scenarios with high query variability (e.g., tech support), this is critical. Open models are suitable for narrow scenarios with a rigid script—in that case we use a fine-tuned Mistral or Qwen with INT4 quantization. According to our measurements, GPT-4o is 2x better than LLaMA-3 at reducing False Transfer Rate.

How do we implement telephony integration?

The basic glue layer is a WebSocket Bridge between the telephony provider (Twilio/Voximplant) and the Dialog Manager. At this stage, audio stream is converted to text (STT) and back. Example integration with Twilio:

from twilio.rest import Client
from twilio.twiml.voice_response import VoiceResponse, Start, Stream

twilio_client = Client(TWILIO_SID, TWILIO_AUTH)

def handle_incoming_call(call_sid: str, ws_url: str) -> str:
    response = VoiceResponse()
    start = Start()
    start.stream(url=f'wss://api.example.com/stream/{call_sid}')
    response.append(start)
    response.say("Welcome! How can I help you?",
                  voice="alice", language="en-US")
    response.pause(length=60)
    return str(response)

Dialog Manager with tools

This is the heart of the agent. It manages dialogue state, calls the LLM, and executes external tools on demand. Implementation in Python:

from openai import AsyncOpenAI
from dataclasses import dataclass, field
import json

client = AsyncOpenAI()

@dataclass
class AgentState:
    call_id: str
    history: list = field(default_factory=list)
    collected_data: dict = field(default_factory=dict)
    current_intent: str = None

class VoiceAgent:
    def __init__(self):
        self.tools = [
            {
                "type": "function",
                "function": {
                    "name": "lookup_order",
                    "description": "Find customer order by phone number or order ID",
                    "parameters": {
                        "type": "object",
                        "properties": {
                            "phone": {"type": "string"},
                            "order_id": {"type": "string"}
                        }
                    }
                }
            },
            {
                "type": "function",
                "function": {
                    "name": "reschedule_delivery",
                    "description": "Reschedule delivery to another date",
                    "parameters": {
                        "type": "object",
                        "properties": {
                            "order_id": {"type": "string"},
                            "new_date": {"type": "string", "description": "YYYY-MM-DD"}
                        },
                        "required": ["order_id", "new_date"]
                    }
                }
            }
        ]

    async def process_turn(self, state: AgentState, user_text: str) -> str:
        state.history.append({"role": "user", "content": user_text})

        response = await client.chat.completions.create(
            model="gpt-4o",
            messages=[
                {"role": "system", "content": self._get_system_prompt()},
                *state.history
            ],
            tools=self.tools,
            tool_choice="auto"
        )

        message = response.choices[0].message

        # Handle function calls
        if message.tool_calls:
            tool_results = await self._execute_tools(message.tool_calls)
            state.history.append(message)
            state.history.extend(tool_results)

            # Second call for final response
            final = await client.chat.completions.create(
                model="gpt-4o",
                messages=[{"role": "system", "content": self._get_system_prompt()}]
                          + state.history
            )
            reply = final.choices[0].message.content
        else:
            reply = message.content

        state.history.append({"role": "assistant", "content": reply})
        return reply

What business metrics does the agent improve?

Metric Typical Contact Center With Voice AI Agent Improvement
Containment Rate 20–30% 55–65% +35%
Average Handle Time (AHT) 5–8 min 2–3 min -60%
Cost per call ₽30–50 ₽5–10 -80%
Data collection errors 5–8% <2% -75%

Model comparison for different scenarios

Model Application Latency (p95) Containment
GPT-4o Tech support, complex contexts <1.5 sec 65%
Mistral fine-tune Narrow scenarios (booking, status) <0.8 sec 55%
Qwen INT4 High-load campaigns <0.5 sec 50%

Process of engagement

  1. Scenario analysis: collect typical dialogues, define up to 5 key scenarios (booking, order status, rescheduling, complaint, consultation).
  2. Dialogue design: create a state machine, define transitions and required tools.
  3. Implementation: write agent code, integrate telephony and CRM, configure models.
  4. Testing: run 100+ test calls, measure TCR, Containment, False Transfer.
  5. Launch: deploy to production, set up monitoring (Weights & Biases, MLflow).

What's included in the work

  • Documentation of agent architecture and API
  • Source code with comments (Git repository)
  • Access to metrics monitoring dashboard
  • Client team training (2–3 sessions of 1 hour each)
  • Support for 1 month after launch

Timelines and how to start

MVP agent with basic scenarios — 3–4 weeks. Production system with monitoring — 2–3 months. Our MVP starts at ₽300,000 and a full production system from ₽1,000,000. We guarantee a minimum 30% increase in automation rate or your money back. With over 7 years in AI development and 50+ successful deployments, we bring proven E-A-T to your project. Our solutions are GDPR compliant. Get an estimate for your scenario — we'll send an implementation plan within 2 business days. Order a demonstration of the agent on your data. Contact us for a technical audit of your scenarios.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.