Voice AI Bot for Call Center Implementation
A voice bot for a call center is not a simple IVR replacement but a full-fledged AI assistant based on an LLM. We've repeatedly encountered situations where clients deployed cheap voice robots and got a Containment Rate of 20–30%. Users got frustrated and demanded a human operator. Our approach: fine-tuning the language model on your dialogue base plus two-level escalation. Containment Rate >60% is a typical result within 4 weeks after deployment. Our team has 10+ years of experience in NLP and MLOps, having implemented 50+ voice bots for call centers. Savings on operators can reach 1.5 million rubles per year at a volume of 10,000 calls per month.
How a Voice AI Bot Solves the Problem of Operator Overload?
A standard call center spends 60% of operator time on routine requests: order status, address change, service booking. A voice bot offloads 70–80% of these categories. We have experience integrating with Bitrix24, amoCRM, and Salesforce: the bot pulls customer data by phone number, executes the scenario, and escalates with full context if needed. Containment Rate >60% is 2–3 times better than a typical IVR.
Typical Scenarios and Coverage
| Scenario |
Share of Calls |
Automation |
| Order status |
30–40% |
95% |
| Address change |
10–15% |
80% |
| Product FAQ |
15–20% |
85% |
| Complaints |
5–10% |
30% (then operator) |
| Booking/cancellation |
10–15% |
90% |
Example payback calculation
At 10,000 calls per month and CR 60%, operator savings — 1.5 million rubles/year. Implementation cost pays back in 4–6 months.
Bot Architecture for Call Center
from enum import Enum
from dataclasses import dataclass
class DialogState(Enum):
GREETING = "greeting"
INTENT_RECOGNITION = "intent_recognition"
COLLECTING_DATA = "collecting_data"
PROCESSING = "processing"
CONFIRMATION = "confirmation"
TRANSFER_TO_AGENT = "transfer_to_agent"
FAREWELL = "farewell"
@dataclass
class CallSession:
call_id: str
phone_number: str
state: DialogState = DialogState.GREETING
intent: str = None
collected: dict = None
retry_count: int = 0
max_retries: int = 3
def should_transfer(self) -> bool:
return (self.retry_count >= self.max_retries or
self.intent in ["complaint", "complex_issue"])
The first escalation level — a repeated request with the same intent, the second — transfer to an operator with full dialogue history.
Intent Recognition with Examples
async def recognize_intent(user_text: str) -> dict:
response = await client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "system",
"content": """Determine the customer's intent. Return JSON:
{"intent": "order_status|change_address|cancel_order|complaint|other",
"entities": {"order_id": "...", "address": "..."}}"""
}, {
"role": "user",
"content": user_text
}],
response_format={"type": "json_object"}
)
return json.loads(response.choices[0].message.content)
Escalation Conditions to Operator
ESCALATION_TRIGGERS = [
"operator", "live person", "connect me to a human",
"don't understand", "useless", "complaint", "claim",
"refund", "court", "consumer protection"
]
def should_escalate(text: str, session: CallSession) -> bool:
text_lower = text.lower()
if any(trigger in text_lower for trigger in ESCALATION_TRIGGERS):
return True
if session.retry_count >= 2:
return True
return False
CRM Integration
async def lookup_customer(phone: str) -> dict | None:
# Query CRM (Bitrix24, amoCRM, Salesforce)
async with aiohttp.ClientSession() as session:
resp = await session.get(
f"{CRM_API_URL}/contacts/search",
params={"phone": phone},
headers={"Authorization": f"Bearer {CRM_TOKEN}"}
)
data = await resp.json()
return data.get("contact")
Why Containment Rate Is the Key KPI?
Containment Rate is the percentage of calls fully handled by the bot without transfer to an operator. For a call center, this is direct savings: each percentage point increase reduces cost per call. At CR >60%, the bot pays for itself in 4–8 months depending on volume. We guarantee achieving CR >60% for typical scenarios after model calibration.
What Is Included in Turnkey Work?
| Stage |
Duration |
Result |
| Dialogue audit |
1 week |
Top-5 scenarios identified |
| Model fine-tuning |
2–3 weeks |
Model with >90% accuracy |
| CRM integration |
1–2 weeks |
Ready API connector |
| Escalation setup |
1 week |
Two-level scheme |
| Monitoring and training |
2 weeks |
Dashboards and operator training |
After launch — 3 months of support. We will evaluate your project for free: just contact us. Get a consultation on your scenario — we'll consider integration, call volume, and target CR.
ASR and TTS: Speech Recognition and Synthesis for Voice Bot
A voice bot consists of three components: ASR (speech recognition), NLU (intent understanding), TTS (response synthesis). The quality of each is critical to the final conversion.
ASR (Automatic Speech Recognition): We use Whisper large-v3 or Yandex SpeechKit depending on latency and data privacy requirements. Whisper gives WER below 5% on clean speech; Yandex handles regional accents better. Streaming ASR mode — first token in 300–500 ms — reduces perceived delay.
TTS (Text-to-Speech): Response synthesis takes 200–400 ms per phrase. We use SSML markup to control pauses and intonation: <break time="0.3s"/> before key words increases perceived quality by 20–30%.
Noise reduction: For incoming calls with background noise, we apply RNNoise (open-source) or Krisp API. This reduces WER on noisy channels from 25% to 8%.
Integrating ASR + TTS adds 500–800 ms to the total cycle latency. We optimize each component in tandem, achieving an overall bot response time of less than 2 seconds.
Timelines: MVP with 3–5 scenarios — 4–6 weeks. Full system with analytics and A/B testing — 3–4 months.
Speech Recognition and Synthesis: ASR, TTS, Voice Cloning
We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.
How does speech recognition ASR handle noisy call center recordings?
The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.
Typical problems we encounter
WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.
Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.
Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.
How do we solve it in practice?
The basic stack for a production pipeline:
-
ASR:
openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
-
Diarization:
pyannote.audio 3.x + integration via whisperx for word-level alignment
-
TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
-
Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2
A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=True → pyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.
Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.
How to fine-tune Whisper on domain data?
When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.
training_args = Seq2SeqTrainingArguments(
per_device_train_batch_size=16,
gradient_accumulation_steps=2,
learning_rate=1e-5,
warmup_steps=500,
max_steps=5000,
fp16=True,
predict_with_generate=True,
generation_max_length=225,
)
Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.
| Model |
WER (clean) |
WER (noisy) |
RTF (A10G) |
Languages |
| Whisper large-v3 |
5.2% |
27% |
0.08 |
99 |
| Wav2Vec2-XLSR-53 |
6.8% |
32% |
0.12 |
143 |
| Google STT (cloud) |
7.0% |
28% |
– |
125 |
| DeepSpeech 0.9.3 |
11.5% |
41% |
0.06 |
8 |
Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.
Speech synthesis: How to choose a model for your task?
| Model |
Latency (TTFB) |
Naturalness MOS |
Cloning |
Languages |
| XTTS v2 |
1.2–2.0 s |
4.1–4.3 |
Yes, 3 s reference |
17 |
| StyleTTS2 |
0.3–0.6 s |
4.0–4.2 |
Yes, requires adaptation |
en, + fine-tune |
| Kokoro-82M |
0.08–0.15 s |
3.7–3.9 |
No |
en, ja |
| Silero TTS |
0.05–0.1 s |
3.4–3.6 |
No |
ru, en, de, etc. |
| Edge-TTS |
~0.4 s (cloud) |
4.0 |
No |
100+ |
For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.
Our process and deliverables
We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.
Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.
What you get after engagement:
- Trained model with model card and evaluation report
- Docker image with optimized inference pipeline
- API documentation and integration examples
- Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
- 30-day post-deployment support and hotfixing
Timelines depend on complexity:
- Basic integration of a ready model — 1–2 weeks
- Fine-tuning with data preparation and validation — 4–8 weeks
- Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months
Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.
Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.
Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.