Voice AI Bot Automates 70% of Inbound Calls

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Voice AI Bot Automates 70% of Inbound Calls
Medium
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Customers waste time in queues, and operators burn out on repetitive questions. 35% of calls are about order status; 25% are FAQs. Each such call lasts 3–5 minutes, with 2 minutes spent on identification and data lookup. Operators spend up to 70% of their time on repeated requests, leading to high turnover and rising costs. We automate inbound call processing using a voice chatbot built on ASR (Yandex SpeechKit, Silero) → LLM (GPT-4o-mini, LLaMA 3) → TTS. The bot understands natural speech, extracts intent from the first sentence, and resolves 60–75% of tasks without escalation. The target containment rate is 70%, and for typical requests it reaches up to 95%. Our team has 5 years of experience in voice AI solutions and certifications on leading platforms. According to internal tests, fine-tuning on 500 examples increases containment rate by 15%. This solution reduces call center costs by up to 50%, saving $200,000 annually for a mid-size call center. It is 2 times more effective than traditional IVR and 3 times faster in resolving simple requests.

Typology of Inbound Inquiries

Type Share Automatability
Status check 35% 95%
Data changes 20% 75%
FAQ / information 25% 90%
Complaints 10% 20%
Urgent questions 10% 40%

How does the AI bot handle complex requests?

For compound requests, the bot uses a chain of intents. Context is stored in dialog memory based on DialogScenario. If the model is unsure, it asks a clarifying question rather than escalating immediately. This allows automation of 40% of inquiries that seem complex. For example, a customer asks "when will my parcel arrive and can I change the address?" — the bot uses RAG for calls to retrieve the order number from CRM, offers delivery options, and in 80% of cases the customer confirms the solution without transfer to an operator.

Why is context transfer to the operator critical?

Upon escalation, the bot passes structured context: intent, extracted entities, and a 2–3 sentence dialog summary. The operator sees this before connecting and does not need to ask again — call time is reduced by 50%. We use the AGENT_BRIEFING_TEMPLATE for formatting. This reduces operator load and speeds up problem resolution.

How we implement the voice AI bot

We build the bot on the ASR → LLM → TTS stack. For speech recognition we use Yandex SpeechKit or Silero, for response generation — GPT-4o-mini or LLaMA 3, for synthesis — the same TTS stack. Scenarios are described as graphs with transitions by intent and entities. Our speech analytics module monitors sentiment to adapt responses in real time.

Multi-scenario routing

from typing import Callable

@dataclass
class DialogScenario:
    name: str
    triggers: list[str]
    handler: Callable
    priority: int = 0

SCENARIOS = [
    DialogScenario(
        name="order_status",
        triggers=["order", "status", "where is my parcel", "when will it be delivered"],
        handler=handle_order_status_scenario,
        priority=10
    ),
    DialogScenario(
        name="account_balance",
        triggers=["balance", "remaining", "how much money", "account"],
        handler=handle_balance_scenario,
        priority=10
    ),
    DialogScenario(
        name="technical_issue",
        triggers=["not working", "error", "broken", "problem"],
        handler=handle_tech_support_scenario,
        priority=5
    ),
]

async def route_to_scenario(user_text: str) -> DialogScenario:
    response = await client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{
            "role": "system",
            "content": f"Determine the scenario from: {[s.name for s in SCENARIOS]}. JSON: {{'scenario': '...'}}"
        }, {"role": "user", "content": user_text}],
        response_format={"type": "json_object"}
    )
    scenario_name = json.loads(response.choices[0].message.content)["scenario"]
    return next((s for s in SCENARIOS if s.name == scenario_name), SCENARIOS[-1])

Context transfer to operator

async def transfer_to_agent(session: CallSession, reason: str):
    context = {
        "call_id": session.call_id,
        "phone": session.phone,
        "customer": await lookup_customer(session.phone),
        "dialog_summary": await summarize_dialog(session.history),
        "intent": session.current_intent,
        "escalation_reason": reason,
        "timestamp": datetime.utcnow().isoformat()
    }
    await crm.create_case(context)
    await telephony.transfer_call(session.call_id, agent_queue="support")

Operator notification before connection

AGENT_BRIEFING_TEMPLATE = """
Transferring client {phone}.
Reason for contact: {intent}.
Client has already reported: {summary}.
No need to ask: order number and name — already obtained.
"""

AI bot vs traditional IVR

Parameter Tone-based IVR AI bot
Request understanding Only digits and yes/no Natural speech, hesitations, compound requests
Containment rate ≤30% 60–75%
Scenario time Fixed menu Dynamic, context-aware
Dialog recording Only audio Audio + structured data
Learning from data No Few-shot, fine-tuning on call recordings

This AI bot resolves 2 times more inquiries than IVR and reduces call center costs up to 50%. We guarantee containment rate of at least 60% after two weeks of pilot.

What's included in the work

We provide the full development cycle:

  • Audit of current calls and identification of typical requests
  • Writing dialog scenarios for each type of request
  • Integration with telephony (Asterisk, MTT, Rostelecom) and CRM (AmoCRM, Bitrix24)
  • Model training on your data (fine-tuning, few-shot)
  • Testing on pilot traffic (minimum 500 calls)
  • Documentation of scenarios and API
  • Training operators on working with the bot
  • Support and refinement after launch

We guarantee quality: after implementation we conduct an audit and adjust scenarios to achieve the target containment rate.

Implementation process

  1. Analytics — analysis of call recordings, identification of typical scenarios
  2. Design — prototype scenarios on the chosen stack
  3. Development — scenario implementation, CRM integration, escalation setup
  4. Testing — A/B test with a pilot group, scenario adjustment
  5. Deployment — launch on production traffic, p99 latency monitoring

Timeline

A single scenario (e.g., order status) takes from 2 weeks. A full multi-scenario bot with integration takes from 2 to 3 months. Timeline is determined after an audit.

Want to see the result on your own calls? Get a consultation: we will analyze 50–100 recordings, prepare a scenario prototype, and show the containment rate on your data. Order a pilot — 500 calls in test mode. Contact us for a free audit. See the effectiveness of the AI bot in practice.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.