We often see customers losing patience while navigating 4–7 layers of touch-tone menus. "Press 1 for... press 2..." — such DTMF IVR irritates and filters out up to 40% of calls. Studies show 60% of users prefer a voice interface. Our approach is AI-IVR: the system understands free speech, determines intent in 1–2 questions, and routes the call without buttons. We have been implementing such solutions for the last 5 years, and customers report a 30–50% reduction in call handling time and a 20 percentage point increase in satisfaction. This is not just replacing buttons — it is a shift from a rigid decision tree to an adaptive dialogue based on LLM.
Problems AI-IVR Solves
Traditional DTMF menus struggle with non-standard requests, are difficult to update, and require memorizing sequences. AI-IVR based on NLP and ASR replaces the rigid decision tree with flexible dialogue. The customer says "I have a payment issue" — the system itself understands it needs to go to billing and transfers.
Comparison of DTMF-IVR and AI-IVR
| Parameter | DTMF-IVR | AI-IVR |
|---|---|---|
| Navigation | 4–7 levels, buttons | 1–2 questions, speech |
| Scenario coverage | Limited by tree | Unlimited (LLM) |
| Updating | Difficult, requires development | Prompt file |
| For non-mobile users | Problematic | Normal |
| Development cost | Low | Medium (pays back in 3–6 months) |
Important: AI-IVR does not require full infrastructure replacement — we integrate it on top of your existing PBX via SIP or API.
How AI-IVR Understands Customer Intent
At its core is an LLM (e.g., GPT-4o or LLaMA 3) that receives the speech transcription (ASR) and determines intent. We use few-shot prompting: we pass a list of available destinations and example phrases in the prompt. The model returns a JSON with destination and confidence. If confidence is below 0.75, the system asks a clarifying question. This approach handles 95% of requests without involving a human agent.
class AIIVR:
def __init__(self, routing_config: dict):
self.destinations = routing_config["destinations"]
self.llm = AsyncOpenAI()
async def handle_call(self, call: IncomingCall) -> str:
"""Process incoming call — return destination"""
# Greeting
await call.say(
"Hello! You have reached {company}. How can I help you?"
)
# Listen for intent (up to 8 seconds)
user_input = await call.listen(timeout_sec=8, silence_threshold_ms=800)
if not user_input:
return await self.handle_silence(call)
# Recognize intent and route
route = await self.recognize_intent(user_input)
if route["confidence"] >= 0.75:
return await self.route_call(call, route["destination"])
else:
return await self.clarify_intent(call, user_input)
async def recognize_intent(self, user_input: str) -> dict:
destinations_description = "\n".join(
f"- {d['id']}: {d['description']}"
for d in self.destinations
)
response = await self.llm.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "system",
"content": f"""Determine where to route the call.\nAvailable destinations:\n{destinations_description}\nReturn JSON: {{"destination": "id", "confidence": 0.0-1.0}}"""
}, {"role": "user", "content": user_input}],
response_format={"type": "json_object"}
)
return json.loads(response.choices[0].message.content)
async def route_call(self, call: IncomingCall, destination: str) -> str:
dest = next(d for d in self.destinations if d["id"] == destination)
# Confirm routing
await call.say(dest.get("routing_message",
f"Transferring you to {dest['name']}..."))
if dest["type"] == "queue":
await call.transfer_to_queue(dest["queue_id"])
elif dest["type"] == "extension":
await call.transfer(dest["extension"])
elif dest["type"] == "bot":
await call.transfer_to_bot(dest["bot_id"])
return destination
Destination Configuration (YAML/JSON)
Routing paths are defined in a simple config — adding a new destination can be done without redeployment.
destinations:
- id: technical_support
name: "Technical Support"
description: "Problems with service, errors, not working"
type: queue
queue_id: tech_support_q
routing_message: "Connecting to technical support. Please wait."
- id: billing
name: "Billing and Invoices"
description: "Payment issues, invoices, debts, tariffs"
type: bot
bot_id: billing_bot
- id: sales
name: "Sales and New Connections"
description: "Subscribe to service, new contract, tariffs"
type: queue
queue_id: sales_q
What Business Metrics Does AI-IVR Improve?
According to Gartner, implementing intelligent IVR reduces load on first-line support by 40–60%. AI-IVR handles up to 80% of calls without agent involvement — 3 times more than DTMF. Average call time is cut in half. Savings on agents can reach 50% at scale. By reducing load, you can reassign staff to more complex tasks.
How AI-IVR Integrates with Your Existing PBX
Integration is done via SIP trunk or REST API. We connect to Asterisk, 1C-Bitrix24, Cisco, Genesys, and others without full infrastructure replacement. The system acts as an intermediary layer: it accepts the call, processes the dialogue, and hands control back to the PBX. For on-premise, we use vLLM with INT4 quantization — this reduces GPU costs by up to 4x compared to FP16.
Why Latency Is Critical for AI-IVR
If ASR + LLM takes more than 3 seconds, the user hangs up. We use streaming ASR (e.g., real-time Whisper) and intent caching for frequent requests. Typical p99 latency is 1.5–2 seconds. To reduce delays, we use model quantization and inference on Triton Inference Server.
Process of Work
We implement AI-IVR in several stages:
| Stage | Duration |
|---|---|
| Scenario analysis and data collection | 1–2 weeks |
| Dialogue design (prompt engineering) | 1 week |
| Development and PBX integration | 2–3 weeks |
| Testing (A/B, load) | 1–2 weeks |
| Deployment and monitoring | 1 week |
Total time for a typical project is 4–6 weeks. For a pilot with 3–5 destinations, 2–3 weeks.
What's Included
- Documentation: architecture diagram, prompt descriptions, instructions for adding destinations.
- Access: to logging and monitoring systems (Grafana, ELK), to API for self-updating configs.
- Training: a 4-hour workshop for the team on managing AI-IVR and fine-tuning new scenarios.
- Support: 2 weeks post-launch with daily standups and priority bug fixes.
Common Implementation Mistakes
- Too many destinations in the prompt (more than 10) — reduces accuracy. Optimal is 5–7.
- Lack of a fallback scenario for long silences — the system should re-ask or transfer to an agent.
- Ignoring latency: if ASR + LLM takes more than 3 seconds, the user hangs up. We use streaming ASR and intent caching.
- Poor handling of ambiguous requests — without a clarifying question, the customer may end up in the wrong place.
Request a demonstration of AI-IVR for your scenario — we will assess the project in one day and offer a solution with a result guarantee. Get a consultation from an engineer: our specialists will help you choose the optimal architecture.







