Voice AI Development on Voximplant: VoxEngine & STT/TTS Integration

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Voice AI Development on Voximplant: VoxEngine & STT/TTS Integration
Medium
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Voice AI Development on Voximplant: VoxEngine & STT/TTS Integration

Outbound calling to 500 contacts — the bot answers with a 4-second delay, customers hang up. We encountered this when a real estate company was losing 30% of leads due to pauses. After integrating Voximplant with an AI backend, we reduced latency to 600 ms, and conversation conversion increased by 40%. One project saved the client 1.2 million rubles per year by reducing operator headcount.

The Russian platform Voximplant with VoxEngine gives full control over the call, but its integration with AI models requires fine-tuning of streaming audio, VAD, and queue management. Over 5 years of work, we have implemented more than 30 voice AI solutions on Voximplant. Our engineers are certified Voximplant Developers and know how to squeeze at least 200 ms latency at the STT stage.

What Problems We Solve

High response latency is the main pain. With sequential audio processing (call → recording → recognition → response), latency reaches 3–5 seconds. A pause of 1.5 seconds is already noticeable to the user. The solution is streaming processing via WebSocket: audio is transmitted in 30 ms frames, and voice VAD runs in parallel.

Unstable recognition with background noise or fast speakers. We use models with adaptive noise reduction (e.g., Silero VAD) and adjust sensitivity for each channel.

Scaling issues — VoxEngine scripts run in a single thread, and with 100+ simultaneous calls you easily hit limits. We design architecture with a message queue (RabbitMQ/NATS) and a pool of AI workers to distribute load.

How the Voximplant with AI Integration Works

The main pattern: VoxEngine opens a WebSocket connection to our Python backend for each call. The entire audio stream goes through it, controlled by events.

// VoxEngine scenario (JavaScript)
VoxEngine.addEventListener(AppEvents.CallAlerting, (e) => {
    const call = e.call;
    call.answer();

    // Create a WebSocket connection to our AI backend
    const wsConn = VoxEngine.createWSClient(`wss://api.yourapp.com/voxi-stream`);

    // Bind call audio to WebSocket
    call.sendMediaTo(wsConn);
    wsConn.sendMediaTo(call);

    wsConn.addEventListener(WSClientEvents.ConnectionClosed, () => {
        call.hangup();
    });

    call.addEventListener(CallEvents.Disconnected, () => {
        wsConn.close();
    });
});

On the backend, FastAPI receives the stream, accumulates audio in a buffer, detects the end of a phrase (VAD), and sends it to an ASR model. The result goes to TTS, and the synthesized speech is sent back.

from fastapi import FastAPI, WebSocket
import asyncio

@app.websocket("/voxi-stream")
async def voximplant_stream(websocket: WebSocket):
    await websocket.accept()
    session = VoiceSession()

    # Send greeting
    greeting_audio = await tts.synthesize("Hello! How can I help you?")
    await websocket.send_bytes(greeting_audio)

    audio_buffer = bytearray()
    silence_frames = 0

    async for chunk in websocket.iter_bytes():
        audio_buffer.extend(chunk)
        silence_frames = 0  # reset on audio receipt

        # Process every 1.5 seconds of accumulated audio
        if len(audio_buffer) >= 24000 * 2:  # 1.5 sec @ 16kHz 16-bit
            response = await process_utterance(bytes(audio_buffer), session)
            if response:
                audio_response = await tts.synthesize(response)
                await websocket.send_bytes(audio_response)
            audio_buffer = bytearray()

For more details on the protocol, see WebSocket. Contact us to discuss your scenario and get a demo of streaming processing.

What Using VoxEngine Provides

VoxEngine is not a proxy — it's a full JavaScript engine inside the call. It allows flexible media flow control: mixing audio, switching channels, adding DTMF. For AI scenarios, this means we can play pre-recorded messages (e.g., "Please wait") while the AI model processes the request — without delay.

Comparison VoxEngine + WebSocket REST API (request-response)
Latency (p99) 400–800 ms 2–5 s
Scaling 500+ simultaneous calls 50–100 calls
Infrastructure cost Moderate (one WebSocket server) High (depends on number of HTTP workers)
Dialog flexibility Real-time, interruptible Only sequential requests

The table shows that VoxEngine + WebSocket provides latency of 400–800 ms, which is 5–10 times faster than REST API.

Comparison of Popular ASR/TTS Models

Model Latency (p50) Accuracy (WER) Languages
Whisper (large) 300 ms 5% 99
Silero 150 ms 8% 2
Google STT 200 ms 6% 125
Yandex SpeechKit 250 ms 7% ru/en

Model selection depends on required accuracy and budget. For Russian, we often recommend Yandex SpeechKit or Whisper.

Why Voximplant Is Better for Voice AI

Voximplant documentation confirms: the platform was originally designed for real-time communications. Unlike ordinary SIP trunks, VoxEngine allows processing audio at the JavaScript level, achieving sub-second delays with proper integration.

How to Reduce Latency to 200 ms

Achieving 200 ms at the STT stage is only possible with a comprehensive approach:

  • Use VAD with a low threshold (Silero VAD with threshold 0.3)
  • Preload ASR model into GPU memory
  • Use streaming ASR (e.g., Whisper cpp in real-time mode)
  • Optimize audio frame size (30–50 ms)
More on VAD configuration VAD (Voice Activity Detection) is critical for reducing latency. We recommend using Silero VAD with threshold 0.3 and min_silence_duration_ms 150. This cuts pauses and avoids wasting time transmitting silence.

Outbound Calling: How to Automate Mass Campaigns

For outbound calls, Voximplant provides the StartScenarios API. We run campaigns with personalization: passing the client's name and context of previous interactions into the scenario.

import requests

def start_outbound_campaign(contacts: list[dict]):
    """Start mass outbound calling via Voximplant API"""
    for contact in contacts:
        response = requests.post(
            "https://api.voximplant.com/platform_api/StartScenarios/",
            data={
                "account_name": VOXI_ACCOUNT,
                "api_key": VOXI_API_KEY,
                "rule_name": "outbound_bot",
                "script_custom_data": json.dumps({
                    "phone": contact["phone"],
                    "customer_name": contact["name"],
                    "context": contact.get("context", {})
                }),
                "reference_to_call_id": contact["phone"]
            }
        )

Typical errors: not handling busy lines/unavailability — we add retry with exponential backoff and log hang-up reasons.

What the Work Includes

  • Audit of current telephony and AI scenario requirements.
  • Architecture design: ASR/TTS selection, VAD, WebSocket connection setup.
  • Development of VoxEngine scenario and Python backend (FastAPI, asyncio).
  • Integration with CRM and accounting systems (if needed).
  • Load testing (1000+ virtual calls).
  • Scenario and API documentation.
  • Operator training and 3 months of post-launch support.

Work Process

  1. Analytics — we study your dialogues, identify intents and slots.
  2. Prototype — in 5 days we make an MVP on one scenario (e.g., answers to frequent questions).
  3. Development — write VoxEngine code, connect selected ASR/TTS, configure VAD.
  4. Testing — run 100+ test calls, measure latency and recognition quality.
  5. Production launch — deploy production infrastructure, set up monitoring (Grafana, Alertmanager).
  6. Support — monitoring, model retraining, scenario refinement based on statistics.

Estimated Timelines

  • Basic integration (one scenario, out-of-the-box ASR/TTS) — from 1 to 2 weeks.
  • Full production with outbound calling, 3+ intents, CRM integration — from 1.5 to 2 months.
  • Cost is calculated individually, depending on dialog complexity and latency requirements.

Describe your task — we will prepare a commercial proposal with architecture and timelines within 1 day. We guarantee 99.9% SLA and a reduction in bot response time to 200 ms. Request a consultation now and get an analysis of your current telephony setup. For a quick start, contact us through the form on the website.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.