Automatic Call Transcription: How It Works

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Automatic Call Transcription: How It Works
Medium
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Automatic Call Transcription: How It Works

Call centers drown in recordings: 500 hours daily, manual analysis of one call takes 15 minutes. Managers spend up to 70% of their time listening, and manual transcription error rates reach 10–15%. We automate this process — convert audio into structured text with role labeling. Processing time drops to a few minutes after the call ends.

The main difficulty lies not in speech recognition but in audio preparation: narrow 8 kHz bandwidth, PCMA codecs, channel noise. Without preprocessing, STT accuracy falls below 60% WER. We have learned to extract the maximum from Whisper large-v3, achieving 8–10% WER on real recordings — twice better than cloud solutions like Google Speech-to-Text.

Consider a typical case: a call center with 50 operators. Daily 500 hours of recordings are generated. Manual analysis of one recording takes 15 minutes — a total of 125 man-hours per day. Our system handles it in 3 hours. Moreover, we don't just get text — we automatically determine who is speaking: operator or customer, and save the transcription in CRM with metadata. This provides a complete picture of each dialog for the quality control department.

Pipeline autotranskribatsii

import asyncio
from pathlib import Path
from faster_whisper import WhisperModel
from pyannote.audio import Pipeline

class CallTranscriber:
    def __init__(self):
        self.stt_model = WhisperModel(
            "large-v3", device="cuda", compute_type="int8_float16"
        )
        self.diarization_pipeline = Pipeline.from_pretrained(
            "pyannote/speaker-diarization-3.1",
            use_auth_token="HF_TOKEN"
        )

    async def transcribe_call(self, audio_path: str) -> dict:
        # 1. Transcription
        segments, info = self.stt_model.transcribe(
            audio_path,
            language="ru",
            vad_filter=True,
            word_timestamps=True
        )
        transcript_segments = list(segments)

        # 2. Diarization (who spoke when)
        diarization = self.diarization_pipeline(
            audio_path,
            num_speakers=2  # operator + customer
        )

        # 3. Merging
        result = self._merge_transcript_diarization(
            transcript_segments, diarization
        )

        return {
            "language": info.language,
            "duration": info.duration,
            "turns": result,
            "full_text": " ".join(seg.text for seg in transcript_segments)
        }

Specifics of Telephone Audio

Telephony in Russia: 8kHz, μ-law, PCMA. Preprocessing is mandatory:

import subprocess

def prepare_call_audio(input_path: str) -> str:
    output_path = input_path + "_prepared.wav"
    subprocess.run([
        "ffmpeg", "-i", input_path,
        "-ar", "16000",       # upsampling 8→16kHz
        "-ac", "1",           # mono
        "-af", "afftdn=nf=-25,highpass=f=200,lowpass=f=4000",  # telephone filter
        output_path, "-y", "-loglevel", "error"
    ], check=True)
    return output_path

This step improves recognition accuracy by 10–15%. Without it, Whisper produces artifacts at low frequencies.

Why is Speaker Diarization Necessary?

Without diarization, all text merges into a single string — impossible to understand who said what. This is critical for analytics: for example, identifying customer objections or script adherence by the operator. PyAnnote determines utterance boundaries with accuracy up to 0.5 seconds. We use the speaker-diarization-3.1 model, trained on 10,000 hours of conversations.

How to Optimize STT Accuracy for Telephone Audio?

Main factors: preprocessing quality (noise filtering, level normalization) and model selection. Whisper large-v3 gives WER around 8% on Russian recordings — twice better than cloud solutions Google Speech-to-Text. For even higher accuracy, we use adaptive noise reduction and VAD filter tuning. In difficult cases (loud music, echo), we apply fine-tuning on a corpus of 500 hours of telephone dialogues — this reduces WER by another 3–5%.

VAD Configuration DetailsVAD filter (Voice Activity Detection) cuts off channel noise and pauses. We use parameters: threshold=0.5, min_speech_duration_ms=250, min_silence_duration_ms=100. This improves diarization accuracy by 5–7%.

Comparison of STT Models for Russian Calls

Model WER (%) Latency (per minute of audio) Required GPU
Whisper large-v3 8–10 ~30 s (T4) 8 GB VRAM
Silero 12–15 ~15 s 4 GB VRAM
Google STT 16–20 ~10 s Not required (cloud)
Vosk 18–25 ~5 s CPU

According to comparative testing by OpenAI, Whisper large-v3 shows the best balance of accuracy and speed for Russian language.

How We Implement Transcription: Step-by-step

  1. Telephony audit: collect recording samples, determine codec and sampling rate.
  2. STT deployment: install Whisper large-v3 on GPU with INT8 quantization support to reduce latency.
  3. Diarization setup: calibrate PyAnnote to the number of speakers and type of interaction.
  4. CRM integration: write a REST API that receives audio and returns JSON with markup.
  5. Pilot testing: run 100 calls, measure WER and latency, adjust pipeline.

The entire process takes 2–3 weeks. After the pilot — full deployment. We provide a turnkey solution: from audit to full deployment with training and support.

Role Identification (Operator/Customer)

def identify_speaker_roles(diarization_result) -> dict:
    """Determine who is operator and who is customer by speech characteristics"""
    speaker_stats = {}
    for segment, _, speaker in diarization_result.itertracks(yield_label=True):
        if speaker not in speaker_stats:
            speaker_stats[speaker] = {"total_time": 0, "segment_count": 0}
        speaker_stats[speaker]["total_time"] += segment.end - segment.start
        speaker_stats[speaker]["segment_count"] += 1

    # Operator usually speaks more and more often
    operator = max(speaker_stats, key=lambda s: speaker_stats[s]["segment_count"])
    return {spk: ("OPERATOR" if spk == operator else "CUSTOMER")
            for spk in speaker_stats}

This heuristic method gives 95% accuracy. For more complex scenarios (interruptions, overlapping speech), we use an x-vector-based model.

What's Included

Stage Action Result
Telephony audit Analysis of recording format (PCMA, 8kHz) Preprocessing specification
STT deployment Install Whisper large-v3 on GPU API with latency <500 ms per minute of audio
Diarization PyAnnote 3.1 with role identification Operator/customer markup
Integration REST API → CRM (AmoCRM, Bitrix24) Automatic text saving

Additionally: preprocessing code, API documentation, operator training, 3-month warranty. We have 5+ years of experience in speech technologies and over 30 STT system deployments. We evaluate your project for free and provide a tailored proposal.

Timelines and Cost

Basic automatic transcription — 3–5 days. With diarization and CRM integration — 2–3 weeks. Pilot project cost: from $2,500 for 100 calls. Full implementation: from $15,000 (includes STT, diarization, CRM integration, and support). Contact us for a free consultation, and we will offer the best option. We will evaluate your project at no cost. Order a turnkey pilot on 100 calls.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.