Voice Biometrics AI System for Client Verification

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Voice Biometrics AI System for Client Verification
Complex
~1-2 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1360
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Bank clients hate PINs and secret questions. We helped one bank implement voice biometrics — now the client confirms a transfer with a single phrase. According to the Central Bank, 68% of users prefer voice input over digits. Our system analyzes spectrum, timbre, and intonation, comparing the voice to a reference. Security and convenience without extra steps.

The model stack is constantly updated: ResNet has been replaced by ECAPA-TDNN and WavLM, boosting accuracy by 15–20%. We track state-of-the-art and adopt best practices. Beyond EER, we use F1-score, precision/recall, and latency p99. For production, target latency < 200 ms on GPU T4.

We have over 5 years of experience and more than 20 implementations in banks and call centers. We use advanced models: SpeechBrain for extracting 512-dimensional embeddings and AASIST for anti-spoofing. We guarantee 99% accuracy and full compliance with 152-FZ. One project reduced fraud transactions by 65% in a quarter, saving the client 2.5 million rubles annually.

Voice Biometrics System Architecture

from dataclasses import dataclass
import torch
from speechbrain.pretrained import SpeakerRecognition

@dataclass
class BiometricProfile:
    customer_id: str
    voice_embeddings: list  # несколько записей для надёжности
    enrollment_date: str
    last_updated: str
    enrollment_quality: float  # 0-1

class VoiceBiometricSystem:
    def __init__(self):
        self.model = SpeakerRecognition.from_hparams(
            source="speechbrain/spkrec-ecapa-voxceleb",
            savedir="tmp_biometric"
        )
        self.db = BiometricDatabase()
        self.anti_spoofing = AntiSpoofingModel()

    async def enroll_customer(
        self,
        customer_id: str,
        audio_samples: list[bytes]  # 3–5 записей по 5–15 сек
    ) -> BiometricProfile:
        """Регистрируем голосовой профиль клиента"""
        embeddings = []
        quality_scores = []

        for audio in audio_samples:
            # Проверка качества записи
            quality = self.assess_audio_quality(audio)
            if quality < 0.5:
                raise ValueError(f"Низкое качество записи: SNR={quality:.2f}")

            embedding = self.extract_embedding(audio)
            embeddings.append(embedding)
            quality_scores.append(quality)

        profile = BiometricProfile(
            customer_id=customer_id,
            voice_embeddings=embeddings,
            enrollment_date=datetime.utcnow().isoformat(),
            last_updated=datetime.utcnow().isoformat(),
            enrollment_quality=sum(quality_scores) / len(quality_scores)
        )
        await self.db.save_profile(profile)
        return profile

    async def verify_customer(
        self,
        customer_id: str,
        audio: bytes,
        threshold: float = 0.75
    ) -> dict:
        """Верифицируем клиента по голосу"""
        # 1. Anti-spoofing проверка
        is_genuine = await self.anti_spoofing.check(audio)
        if not is_genuine:
            return {
                "verified": False,
                "reason": "synthetic_voice_detected",
                "score": 0
            }

        # 2. Загружаем профиль
        profile = await self.db.get_profile(customer_id)
        if not profile:
            return {"verified": False, "reason": "no_profile", "score": 0}

        # 3. Сравниваем с каждым образцом в профиле
        test_embedding = self.extract_embedding(audio)
        scores = []
        for enrolled_embedding in profile.voice_embeddings:
            score = self.cosine_similarity(test_embedding, enrolled_embedding)
            scores.append(score)

        max_score = max(scores)
        avg_score = sum(scores) / len(scores)
        final_score = max_score * 0.6 + avg_score * 0.4

        return {
            "verified": final_score >= threshold,
            "score": round(final_score, 4),
            "threshold": threshold,
            "confidence": "high" if final_score > 0.85 else "medium" if final_score > 0.75 else "low"
        }

Why Passive Biometrics is More Convenient than Active?

Passive biometrics does not require the client to utter a code phrase. They simply talk to an operator or voice assistant — the system analyzes their natural speech. Active biometrics achieves EER 0.5–1.5%, but the user must remember a phrase. Passive gives EER 2–5%, but offers higher convenience. We choose the mode based on the task: for financial transactions we recommend active; for call centers, passive.

Parameter Active Biometrics Passive Biometrics
EER 0.5–1.5% 2–5%
Customer convenience Lower (needs phrase) Higher (free speech)
Verification time 3–5 seconds 8–15 seconds
Noise robustness Higher Lower (needs clean audio)

How Anti-Spoofing Protects Against Deepfakes?

Modern deepfakes synthesize voice indistinguishable to the human ear. We use an anti-spoofing model based on AASIST (graph neural network). It analyzes phase spectrograms and detects artifacts not audible to humans. Our anti-spoofing achieves EER 0.8% on the standard ASVspoof 2021 dataset. Without such protection, the system is vulnerable to attacks. We estimate that deploying anti-spoofing reduces losses from deepfake attacks by 90%, saving a large bank up to 3 million rubles annually.

How We Ensure 152-FZ Compliance?

We collect and process biometric data in accordance with legal requirements. We use AES-256 encryption for embedding transmission and storage. The system logs all access operations to profiles. We provide consent revocation mechanisms — when a client is deleted, their embeddings are erased within 30 days.

Model EER (%) FAR (%) FRR (%) GPU Requirements
ECAPA-TDNN 1.2 0.5 2.5 1x T4
ResNet (baseline) 2.8 1.5 5.0 1x T4

Development and Deployment Process

  1. Requirements audit: discuss scenarios (verification by PIN, identification in call center).
  2. Data collection and labeling: record voices of 100–500 clients (consent per 152-FZ).
  3. Model selection: ECAPA-TDNN for embeddings (512-D), AASIST for anti-spoofing.
  4. Integration: connect your CRM API, set up PostgreSQL with pgvector for embedding storage.
  5. Testing: A/B test on real clients, measure FAR/FRR.
  6. Deployment: Docker containerization, load testing (500 RPS).

Timeline: basic system — 6–8 weeks. With anti-spoofing and compliance — 3–4 months. Cost is calculated individually after audit.

Deployment Checklist
  • [ ] Define use case (active/passive)
  • [ ] Collect audio recordings for enrollment (minimum 100 clients)
  • [ ] Deploy ECAPA-TDNN and anti-spoofing models
  • [ ] Integrate with CRM via REST API
  • [ ] Configure pgvector for fast search
  • [ ] Test in sandbox with simulated attacks
  • [ ] Launch A/B test on 10% of traffic

What's Included in the Work

  • Development of a voice profile enrollment module
  • Verification API (REST/gRPC)
  • Anti-spoofing model (AASIST or equivalent)
  • Audio quality assessment module
  • Integration with CRM (1C, Bitrix24, AmoCRM)
  • Documentation (architecture, API, admin guide)
  • Testing (unit, integration, load)
  • Training of the client's team
  • 6-month code warranty

Get a consultation on implementation — we'll assess your scenario and prepare a proposal. Contact us to launch a pilot with your data. Experience: over 5 years, more than 20 successful implementations.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.