Every third response from your voice assistant sounds unnatural—trembling timbre, missing phonemes. We solve this by fine-tuning a TTS model on the client's voice. After fine-tuning on 30–60 minutes of recordings, the model steadily reads any text: MOS rises to 4.3+ from 3.8 in zero-shot, and reverse recognition WER drops by 5–10%. Result: the assistant stops 'stuttering' even on complex queries.
Why fine-tuning over zero-shot?
Zero-shot cloning (e.g., XTTSv2 in speaker encoder mode) gives acceptable results but suffers from timbre trembling, artifacts on rare phonemes, and instability on long texts. Fine-tuning on 30–60 minutes of the target voice locks in the speaker's acoustic space, reduces reverse recognition WER by 5–10%, and increases UTMOS by 0.3–0.5. Main advantages: predictable quality on any input, ability to augment data (noise, reverberation), and control over intonation via conditioning.
What goes into dataset preparation for TTS fine-tuning?
Minimum volume: 30 minutes of clean recordings. Optimal: 1–2 hours. Audio requirements: sampling rate 22050 or 24000 Hz, signal level –18…–12 dBFS, signal-to-noise ratio >30 dB, clip lengths 3–15 seconds.
Preparation steps:
- Record in a studio or quiet room (check for background noise).
- Clean noise: use HPSS filter or spectral subtraction.
- Sentence-level segmentation: force alignment with Montreal Forced Aligner.
- Validate duration and quality: run through a validation script.
Example dataset validation script
import pandas as pd
from pathlib import Path
import soundfile as sf
import numpy as np
def validate_dataset(dataset_dir: str) -> dict:
"""Check dataset before training"""
metadata = pd.read_csv(f"{dataset_dir}/metadata.csv",
sep="|", names=["file", "text"])
stats = {
"total_files": len(metadata),
"total_duration": 0,
"errors": []
}
for _, row in metadata.iterrows():
wav_path = f"{dataset_dir}/wavs/{row['file']}.wav"
if not Path(wav_path).exists():
stats["errors"].append(f"Missing: {wav_path}")
continue
audio, sr = sf.read(wav_path)
duration = len(audio) / sr
stats["total_duration"] += duration
if sr != 22050:
stats["errors"].append(f"Wrong SR {sr}: {wav_path}")
if duration < 1.0 or duration > 15.0:
stats["errors"].append(f"Bad duration {duration:.1f}s: {wav_path}")
stats["total_duration_min"] = stats["total_duration"] / 60
return stats
Fine-tuning XTTS v2—stack and configuration
We use the official Coqui TTS repository with modifications for commercial tasks. Below is the config for fine-tuning only the decoder (faster, less noise).
from trainer import Trainer, TrainerArgs
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts
config = XttsConfig()
config.load_json("base_xtts_config.json")
# Fine-tuning parameters
config.audio.output_sample_rate = 24000
config.batch_size = 4
config.eval_batch_size = 2
config.num_loader_workers = 4
# Fine-tuning only decoder (faster, less data)
config.trainer_args = {
"epochs": 100,
"save_step": 1000,
"print_step": 50,
"eval_split_size": 0.1
}
Variations: you can fine-tune the entire encoder+decoder if dataset >2 hours, but this increases training time 2–3x and requires caution with overfitting.
How to evaluate synthesized voice quality?
The primary metric is MOS (Mean Opinion Score) per ITU-T P.800. We use an internal panel of 10–15 listeners, each evaluating 50–80 samples. Results:
| Configuration |
MOS (95% CI) |
| XTTS zero-shot |
3.7–3.9 |
| Fine-tuned 30 min |
4.1–4.3 |
| Fine-tuned 60+ min |
4.3–4.5 |
Objective metrics:
-
UTMOS: automatic naturalness score (MOS-predictor model)
-
SECS (Speaker Embedding Cosine Similarity): similarity to donor voice >0.95
-
WER on reverse recognition: no more than 5% at medium pace
Infrastructure and training cost
GPU selection depends on budget and required speed. We recommend configurations with minimal FLOPS:
| Configuration |
Time (30 min of data) |
Note |
| 1x A100 80GB |
~3–4 hours |
Optimal for batch size 8 |
| 1x A10G |
~6–8 hours |
Price/performance balance |
| 1x RTX 4090 |
~8–12 hours |
Local training |
Training cost depends on the chosen configuration and data volume. Savings compared to buying a ready-made TTS solution can reach 30–50%. We help select a configuration within your budget.
What's included in our TTS fine-tuning project?
- Source material audit—evaluate recording quality, noise, diction.
- Dataset preparation—cleaning, volume normalization, segmentation (force alignment).
- Model training—choose architecture (XTTS, IhreTTS, YourTTS), tune hyperparameters.
- Quality evaluation—MOS, UTMOS, SECS, WER.
- Model export—ONNX / TorchScript for inference.
- Integration—API wrapper, testing in your product.
- Documentation and team training—how to update the voice, extend fine-tuning.
We guarantee: final MOS at least 4.0 with a dataset from 30 minutes. If not met, we redo at our cost.
Estimated timelines
| Stage |
Duration |
| Dataset collection and cleaning |
1–2 weeks |
| Training and evaluation |
3–5 days |
| Integration and testing |
3–5 days |
| Total |
3–4 weeks |
How to avoid common fine-tuning pitfalls
Recordings with background noise are the main enemy of quality. We apply HPSS filter and VAD segmentation. Phoneme imbalance (e.g., missing unvoiced or sibilant sounds) is compensated by a specialized script to create a balanced dataset. On small data (<30 minutes), L2 regularization and early stopping help. All these measures ensure stable results without overfitting.
If you have questions about dataset, architecture, or budget—contact us for a consultation. Request a cost estimate for your project—we will find the optimal solution for your needs.
Speech Recognition and Synthesis: ASR, TTS, Voice Cloning
We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.
How does speech recognition ASR handle noisy call center recordings?
The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.
Typical problems we encounter
WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.
Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.
Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.
How do we solve it in practice?
The basic stack for a production pipeline:
-
ASR:
openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
-
Diarization:
pyannote.audio 3.x + integration via whisperx for word-level alignment
-
TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
-
Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2
A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=True → pyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.
Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.
How to fine-tune Whisper on domain data?
When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.
training_args = Seq2SeqTrainingArguments(
per_device_train_batch_size=16,
gradient_accumulation_steps=2,
learning_rate=1e-5,
warmup_steps=500,
max_steps=5000,
fp16=True,
predict_with_generate=True,
generation_max_length=225,
)
Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.
| Model |
WER (clean) |
WER (noisy) |
RTF (A10G) |
Languages |
| Whisper large-v3 |
5.2% |
27% |
0.08 |
99 |
| Wav2Vec2-XLSR-53 |
6.8% |
32% |
0.12 |
143 |
| Google STT (cloud) |
7.0% |
28% |
– |
125 |
| DeepSpeech 0.9.3 |
11.5% |
41% |
0.06 |
8 |
Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.
Speech synthesis: How to choose a model for your task?
| Model |
Latency (TTFB) |
Naturalness MOS |
Cloning |
Languages |
| XTTS v2 |
1.2–2.0 s |
4.1–4.3 |
Yes, 3 s reference |
17 |
| StyleTTS2 |
0.3–0.6 s |
4.0–4.2 |
Yes, requires adaptation |
en, + fine-tune |
| Kokoro-82M |
0.08–0.15 s |
3.7–3.9 |
No |
en, ja |
| Silero TTS |
0.05–0.1 s |
3.4–3.6 |
No |
ru, en, de, etc. |
| Edge-TTS |
~0.4 s (cloud) |
4.0 |
No |
100+ |
For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.
Our process and deliverables
We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.
Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.
What you get after engagement:
- Trained model with model card and evaluation report
- Docker image with optimized inference pipeline
- API documentation and integration examples
- Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
- 30-day post-deployment support and hotfixing
Timelines depend on complexity:
- Basic integration of a ready model — 1–2 weeks
- Fine-tuning with data preparation and validation — 4–8 weeks
- Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months
Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.
Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.
Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.