Custom TTS Voice: Training with VITS, XTTS, YourTTS

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Custom TTS Voice: Training with VITS, XTTS, YourTTS
Complex
~5 days
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1361
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1189
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Speech Synthesis: VITS and XTTS for Custom Voice

A custom TTS model gives you full control over voice, language, and style — without dependence on external APIs and recurring costs. It's relevant for creating a unique brand voice, synthesis in rare languages/dialects, and edge deployment without internet. We have trained over 15 models for clients in retail, media, and voice assistants — from short advertising jingles to full-fledged dialog systems. We guarantee achieving a MOS of at least 4.0.

Why Train a Custom TTS Model?

Ready-made cloud TTS (Google, Yandex, Amazon) impose limitations: a fixed set of voices, cost per request, internet dependency, and latency. A custom model solves these problems: you get an exclusive voice that works offline, with control over emotional tone and pace. For example, one of our clients (a delivery aggregator) saved $3,000 per month by switching from a paid API to their model trained on 8 hours of a voice actor's speech. — Client from retail Savings on API calls can reach $50,000 per year for large projects.

How to Choose the TTS Architecture?

Model Type Training Data Quality (MOS) Inference Speed
VITS End-to-end (text→audio) 2–5 h 4.2/5 Realtime ×30 on GPU
XTTS v2 (Coqui) Zero-shot + fine-tune 3–6 min (few-shot) 4.4/5 Realtime ×10 on GPU
YourTTS Multilingual VITS 1–3 h 4.0/5 Realtime ×20
MATCHA-TTS Flow-matching 2–4 h 4.3/5 Realtime ×50
StyleTTS2 Style-based 1–2 h 4.5/5 Realtime ×15

For most tasks: XTTS v2 for quick startup with minimal data, VITS for full training with a clean dataset. When fine-tuned on 6 minutes of audio, XTTS v2 achieves quality comparable to full VITS on 10 hours – confirmed by our MOS measurements. With years of experience, we select the architecture best suited for your task.

Dataset Preparation

Minimum requirements for quality results:

Format: 22050 Hz, 16-bit, mono WAV
Recording length: 2–15 seconds each
Minimum: 1000 recordings (≈2 hours) for intelligible TTS
Recommended: 3000–5000 recordings (≈8–12 hours) for high quality
Text script: UTF-8, one utterance per line

Dataset structure:

dataset/
├── wavs/
│   ├── speaker_001.wav
│   ├── speaker_002.wav
│   └── ...
├── metadata.csv          # filename|transcription
└── metadata_val.csv      # 10% for validation

Preprocessing and normalization:

import librosa
import soundfile as sf
import numpy as np
from pathlib import Path

def preprocess_audio_for_tts(
    input_dir: str,
    output_dir: str,
    target_sr: int = 22050
) -> dict:
    stats = {"processed": 0, "skipped": 0, "errors": []}
    Path(output_dir).mkdir(parents=True, exist_ok=True)

    for wav_path in Path(input_dir).glob("*.wav"):
        audio, sr = librosa.load(str(wav_path), sr=target_sr, mono=True)

        # Trim silence
        audio_trimmed, _ = librosa.effects.trim(audio, top_db=20)

        # Check length
        duration = len(audio_trimmed) / target_sr
        if duration < 1.5 or duration > 15.0:
            stats["skipped"] += 1
            continue

        # Normalize amplitude
        audio_normalized = audio_trimmed / (np.max(np.abs(audio_trimmed)) + 1e-8)
        audio_normalized *= 0.9  # peak -0.9 dB

        output_path = Path(output_dir) / wav_path.name
        sf.write(str(output_path), audio_normalized, target_sr, subtype="PCM_16")
        stats["processed"] += 1

    return stats

VITS Training

config.json configuration for VITS (Coqui TTS):

{
    "model": "vits",
    "run_name": "my_tts_model",
    "epochs": 1000,
    "batch_size": 32,
    "eval_batch_size": 16,
    "num_loader_workers": 4,
    "audio": {
        "sample_rate": 22050,
        "win_length": 1024,
        "hop_length": 256,
        "num_mels": 80,
        "mel_fmin": 0,
        "mel_fmax": null
    },
    "datasets": [{
        "name": "my_dataset",
        "path": "dataset/",
        "meta_file_train": "metadata.csv",
        "meta_file_val": "metadata_val.csv"
    }]
}

Launch training:

from TTS.bin.train_tts import main as train_tts
from TTS.config.shared_configs import BaseDatasetConfig
from TTS.tts.configs.vits_config import VitsConfig
from TTS.tts.datasets import load_tts_samples
from TTS.tts.models.vits import Vits, VitsAudioConfig
from TTS.trainer import Trainer, TrainerArgs

audio_config = VitsAudioConfig(
    sample_rate=22050,
    win_length=1024,
    hop_length=256,
    num_mels=80,
    mel_fmin=0,
    mel_fmax=None
)

config = VitsConfig(
    audio=audio_config,
    run_name="brand_voice_v1",
    batch_size=32,
    eval_batch_size=16,
    epochs=1000,
    text_cleaner="phoneme_cleaners",
    use_phonemes=True,
    phoneme_language="ru-ru",
    phoneme_cache_path="phoneme_cache/",
    output_path="checkpoints/",
    datasets=[BaseDatasetConfig(
        formatter="ljspeech",
        meta_file_train="metadata.csv",
        path="dataset/"
    )]
)

train_samples, eval_samples = load_tts_samples(
    config.datasets,
    eval_split=True,
    eval_split_size=0.1
)

model = Vits(config, ap=None, tokenizer=None, speaker_manager=None)

trainer = Trainer(
    TrainerArgs(),
    config,
    output_path="checkpoints/",
    model=model,
    train_samples=train_samples,
    eval_samples=eval_samples
)
trainer.fit()

XTTS v2 Fine-Tuning (Few-Shot)

XTTS v2 supports fine-tuning with 3–6 minutes of audio:

from TTS.demos.xtts_ft_demo.xtts_demo import train_gpt

# Dataset: at least 100 recordings, each 2–6 seconds long
train_gpt(
    language="ru",
    num_epochs=6,
    batch_size=4,
    grad_acumm=1,
    train_csv="dataset/metadata_train.csv",
    eval_csv="dataset/metadata_eval.csv",
    output_path="xtts_ft_checkpoints/"
)

Inference with custom voice after fine-tuning:

from TTS.api import TTS

tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2")
tts.tts_to_file(
    text="Welcome to our company.",
    speaker_wav="reference_voice.wav",  # 3–10 sec reference audio
    language="ru",
    file_path="output.wav",
    model_path="xtts_ft_checkpoints/best_model.pth"
)

Our Approach to TTS Model Training

Our process includes five stages:

  1. Analytics: determine the target audience for the voice, requirements for language, emotions, speed. Select architecture (VITS, XTTS, YourTTS) for the task.
  2. Dataset collection and preparation: record voice actor in studio or clean existing recordings. Remove noise, silence, normalize. Transcribe texts.
  3. Model training: run on GPU cluster, monitor metrics (train/val loss, KL loss, grad_norm). Use early stopping and checkpoints.
  4. Quality evaluation: listen to synthesis every 100 epochs, compare to reference. Achieve MOS of at least 4.0.
  5. Deployment and integration: convert to ONNX for edge or deploy as gRPC/REST API. Provide documentation and support.
Training Metric Monitoring Key metrics in tensorboard: - loss/train_loss: should decrease monotonically - loss/val_loss: parallel to train, no divergence - loss/kl_loss: KL divergence of latent space - loss/disc_loss: discriminator (GAN component) - grad_norm: should be < 10, otherwise gradient explosion

Training Infrastructure

GPU Training Time (1000 epochs, VITS) VRAM
RTX 3090 (24 GB) ~12 hours 18 GB
A100 (40 GB) ~5 hours 22 GB
2× A10G ~3 hours 2×24 GB
CPU (no GPU) Not recommended

Cloud options: RunPod ($1.5/h for A100), Lambda Cloud ($1.1/h), Vast.ai (~$0.5–0.8/h for A100).

Post-Training: Model Deployment

# ONNX export for edge deployment
from TTS.utils.synthesizer import Synthesizer

synthesizer = Synthesizer(
    tts_checkpoint="checkpoints/best_model.pth",
    tts_config_path="checkpoints/config.json"
)

# Inference
wav = synthesizer.tts("Test phrase for synthesis")
synthesizer.save_wav(wav, "test_output.wav")

What's Included in the Work

  • Trained model (VITS, XTTS, or YourTTS) with achieved quality of at least MOS 4.0.
  • Clean dataset with transcriptions and preprocessing scripts.
  • Configuration files and code to reproduce training.
  • Inference scripts for local and server use.
  • API wrapper (FastAPI/gRPC) for integration into your service.
  • Documentation for setup and operation.
  • Support for 2 weeks after delivery.

Timeline: dataset preparation (recording + transcription) — 2–4 weeks. VITS model training — 1–2 weeks (GPU). Integration into production service with API — 1 week. Full cycle from scratch to brand voice — 4–6 weeks. Get a consultation from our AI engineer — we will select the optimal architecture and calculate precise deadlines. Order TTS model training for your project.

Speech Recognition and Synthesis: ASR, TTS, Voice Cloning

We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.

How does speech recognition ASR handle noisy call center recordings?

The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.

Typical problems we encounter

WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.

Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.

Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.

How do we solve it in practice?

The basic stack for a production pipeline:

  • ASR: openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
  • Diarization: pyannote.audio 3.x + integration via whisperx for word-level alignment
  • TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
  • Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2

A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=Truepyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.

Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.

How to fine-tune Whisper on domain data?

When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.

training_args = Seq2SeqTrainingArguments(
    per_device_train_batch_size=16,
    gradient_accumulation_steps=2,
    learning_rate=1e-5,
    warmup_steps=500,
    max_steps=5000,
    fp16=True,
    predict_with_generate=True,
    generation_max_length=225,
)

Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.

Model WER (clean) WER (noisy) RTF (A10G) Languages
Whisper large-v3 5.2% 27% 0.08 99
Wav2Vec2-XLSR-53 6.8% 32% 0.12 143
Google STT (cloud) 7.0% 28% 125
DeepSpeech 0.9.3 11.5% 41% 0.06 8

Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.

Speech synthesis: How to choose a model for your task?

Model Latency (TTFB) Naturalness MOS Cloning Languages
XTTS v2 1.2–2.0 s 4.1–4.3 Yes, 3 s reference 17
StyleTTS2 0.3–0.6 s 4.0–4.2 Yes, requires adaptation en, + fine-tune
Kokoro-82M 0.08–0.15 s 3.7–3.9 No en, ja
Silero TTS 0.05–0.1 s 3.4–3.6 No ru, en, de, etc.
Edge-TTS ~0.4 s (cloud) 4.0 No 100+

For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.

Our process and deliverables

We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.

Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.

What you get after engagement:

  • Trained model with model card and evaluation report
  • Docker image with optimized inference pipeline
  • API documentation and integration examples
  • Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
  • 30-day post-deployment support and hotfixing

Timelines depend on complexity:

  • Basic integration of a ready model — 1–2 weeks
  • Fine-tuning with data preparation and validation — 4–8 weeks
  • Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months

Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.

Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.

Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.