Speech Synthesis: VITS and XTTS for Custom Voice
A custom TTS model gives you full control over voice, language, and style — without dependence on external APIs and recurring costs. It's relevant for creating a unique brand voice, synthesis in rare languages/dialects, and edge deployment without internet. We have trained over 15 models for clients in retail, media, and voice assistants — from short advertising jingles to full-fledged dialog systems. We guarantee achieving a MOS of at least 4.0.
Why Train a Custom TTS Model?
Ready-made cloud TTS (Google, Yandex, Amazon) impose limitations: a fixed set of voices, cost per request, internet dependency, and latency. A custom model solves these problems: you get an exclusive voice that works offline, with control over emotional tone and pace. For example, one of our clients (a delivery aggregator) saved $3,000 per month by switching from a paid API to their model trained on 8 hours of a voice actor's speech. — Client from retail Savings on API calls can reach $50,000 per year for large projects.
How to Choose the TTS Architecture?
| Model | Type | Training Data | Quality (MOS) | Inference Speed |
|---|---|---|---|---|
| VITS | End-to-end (text→audio) | 2–5 h | 4.2/5 | Realtime ×30 on GPU |
| XTTS v2 (Coqui) | Zero-shot + fine-tune | 3–6 min (few-shot) | 4.4/5 | Realtime ×10 on GPU |
| YourTTS | Multilingual VITS | 1–3 h | 4.0/5 | Realtime ×20 |
| MATCHA-TTS | Flow-matching | 2–4 h | 4.3/5 | Realtime ×50 |
| StyleTTS2 | Style-based | 1–2 h | 4.5/5 | Realtime ×15 |
For most tasks: XTTS v2 for quick startup with minimal data, VITS for full training with a clean dataset. When fine-tuned on 6 minutes of audio, XTTS v2 achieves quality comparable to full VITS on 10 hours – confirmed by our MOS measurements. With years of experience, we select the architecture best suited for your task.
Dataset Preparation
Minimum requirements for quality results:
Format: 22050 Hz, 16-bit, mono WAV
Recording length: 2–15 seconds each
Minimum: 1000 recordings (≈2 hours) for intelligible TTS
Recommended: 3000–5000 recordings (≈8–12 hours) for high quality
Text script: UTF-8, one utterance per line
Dataset structure:
dataset/
├── wavs/
│ ├── speaker_001.wav
│ ├── speaker_002.wav
│ └── ...
├── metadata.csv # filename|transcription
└── metadata_val.csv # 10% for validation
Preprocessing and normalization:
import librosa
import soundfile as sf
import numpy as np
from pathlib import Path
def preprocess_audio_for_tts(
input_dir: str,
output_dir: str,
target_sr: int = 22050
) -> dict:
stats = {"processed": 0, "skipped": 0, "errors": []}
Path(output_dir).mkdir(parents=True, exist_ok=True)
for wav_path in Path(input_dir).glob("*.wav"):
audio, sr = librosa.load(str(wav_path), sr=target_sr, mono=True)
# Trim silence
audio_trimmed, _ = librosa.effects.trim(audio, top_db=20)
# Check length
duration = len(audio_trimmed) / target_sr
if duration < 1.5 or duration > 15.0:
stats["skipped"] += 1
continue
# Normalize amplitude
audio_normalized = audio_trimmed / (np.max(np.abs(audio_trimmed)) + 1e-8)
audio_normalized *= 0.9 # peak -0.9 dB
output_path = Path(output_dir) / wav_path.name
sf.write(str(output_path), audio_normalized, target_sr, subtype="PCM_16")
stats["processed"] += 1
return stats
VITS Training
config.json configuration for VITS (Coqui TTS):
{
"model": "vits",
"run_name": "my_tts_model",
"epochs": 1000,
"batch_size": 32,
"eval_batch_size": 16,
"num_loader_workers": 4,
"audio": {
"sample_rate": 22050,
"win_length": 1024,
"hop_length": 256,
"num_mels": 80,
"mel_fmin": 0,
"mel_fmax": null
},
"datasets": [{
"name": "my_dataset",
"path": "dataset/",
"meta_file_train": "metadata.csv",
"meta_file_val": "metadata_val.csv"
}]
}
Launch training:
from TTS.bin.train_tts import main as train_tts
from TTS.config.shared_configs import BaseDatasetConfig
from TTS.tts.configs.vits_config import VitsConfig
from TTS.tts.datasets import load_tts_samples
from TTS.tts.models.vits import Vits, VitsAudioConfig
from TTS.trainer import Trainer, TrainerArgs
audio_config = VitsAudioConfig(
sample_rate=22050,
win_length=1024,
hop_length=256,
num_mels=80,
mel_fmin=0,
mel_fmax=None
)
config = VitsConfig(
audio=audio_config,
run_name="brand_voice_v1",
batch_size=32,
eval_batch_size=16,
epochs=1000,
text_cleaner="phoneme_cleaners",
use_phonemes=True,
phoneme_language="ru-ru",
phoneme_cache_path="phoneme_cache/",
output_path="checkpoints/",
datasets=[BaseDatasetConfig(
formatter="ljspeech",
meta_file_train="metadata.csv",
path="dataset/"
)]
)
train_samples, eval_samples = load_tts_samples(
config.datasets,
eval_split=True,
eval_split_size=0.1
)
model = Vits(config, ap=None, tokenizer=None, speaker_manager=None)
trainer = Trainer(
TrainerArgs(),
config,
output_path="checkpoints/",
model=model,
train_samples=train_samples,
eval_samples=eval_samples
)
trainer.fit()
XTTS v2 Fine-Tuning (Few-Shot)
XTTS v2 supports fine-tuning with 3–6 minutes of audio:
from TTS.demos.xtts_ft_demo.xtts_demo import train_gpt
# Dataset: at least 100 recordings, each 2–6 seconds long
train_gpt(
language="ru",
num_epochs=6,
batch_size=4,
grad_acumm=1,
train_csv="dataset/metadata_train.csv",
eval_csv="dataset/metadata_eval.csv",
output_path="xtts_ft_checkpoints/"
)
Inference with custom voice after fine-tuning:
from TTS.api import TTS
tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2")
tts.tts_to_file(
text="Welcome to our company.",
speaker_wav="reference_voice.wav", # 3–10 sec reference audio
language="ru",
file_path="output.wav",
model_path="xtts_ft_checkpoints/best_model.pth"
)
Our Approach to TTS Model Training
Our process includes five stages:
- Analytics: determine the target audience for the voice, requirements for language, emotions, speed. Select architecture (VITS, XTTS, YourTTS) for the task.
- Dataset collection and preparation: record voice actor in studio or clean existing recordings. Remove noise, silence, normalize. Transcribe texts.
- Model training: run on GPU cluster, monitor metrics (train/val loss, KL loss, grad_norm). Use early stopping and checkpoints.
- Quality evaluation: listen to synthesis every 100 epochs, compare to reference. Achieve MOS of at least 4.0.
- Deployment and integration: convert to ONNX for edge or deploy as gRPC/REST API. Provide documentation and support.
Training Metric Monitoring
Key metrics in tensorboard: - loss/train_loss: should decrease monotonically - loss/val_loss: parallel to train, no divergence - loss/kl_loss: KL divergence of latent space - loss/disc_loss: discriminator (GAN component) - grad_norm: should be < 10, otherwise gradient explosionTraining Infrastructure
| GPU | Training Time (1000 epochs, VITS) | VRAM |
|---|---|---|
| RTX 3090 (24 GB) | ~12 hours | 18 GB |
| A100 (40 GB) | ~5 hours | 22 GB |
| 2× A10G | ~3 hours | 2×24 GB |
| CPU (no GPU) | Not recommended | — |
Cloud options: RunPod ($1.5/h for A100), Lambda Cloud ($1.1/h), Vast.ai (~$0.5–0.8/h for A100).
Post-Training: Model Deployment
# ONNX export for edge deployment
from TTS.utils.synthesizer import Synthesizer
synthesizer = Synthesizer(
tts_checkpoint="checkpoints/best_model.pth",
tts_config_path="checkpoints/config.json"
)
# Inference
wav = synthesizer.tts("Test phrase for synthesis")
synthesizer.save_wav(wav, "test_output.wav")
What's Included in the Work
- Trained model (VITS, XTTS, or YourTTS) with achieved quality of at least MOS 4.0.
- Clean dataset with transcriptions and preprocessing scripts.
- Configuration files and code to reproduce training.
- Inference scripts for local and server use.
- API wrapper (FastAPI/gRPC) for integration into your service.
- Documentation for setup and operation.
- Support for 2 weeks after delivery.
Timeline: dataset preparation (recording + transcription) — 2–4 weeks. VITS model training — 1–2 weeks (GPU). Integration into production service with API — 1 week. Full cycle from scratch to brand voice — 4–6 weeks. Get a consultation from our AI engineer — we will select the optimal architecture and calculate precise deadlines. Order TTS model training for your project.







