Custom TTS Voice: Training with VITS, XTTS, YourTTS

Speech Synthesis: VITS and XTTS for Custom Voice

AI Development Areas

Frequently Asked Questions

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1441
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1302
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    998
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    714
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1006

Speech Synthesis: VITS and XTTS for Custom Voice

A custom TTS model gives you full control over voice, language, and style — without dependence on external APIs and recurring costs. It's relevant for creating a unique brand voice, synthesis in rare languages/dialects, and edge deployment without internet. We have trained over 15 models for clients in retail, media, and voice assistants — from short advertising jingles to full-fledged dialog systems. We guarantee achieving a MOS of at least 4.0.

Why Train a Custom TTS Model?

Ready-made cloud TTS (Google, Yandex, Amazon) impose limitations: a fixed set of voices, cost per request, internet dependency, and latency. A custom model solves these problems: you get an exclusive voice that works offline, with control over emotional tone and pace. For example, one of our clients (a delivery aggregator) saved $3,000 per month by switching from a paid API to their model trained on 8 hours of a voice actor's speech. — Client from retail Savings on API calls can reach $50,000 per year for large projects.

How to Choose the TTS Architecture?

Model Type Training Data Quality (MOS) Inference Speed
VITS End-to-end (text→audio) 2–5 h 4.2/5 Realtime ×30 on GPU
XTTS v2 (Coqui) Zero-shot + fine-tune 3–6 min (few-shot) 4.4/5 Realtime ×10 on GPU
YourTTS Multilingual VITS 1–3 h 4.0/5 Realtime ×20
MATCHA-TTS Flow-matching 2–4 h 4.3/5 Realtime ×50
StyleTTS2 Style-based 1–2 h 4.5/5 Realtime ×15

For most tasks: XTTS v2 for quick startup with minimal data, VITS for full training with a clean dataset. When fine-tuned on 6 minutes of audio, XTTS v2 achieves quality comparable to full VITS on 10 hours – confirmed by our MOS measurements. With years of experience, we select the architecture best suited for your task.

Dataset Preparation

Minimum requirements for quality results:

Format: 22050 Hz, 16-bit, mono WAV Recording length: 2–15 seconds each Minimum: 1000 recordings (≈2 hours) for intelligible TTS Recommended: 3000–5000 recordings (≈8–12 hours) for high quality Text script: UTF-8, one utterance per line 

Dataset structure:

dataset/ ├── wavs/ │ ├── speaker_001.wav │ ├── speaker_002.wav │ └── ... ├── metadata.csv # filename|transcription └── metadata_val.csv # 10% for validation 

Preprocessing and normalization:

import librosa import soundfile as sf import numpy as np from pathlib import Path def preprocess_audio_for_tts( input_dir: str, output_dir: str, target_sr: int = 22050 ) -> dict: stats = {"processed": 0, "skipped": 0, "errors": []} Path(output_dir).mkdir(parents=True, exist_ok=True) for wav_path in Path(input_dir).glob("*.wav"): audio, sr = librosa.load(str(wav_path), sr=target_sr, mono=True) # Trim silence audio_trimmed, _ = librosa.effects.trim(audio, top_db=20) # Check length duration = len(audio_trimmed) / target_sr if duration < 1.5 or duration > 15.0: stats["skipped"] += 1 continue # Normalize amplitude audio_normalized = audio_trimmed / (np.max(np.abs(audio_trimmed)) + 1e-8) audio_normalized *= 0.9 # peak -0.9 dB output_path = Path(output_dir) / wav_path.name sf.write(str(output_path), audio_normalized, target_sr, subtype="PCM_16") stats["processed"] += 1 return stats 

VITS Training

config.json configuration for VITS (Coqui TTS):

{ "model": "vits", "run_name": "my_tts_model", "epochs": 1000, "batch_size": 32, "eval_batch_size": 16, "num_loader_workers": 4, "audio": { "sample_rate": 22050, "win_length": 1024, "hop_length": 256, "num_mels": 80, "mel_fmin": 0, "mel_fmax": null }, "datasets": [{ "name": "my_dataset", "path": "dataset/", "meta_file_train": "metadata.csv", "meta_file_val": "metadata_val.csv" }] } 

Launch training:

from TTS.bin.train_tts import main as train_tts from TTS.config.shared_configs import BaseDatasetConfig from TTS.tts.configs.vits_config import VitsConfig from TTS.tts.datasets import load_tts_samples from TTS.tts.models.vits import Vits, VitsAudioConfig from TTS.trainer import Trainer, TrainerArgs audio_config = VitsAudioConfig( sample_rate=22050, win_length=1024, hop_length=256, num_mels=80, mel_fmin=0, mel_fmax=None ) config = VitsConfig( audio=audio_config, run_name="brand_voice_v1", batch_size=32, eval_batch_size=16, epochs=1000, text_cleaner="phoneme_cleaners", use_phonemes=True, phoneme_language="ru-ru", phoneme_cache_path="phoneme_cache/", output_path="checkpoints/", datasets=[BaseDatasetConfig( formatter="ljspeech", meta_file_train="metadata.csv", path="dataset/" )] ) train_samples, eval_samples = load_tts_samples( config.datasets, eval_split=True, eval_split_size=0.1 ) model = Vits(config, ap=None, tokenizer=None, speaker_manager=None) trainer = Trainer( TrainerArgs(), config, output_path="checkpoints/", model=model, train_samples=train_samples, eval_samples=eval_samples ) trainer.fit() 

XTTS v2 Fine-Tuning (Few-Shot)

XTTS v2 supports fine-tuning with 3–6 minutes of audio:

from TTS.demos.xtts_ft_demo.xtts_demo import train_gpt # Dataset: at least 100 recordings, each 2–6 seconds long train_gpt( language="ru", num_epochs=6, batch_size=4, grad_acumm=1, train_csv="dataset/metadata_train.csv", eval_csv="dataset/metadata_eval.csv", output_path="xtts_ft_checkpoints/" ) 

Inference with custom voice after fine-tuning:

from TTS.api import TTS tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2") tts.tts_to_file( text="Welcome to our company.", speaker_wav="reference_voice.wav", # 3–10 sec reference audio language="ru", file_path="output.wav", model_path="xtts_ft_checkpoints/best_model.pth" ) 

Our Approach to TTS Model Training

Our process includes five stages:

  1. Analytics: determine the target audience for the voice, requirements for language, emotions, speed. Select architecture (VITS, XTTS, YourTTS) for the task.
  2. Dataset collection and preparation: record voice actor in studio or clean existing recordings. Remove noise, silence, normalize. Transcribe texts.
  3. Model training: run on GPU cluster, monitor metrics (train/val loss, KL loss, grad_norm). Use early stopping and checkpoints.
  4. Quality evaluation: listen to synthesis every 100 epochs, compare to reference. Achieve MOS of at least 4.0.
  5. Deployment and integration: convert to ONNX for edge or deploy as gRPC/REST API. Provide documentation and support.
Training Metric Monitoring Key metrics in tensorboard: - loss/train_loss: should decrease monotonically - loss/val_loss: parallel to train, no divergence - loss/kl_loss: KL divergence of latent space - loss/disc_loss: discriminator (GAN component) - grad_norm: should be < 10, otherwise gradient explosion

Training Infrastructure

GPU Training Time (1000 epochs, VITS) VRAM
RTX 3090 (24 GB) ~12 hours 18 GB
A100 (40 GB) ~5 hours 22 GB
2× A10G ~3 hours 2×24 GB
CPU (no GPU) Not recommended

Cloud options: RunPod ($1.5/h for A100), Lambda Cloud ($1.1/h), Vast.ai (~$0.5–0.8/h for A100).

Post-Training: Model Deployment

# ONNX export for edge deployment from TTS.utils.synthesizer import Synthesizer synthesizer = Synthesizer( tts_checkpoint="checkpoints/best_model.pth", tts_config_path="checkpoints/config.json" ) # Inference wav = synthesizer.tts("Test phrase for synthesis") synthesizer.save_wav(wav, "test_output.wav") 

What's Included in the Work

  • Trained model (VITS, XTTS, or YourTTS) with achieved quality of at least MOS 4.0.
  • Clean dataset with transcriptions and preprocessing scripts.
  • Configuration files and code to reproduce training.
  • Inference scripts for local and server use.
  • API wrapper (FastAPI/gRPC) for integration into your service.
  • Documentation for setup and operation.
  • Support for 2 weeks after delivery.

Timeline: dataset preparation (recording + transcription) — 2–4 weeks. VITS model training — 1–2 weeks (GPU). Integration into production service with API — 1 week. Full cycle from scratch to brand voice — 4–6 weeks. Get a consultation from our AI engineer — we will select the optimal architecture and calculate precise deadlines. Order TTS model training for your project.