Custom voice: TTS model fine-tuning with MOS 4.3+ guarantee

Every third response from your voice assistant sounds unnatural—trembling timbre, missing phonemes. We solve this by fine-tuning a TTS model on the client's voice. After fine-tuning on 30–60 minutes of recordings, the model steadily reads any text: MOS rises to 4.3+ from 3.8 in zero-shot, and revers

AI Development Areas

Frequently Asked Questions

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1441
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    998
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    713
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1003

Every third response from your voice assistant sounds unnatural—trembling timbre, missing phonemes. We solve this by fine-tuning a TTS model on the client's voice. After fine-tuning on 30–60 minutes of recordings, the model steadily reads any text: MOS rises to 4.3+ from 3.8 in zero-shot, and reverse recognition WER drops by 5–10%. Result: the assistant stops 'stuttering' even on complex queries.

Why fine-tuning over zero-shot?

Zero-shot cloning (e.g., XTTSv2 in speaker encoder mode) gives acceptable results but suffers from timbre trembling, artifacts on rare phonemes, and instability on long texts. Fine-tuning on 30–60 minutes of the target voice locks in the speaker's acoustic space, reduces reverse recognition WER by 5–10%, and increases UTMOS by 0.3–0.5. Main advantages: predictable quality on any input, ability to augment data (noise, reverberation), and control over intonation via conditioning.

What goes into dataset preparation for TTS fine-tuning?

Minimum volume: 30 minutes of clean recordings. Optimal: 1–2 hours. Audio requirements: sampling rate 22050 or 24000 Hz, signal level –18…–12 dBFS, signal-to-noise ratio >30 dB, clip lengths 3–15 seconds.

Preparation steps:

  • Record in a studio or quiet room (check for background noise).
  • Clean noise: use HPSS filter or spectral subtraction.
  • Sentence-level segmentation: force alignment with Montreal Forced Aligner.
  • Validate duration and quality: run through a validation script.
Example dataset validation script
import pandas as pd from pathlib import Path import soundfile as sf import numpy as np def validate_dataset(dataset_dir: str) -> dict: """Check dataset before training""" metadata = pd.read_csv(f"{dataset_dir}/metadata.csv", sep="|", names=["file", "text"]) stats = { "total_files": len(metadata), "total_duration": 0, "errors": [] } for _, row in metadata.iterrows(): wav_path = f"{dataset_dir}/wavs/{row['file']}.wav" if not Path(wav_path).exists(): stats["errors"].append(f"Missing: {wav_path}") continue audio, sr = sf.read(wav_path) duration = len(audio) / sr stats["total_duration"] += duration if sr != 22050: stats["errors"].append(f"Wrong SR {sr}: {wav_path}") if duration < 1.0 or duration > 15.0: stats["errors"].append(f"Bad duration {duration:.1f}s: {wav_path}") stats["total_duration_min"] = stats["total_duration"] / 60 return stats 

Fine-tuning XTTS v2—stack and configuration

We use the official Coqui TTS repository with modifications for commercial tasks. Below is the config for fine-tuning only the decoder (faster, less noise).

from trainer import Trainer, TrainerArgs from TTS.tts.configs.xtts_config import XttsConfig from TTS.tts.models.xtts import Xtts config = XttsConfig() config.load_json("base_xtts_config.json") # Fine-tuning parameters config.audio.output_sample_rate = 24000 config.batch_size = 4 config.eval_batch_size = 2 config.num_loader_workers = 4 # Fine-tuning only decoder (faster, less data) config.trainer_args = { "epochs": 100, "save_step": 1000, "print_step": 50, "eval_split_size": 0.1 } 

Variations: you can fine-tune the entire encoder+decoder if dataset >2 hours, but this increases training time 2–3x and requires caution with overfitting.

How to evaluate synthesized voice quality?

The primary metric is MOS (Mean Opinion Score) per ITU-T P.800. We use an internal panel of 10–15 listeners, each evaluating 50–80 samples. Results:

Configuration MOS (95% CI)
XTTS zero-shot 3.7–3.9
Fine-tuned 30 min 4.1–4.3
Fine-tuned 60+ min 4.3–4.5

Objective metrics:

  • UTMOS: automatic naturalness score (MOS-predictor model)
  • SECS (Speaker Embedding Cosine Similarity): similarity to donor voice >0.95
  • WER on reverse recognition: no more than 5% at medium pace

Infrastructure and training cost

GPU selection depends on budget and required speed. We recommend configurations with minimal FLOPS:

Configuration Time (30 min of data) Note
1x A100 80GB ~3–4 hours Optimal for batch size 8
1x A10G ~6–8 hours Price/performance balance
1x RTX 4090 ~8–12 hours Local training

Training cost depends on the chosen configuration and data volume. Savings compared to buying a ready-made TTS solution can reach 30–50%. We help select a configuration within your budget.

What's included in our TTS fine-tuning project?

  1. Source material audit—evaluate recording quality, noise, diction.
  2. Dataset preparation—cleaning, volume normalization, segmentation (force alignment).
  3. Model training—choose architecture (XTTS, IhreTTS, YourTTS), tune hyperparameters.
  4. Quality evaluation—MOS, UTMOS, SECS, WER.
  5. Model export—ONNX / TorchScript for inference.
  6. Integration—API wrapper, testing in your product.
  7. Documentation and team training—how to update the voice, extend fine-tuning.

We guarantee: final MOS at least 4.0 with a dataset from 30 minutes. If not met, we redo at our cost.

Estimated timelines

Stage Duration
Dataset collection and cleaning 1–2 weeks
Training and evaluation 3–5 days
Integration and testing 3–5 days
Total 3–4 weeks

How to avoid common fine-tuning pitfalls

Recordings with background noise are the main enemy of quality. We apply HPSS filter and VAD segmentation. Phoneme imbalance (e.g., missing unvoiced or sibilant sounds) is compensated by a specialized script to create a balanced dataset. On small data (<30 minutes), L2 regularization and early stopping help. All these measures ensure stable results without overfitting.

If you have questions about dataset, architecture, or budget—contact us for a consultation. Request a cost estimate for your project—we will find the optimal solution for your needs.