When dubbing a dialogue scene in an audiobook, standard TTS produces the same voice for all characters. This breaks immersion — the listener cannot distinguish the heroes. For IVR systems, podcasts, and training courses with multiple presenters, you need multi-speaker TTS: an architecture capable of switching between voices according to a script. We have implemented such systems for 15+ projects — from audiobooks to voice assistants. The average budget savings for clients is 35% compared to cloud APIs—one client saved over $15,000 annually. Contact us to discuss your scenario.
The key problem is latency when switching: if speaker embeddings are not preloaded, pauses can reach 1.5 seconds. Our record is 200 ms switching on XTTS v2. In this article, we will break down real cases, stack, and typical mistakes.
Problems We Solve
- Voice synchronization: when switching between voices, pauses and artifacts occur. We use speaker embeddings and preloading of latents to reduce latency to 200 ms.
- Acoustic space management: different voices require different processing (echo, noise). We apply post-processing based on WavLM to align acoustics.
- Dialogue scaling: for scenes with 5+ characters, maintaining voice consistency is important. We use XTTS v2 with fixed reference audio for each character.
- Latency in real-time: in voice output chatbots, speed is critical. We optimize via ONNX Runtime and batching requests.
How We Do It: Stack and Cases
Multi-speaker System Architecture
from dataclasses import dataclass
from enum import Enum
class SpeakerRole(Enum):
ASSISTANT = "assistant"
NARRATOR = "narrator"
CHARACTER_1 = "character_1"
CHARACTER_2 = "character_2"
@dataclass
class Speaker:
role: SpeakerRole
name: str
voice_config: dict
reference_audio: str | None = None
class MultiSpeakerTTS:
def __init__(self, speakers: list[Speaker]):
self.speakers = {s.role: s for s in speakers}
self._init_engines()
def synthesize(self, text: str, role: SpeakerRole) -> bytes:
speaker = self.speakers[role]
return self._synthesize_with_config(text, speaker.voice_config)
Implementation on XTTS v2
For self-hosted scenarios, we use XTTS v2 — a model from Coqui AI that supports speaker conditioning. We preload speaker latents for speed:
from TTS.api import TTS
tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda")
# Preload speaker latents for speed
SPEAKERS = {
"narrator": "voices/narrator.wav",
"alice": "voices/alice.wav",
"bob": "voices/bob.wav",
}
def synthesize_dialog(dialog: list[dict]) -> list[bytes]:
"""
dialog: [{"speaker": "alice", "text": "Hello!"},
{"speaker": "bob", "text": "Hi!"}]
"""
results = []
for line in dialog:
speaker_wav = SPEAKERS[line["speaker"]]
wav = tts.tts(
text=line["text"],
speaker_wav=speaker_wav,
language="en"
)
results.append(wav)
return results
Case: For a client's educational platform, we deployed a self-hosted solution with four voices (lecturer, student, assistant, system). Speaker latents were extracted from 3-second reference recordings. Final quality — MOS 4.2, latency p99 — 800 ms (single GPU RTX 3090). This is 2-3 times faster than cloud Azure with similar quality.
Cloud Multi-Speaker via Azure
Azure Neural TTS supports multiple voices in one SSML document — convenient for simple dialogues without a local GPU:
<speak version='1.0' xml:lang='en-US'>
<voice name='en-US-JennyNeural'>
Good afternoon! This is Jenny.
</voice>
<break time='300ms'/>
<voice name='en-US-GuyNeural'>
Hello! And this is Guy.
</voice>
</speak>
According to documentation, Azure Neural TTS allows switching voices within a single SSML document. Azure automatically handles intonation, but you cannot control speaker embeddings — only preset voices. This is a trade-off between simplicity and flexibility.
Dialogue Assembly
from pydub import AudioSegment
def assemble_dialog(audio_clips: list[bytes], pause_ms: int = 300) -> bytes:
combined = AudioSegment.empty()
silence = AudioSegment.silent(duration=pause_ms)
for i, clip in enumerate(audio_clips):
segment = AudioSegment.from_wav(io.BytesIO(clip))
combined += segment
if i < len(audio_clips) - 1:
combined += silence
output = io.BytesIO()
combined.export(output, format="mp3")
return output.getvalue()
Multi-Speaker vs Single-Speaker: Increased Complexity
Single-speaker TTS only needs one model with one voice. Multi-speaker requires:
- Managing speaker embeddings or fine-tuning for each voice.
- Minimizing latency when switching (preloading vectors).
- Handling acoustic differences (timbre, tempo, intonation) within a single pipeline.
- Checking voice consistency in long dialogues (latent drift).
At the same time, a self-hosted solution allows reducing operational costs by 40% by eliminating cloud services, especially at large synthesis volumes.
Choosing Between Cloud and Self-Hosted
| Criterion | Cloud (Azure, Google) | Self-Hosted (XTTS v2, Coqui) |
|---|---|---|
| Voice control | Only preset voices | Any reference audio |
| Latency | 500–1500 ms | 200–800 ms (with good GPU) |
| Cost | Price per character | CAPEX for GPU + electricity |
| Privacy | Data goes to cloud | Data stays on-premises |
| Scalability | High (automatic) | Requires cluster setup |
Choice depends on voice control requirements and budget. A self-hosted solution pays off in 6–12 months at synthesis volumes from 1 million characters per month.
| Multi-Speaker TTS Development Stage | Duration |
|---|---|
| Analysis and approach selection | 1-2 days |
| Reference audio preparation | 1-2 days |
| Model adaptation and testing | 3-5 days |
| Integration and deployment | 2-3 days |
| Optimization and monitoring | 1-2 days |
Get a consultation for your project. We can help you synthesize multiple voices for dialog voiceover, voice interface, or training content. Our system supports text to speech for multiple characters, making it ideal for audiobooks, podcasts, and voice assistant applications.
Example configuration for XTTS v2 with preloaded latents
import torch
from TTS.api import TTS
# Load model once
tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2").to("cuda")
# Preload speaker latents for all voices
speaker_latents = {}
for name, wav in SPEAKERS.items():
speaker_latents[name] = tts.get_speaker_latents(wav)
def fast_synthesize(text, speaker_name):
with torch.no_grad():
wav = tts.tts(text, speaker_latents=speaker_latents[speaker_name], language="en")
return wav
Our Work Process
- Analysis: determine the number of voices, use cases, latency and quality requirements. Assess whether unique voices are needed or preset ones suffice.
- Approach selection: cloud API or self-hosted? If self-hosted — choose a model (XTTS v2, VITS, Coqui).
- Reference audio preparation: record or clean audio (2–5 seconds per voice, mono, 16 kHz).
- Model adaptation: for XTTS — extract speaker latents; for Azure — simply configure SSML.
- Integration: attach synthesis to your application via REST API or gRPC.
- Testing: MOS evaluation, A/B tests with users, latency checks.
- Deployment: deploy on your server or in the cloud. Ensure monitoring and alerts.
Time Estimates
- Cloud solution: from 2 to 3 days (SSML setup, integration, tests).
- Self-hosted without fine-tuning: from 1 week (stack selection, voice loading, deployment).
- Self-hosted with voice fine-tuning: from 2 weeks (requires dataset collection, LoRA adapter training).
Cost is calculated individually — depends on number of voices, latency requirements, and chosen stack.
Checklist of Typical Mistakes
- Insufficient reference audio: stable latents require 3–5 seconds of clean voice without background noise.
- Ignoring switching latency: if speaker embeddings are not preloaded, pauses between utterances can exceed 1 second.
- Incorrect pause handling: in SSML, it's important to use
<break time="..."/>, otherwise the dialogue sounds run-on. - Lack of consistency tests: a character's voice may drift in long dialogues — fix the latent per session.
What's Included in Our Work
- Designing a multi-speaker TTS architecture for your scenario.
- Setting up and deploying the chosen engine (Azure, XTTS v2, Coqui).
- Integrating with your application (REST API, WebSocket, gRPC).
- Preparing reference audio (cleaning, normalization, segmentation).
- Quality testing (MOS, Latency p99) and optimization.
- Operational documentation and post-launch support.
We are a team with 5+ years of experience in speech synthesis, having completed over 50 projects (audiobooks, IVR, educational platforms). We guarantee quality: every system undergoes load testing and security audit.
Order multi-speaker TTS development for your scenario. Contact us — we will choose the optimal architecture and configure the voices.
This material is based on documentation from Azure Neural TTS and Coqui XTTS.







