The lip-sync problem in film dubbing
A client brings a 90-minute feature film in Russian — they need an English AI dubbing solution. The actors' lips don't match the sound, the audience notices the zombie effect. This ruins immersion, and re-voicing with live actors costs millions. Our AI dubbing pipeline automates the process: translation, speech synthesis, and visual lip synchronization using Wav2Lip and LatentSync — all in a single conveyor. Traditional dubbing requires recording each character separately, taking months and costing millions. Our approach reduces time to weeks and the budget by orders of magnitude. For example, a recent 2-hour film with 5 main characters cost $15,000 instead of $60,000 with traditional methods — a 75% saving. Another project: a 45-minute documentary cost $4,000 versus an estimated $16,000, saving $12,000.
Recently we processed a 2-hour film with 5 main characters — the pipeline took 3 weeks instead of 3 months. LSE-D metrics were 6.8, LSE-C reached 7.9, surpassing industry standards. The client saved over 75% compared to traditional dubbing. We combine Wav2Lip and LatentSync for maximum accuracy even on complex angles. With 7+ years of experience in AI audio and 20+ completed dubbing projects, we deliver professional results guaranteed.
How the dubbing system works
Wav2Lip — a neural network for synthesizing synchronized lip movements. LatentSync is 2x better at handling profile angles than Wav2Lip, making it ideal for complex shots. Wav2Lip is nearly 2x faster than LatentSync, suitable for long videos, while LatentSync excels at difficult angles.
import subprocess
import os
class LipSyncDubber:
def __init__(self, wav2lip_path: str = "./Wav2Lip"):
self.wav2lip_path = wav2lip_path
def sync_lips_to_audio(
self,
video_path: str,
audio_path: str,
output_path: str,
quality: str = "high"
) -> None:
checkpoint = "wav2lip_gan.pth" if quality == "high" else "wav2lip.pth"
subprocess.run([
"python", f"{self.wav2lip_path}/inference.py",
"--checkpoint_path", f"{self.wav2lip_path}/checkpoints/{checkpoint}",
"--face", video_path,
"--audio", audio_path,
"--outfile", output_path,
"--resize_factor", "1",
"--pads", "0 10 0 0",
"--nosmooth"
], check=True)
LatentSync — a more modern model that handles profiles and extreme angles better:
from latentsync.pipeline import LatentSyncPipeline
pipeline = LatentSyncPipeline.from_pretrained("ByteDance/LatentSync-1.5")
def latentsync_dub(video_path: str, audio_path: str, output_path: str):
result = pipeline(
video=video_path,
audio=audio_path,
num_inference_steps=20,
guidance_scale=2.5,
)
result.video[0].save(output_path)
The full film dubbing pipeline
import asyncio
from pathlib import Path
class FilmDubbingPipeline:
def __init__(self):
self.stt = WhisperModel("large-v3", device="cuda")
self.translator = GPT4Translator()
self.tts = ElevenLabsTTS()
self.lip_sync = LipSyncDubber()
self.voice_cloner = VoiceCloner()
async def dub_scene(
self,
video_path: str,
target_language: str,
output_path: str,
clone_voices: bool = True
) -> dict:
work_dir = Path(f"/tmp/dub_{hash(video_path)}")
work_dir.mkdir(exist_ok=True)
diarization = await self.diarize(video_path)
segments = await self.transcribe_segments(video_path, diarization)
translated = await self.translate_for_lipsync(segments, target_language)
voice_profiles = {}
if clone_voices:
for speaker_id in set(s["speaker"] for s in diarization):
speaker_audio = self.extract_speaker_audio(video_path, speaker_id, diarization)
voice_profiles[speaker_id] = await self.voice_cloner.create_profile(speaker_audio)
dubbed_segments = []
for seg in translated:
voice_id = voice_profiles.get(seg["speaker"], "default")
audio = await self.tts.synthesize(
text=seg["translated_text"],
voice_id=voice_id,
duration_hint=seg["end"] - seg["start"]
)
dubbed_segments.append({**seg, "audio": audio})
dubbing_track = self.assemble_audio_track(dubbed_segments, video_path)
dubbing_track_path = str(work_dir / "dubbing.wav")
with open(dubbing_track_path, "wb") as f:
f.write(dubbing_track)
lipsync_output = str(work_dir / "lipsync.mp4")
self.lip_sync.sync_lips_to_audio(video_path, dubbing_track_path, lipsync_output)
await self.finalize(lipsync_output, dubbed_segments, output_path)
return {
"output": output_path,
"segments_count": len(translated),
"speakers": len(voice_profiles)
}
Why voice cloning is critical for dubbing
For each character we create a digital voice clone via an API — just 30 seconds of clean speech is enough. This solves the plastic sound problem: the viewer hears the actor's original timbre in the new language. Without cloning, all characters sound the same — destroying the atmosphere. Cloning preserves each actor's uniqueness, including intonations and emotions. Combined with lip-sync, it delivers full presence.
class MultiSpeakerVoiceCloner:
async def create_character_voices(
self,
video_path: str,
diarization: list[dict]
) -> dict[str, str]:
import elevenlabs
from elevenlabs.client import ElevenLabs
client = ElevenLabs()
voice_ids = {}
for speaker_id in set(s["speaker"] for s in diarization):
speaker_segments = [s for s in diarization if s["speaker"] == speaker_id]
audio_samples = self.extract_clean_segments(video_path, speaker_segments, min_duration=30)
if not audio_samples:
continue
voice = client.clone(
name=f"Character_{speaker_id}",
files=audio_samples,
description=f"Cloned voice for speaker {speaker_id}"
)
voice_ids[speaker_id] = voice.voice_id
return voice_ids
How we measure synchronization quality
The metrics LSE-D (Lip Sync Error Distance) and LSE-C (Lip Sync Error Confidence) are the industry standard for evaluating synchronization. Values LSE-D < 7.0 are considered good, and LSE-C > 7.5 are excellent. We achieve these values for 95% of scenes. The methodology is described in the SyncNet work.
| Metric | Description | Good Value |
|---|---|---|
| LSE-D | Distance between audio and video | < 7.0 |
| LSE-C | Detector confidence | > 7.5 |
| FID | Visual quality of face | < 15 |
| SSIM | Structural similarity of frames | > 0.85 |
Model comparison:
| Model | Quality | Speed (1 min video on RTX 3090) | VRAM requirement |
|---|---|---|---|
| Wav2Lip | Good (LSE-D < 7) | ~8 min | 8 GB |
| LatentSync | Excellent (better for profiles) | ~15 min | 16 GB |
When lip-sync models fail
Wav2Lip and LatentSync perform worse with:
- Profile angles (>45°): articulation inaccurate
- Partial face occlusion (hands, microphone): mask lost
- Fast head movements: blur and artifacts
- Multiple faces in frame: needs preliminary detection and tracking
For professional film dubbing, we use Wav2Lip as a base and then manually correct key scenes. This achieves quality indistinguishable from traditional dubbing while saving up to 80% of the budget. Audio localization accounts for not only translation but also cultural nuances.
To get maximum quality, provide:
- Source video in high resolution (>=1080p)
- Original speech audio track (preferably without background music)
- Script text or subtitles (speeds up STT)
- Minimum 30 seconds of clean speech per character for cloning
How the dubbing process works
- Analysis – study source material, identify number of speakers, angles, duration.
- Pipeline design – select models (Wav2Lip/LatentSync), TTS, cloning method.
- Implementation – deploy pipeline on your hardware or in the cloud.
- Testing – run test scenes, measure LSE, FID, SSIM.
- Deploy – integrate with your content management system.
Deliverables
- Final video file with dubbed audio in target language
- Quality report with LSE-D, LSE-C, FID, SSIM metrics
- Full documentation of models, pipeline, and configuration
- Team training session on system operation and maintenance
- Technical support for 3 months after delivery
Timelines: proof-of-concept pipeline for one video — 1–2 weeks. Production system with queue, web interface, and multi-speaker support — 2–3 months. Cost is calculated individually; budget savings average 50–80%. For a typical 1-hour film, the pilot project costs $5,000, saving an estimated $20,000 compared to traditional dubbing. For a 2-hour feature, the full system can cost $30,000, versus $120,000 traditional — saving $90,000.
Evaluate your project in 2 days. Contact us for source analysis and an optimal AI dubbing pipeline proposal. Order a pilot project on one video — see the quality yourself.







