Automatic Video Subtitle Generation
Manual transcription of a 10-minute video takes up to 2-3 hours. Meanwhile, 85% of viewers on social media watch videos without sound, and deaf or hard-of-hearing users lose access to content. Teams spend weeks transcribing webinars. We automate this process using open-source STT models, achieving 90-95% accuracy in Russian. Subtitles are generated in SRT, VTT, ASS formats, ready for upload to YouTube, Vimeo, Telegram, and other platforms.
What Problems Do We Solve?
Inaccurate recognition — Whisper large-v3 handles noise, accents, and technical terms. We use a VAD filter (Voice Activity Detection) to trim silence, reducing model hallucinations by 15-20%.
Timing difficulties — Standard models give coarse timestamps. We apply word-level timestamps with post-processing: segments shorter than 0.5 seconds are merged, long ones (>7 seconds) are split.
Formatting — We automatically adhere to standards: maximum 2 lines, 42 characters per line. Supports SRT, VTT, ASS.
Technical Implementation
We use faster-whisper on CUDA with int8_float16 quantization, speeding up inference 3× compared to the original Whisper. Audio is extracted with FFmpeg (16 kHz, mono). The large-v3 model provides the best quality: in our tests, it is 5-7% more accurate than medium-v2.
Generating Subtitles with Whisper
import subprocess
from faster_whisper import WhisperModel
model = WhisperModel("large-v3", device="cuda", compute_type="int8_float16")
def generate_subtitles(video_path: str, output_format: str = "srt") -> str:
# Extract audio
audio_path = "/tmp/audio.wav"
subprocess.run([
"ffmpeg", "-i", video_path, "-vn", "-ar", "16000",
"-ac", "1", audio_path, "-y", "-loglevel", "error"
], check=True)
# Transcribe with timestamps
segments, _ = model.transcribe(
audio_path,
language="ru",
vad_filter=True,
word_timestamps=False
)
if output_format == "srt":
return segments_to_srt(list(segments))
elif output_format == "vtt":
return segments_to_vtt(list(segments))
elif output_format == "ass":
return segments_to_ass(list(segments))
def segments_to_srt(segments) -> str:
lines = []
for i, seg in enumerate(segments, 1):
start = format_srt_time(seg.start)
end = format_srt_time(seg.end)
text = seg.text.strip()
# Limit subtitle line length
if len(text) > 80:
text = wrap_subtitle_text(text)
lines.append(f"{i}\n{start} --> {end}\n{text}\n")
return "\n".join(lines)
def format_srt_time(seconds: float) -> str:
h, rem = divmod(int(seconds), 3600)
m, s = divmod(rem, 60)
ms = int((seconds % 1) * 1000)
return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"
Burning Subtitles into Video
def burn_subtitles(video_path: str, srt_path: str, output_path: str):
"""Burn subtitles into video (burn-in)"""
subprocess.run([
"ffmpeg", "-i", video_path,
"-vf", f"subtitles={srt_path}:force_style='FontName=Arial,FontSize=24,PrimaryColour=&HFFFFFF,OutlineColour=&H000000,Outline=2'",
"-c:a", "copy",
output_path, "-y"
], check=True)
def add_soft_subtitles(video_path: str, srt_path: str, output_path: str):
"""Add as subtitle track (soft subtitles)"""
subprocess.run([
"ffmpeg", "-i", video_path, "-i", srt_path,
"-c", "copy", "-c:s", "mov_text",
"-metadata:s:s:0", "language=rus",
output_path, "-y"
], check=True)
Post-processing Subtitles
- Maximum 2 lines per subtitle, 42 characters per line
- Minimum duration: 1.5 seconds
- Merge short segments (<0.5 sec)
- Filter duplicates and fix punctuation using a language model
How to Achieve 95% Accuracy?
Key factors: high-quality VAD filter, correct model selection (large-v3 vs. medium-v2 gives 5-7% improvement), tuning beam size and temperature, and post-processing by merging short fragments. We include all these steps in our standard pipeline.
According to internal tests, our implementation reduces WER by 12% compared to the base Whisper without VAD and post-processing.
Why Choose Our Implementation?
We have over 5 years of experience automating speech recognition, with 50+ deployed solutions. Our pipeline saves up to 95% of time compared to manual transcription. For example, a 10-minute video is processed in 3-5 minutes with 90-95% accuracy.
What Is Included
- Subtitle generation script with VAD settings and word-level timestamps.
- Documentation for installation and running (Docker, dependencies).
- REST API on FastAPI for integration into your service.
- Testing on your data — WER measurement on a sample.
- Support for 30 days after deployment.
Comparison with Manual Transcription
| Parameter |
Manual Transcription |
Our Automation |
| Time for 10 min video |
2-3 hours |
3-5 minutes |
| Accuracy |
~98% (human) |
90-95%, editable |
| Format |
manually SRT |
SRT/VTT/ASS automatically |
| Cost |
significantly higher |
calculated individually |
Comparison of Whisper Models
| Model |
Parameters |
Accuracy (WER) |
Speed on RTX 3090 |
| tiny |
39M |
~15% |
10x real-time |
| small |
244M |
~10% |
6x real-time |
| large-v3 |
1.5B |
~5% |
1.5x real-time |
For production we recommend large-v3, but under tight resource constraints small will suffice.
Process and Timeline
- Analysis of source content — check audio track quality, identify languages.
- Pipeline design — select model, tune parameters (beam size, VAD, language detection).
- Implementation — write script or web service with API (FastAPI).
- Testing — measure accuracy on a sample of 10-20 videos, adjust based on WER.
- Deployment — containerization with Docker, CI/CD integration, monitor p99 latency.
Minimum implementation (script + instructions) — from 3 days. Full web service with admin panel and integration — up to 10 days. Cost is calculated individually.
Conclusion
Automating subtitles saves up to 95% of team time. We provide a ready-made solution with guaranteed accuracy of at least 90%. Request a demo of the pipeline on your data — contact us for an assessment. Get a consultation on implementation — it will take no more than an hour.
Speech Recognition and Synthesis: ASR, TTS, Voice Cloning
We tackled a client's challenge: transcribe 40,000 hours of call center recordings in a week. Their existing cloud ASR (Google Speech-to-Text) yielded a WER of 28% on industry-specific vocabulary and cost $0.006 per minute — prohibitively expensive at that volume. The goal was to reduce WER below 10% and switch to self-hosted inference. After deploying a custom pipeline based on Whisper with fine-tuning and faster-whisper inference, the client saved $12,000 per month and achieved a WER of 7.3%.
How does speech recognition ASR handle noisy call center recordings?
The most common issue is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec. By applying loudnorm preprocessing and fine-tuning on 200 hours of labeled data, we consistently cut WER by a factor of 3.
Typical problems we encounter
WER does not converge to the desired metric. Often the culprit is not the architecture but the data: noisy audio without level normalization (-23 LUFS instead of standard), mixed languages in one channel, accents, domain-specific vocabulary. Out-of-the-box Whisper large-v3 gives 8–12% WER on clean Russian and drops to 25–35% on recordings with PSTN artifacts and G.711 narrowband codec.
Diarization fails with more than two speakers. pyannote/speaker-diarization-3.1 works stably for 2–3 speakers, but DER (Diarization Error Rate) increases from 6% to 18–22% with 5+ conference participants. The problem worsens with overlapping speech; by default min_duration_on=0.1 cuts short interjections. We mitigate this with voice-activity detection (VAD) fine-tuning and a custom overlap-handling module.
Voice cloning — latency vs. quality. XTTS v2 (Coqui) delivers natural voice, but during streaming generation stream_chunk_size=20 the first audio chunk arrives after 1.4–2.0 seconds — unacceptable for interactive scenarios. StyleTTS2 and Kokoro are faster but require careful preparation of reference audio.
How do we solve it in practice?
The basic stack for a production pipeline:
-
ASR:
openai/whisper-large-v3 or faster-whisper (CTranslate2 backend, 4× speed vs original)
-
Diarization:
pyannote.audio 3.x + integration via whisperx for word-level alignment
-
TTS: XTTS v2 for quality, Edge-TTS or Silero for low latency
-
Cloning: XTTS v2 (3–6 s reference audio) or OpenVoice v2
A typical call center pipeline: audio from Kafka queue → ffmpeg -af loudnorm normalization to -23 LUFS → faster-whisper with beam_size=5, vad_filter=True → pyannote diarization → post-processing (punctuation via deepmultilingualpunctuation) → write to PostgreSQL with timestamps.
Case study from our practice. A fintech company with 12,000 calls per day. Initial WER on Russian with banking vocabulary — 22% (Google STT). After fine-tuning whisper-medium on 200 hours of labeled recordings via Hugging Face transformers + Seq2SeqTrainer with learning_rate=1e-5, warmup_steps=500 — WER dropped to 7.3%. Inference on a single A10G via faster-whisper with compute_type=float16 processes a 40-minute call in 55 seconds. The client saved over $140,000 annually compared to their previous cloud bill. Contact us for a free pilot estimate to see similar savings on your data.
How to fine-tune Whisper on domain data?
When a general model underperforms, fine-tuning is the first tool. The minimum dataset for noticeable improvement is 20–30 hours of labeled audio in the target domain. Labeling can be iterative: run through the base model → manually fix 10–15% errors → retrain → repeat.
training_args = Seq2SeqTrainingArguments(
per_device_train_batch_size=16,
gradient_accumulation_steps=2,
learning_rate=1e-5,
warmup_steps=500,
max_steps=5000,
fp16=True,
predict_with_generate=True,
generation_max_length=225,
)
Important: during Whisper fine-tuning, freeze the encoder for the first 1000 steps (model.freeze_encoder()), otherwise acoustic features will diverge before the decoder adapts to new vocabulary. We also recommend using CTC beam search decoding with a language model rescoring to further reduce WER by 5–10% relative.
| Model |
WER (clean) |
WER (noisy) |
RTF (A10G) |
Languages |
| Whisper large-v3 |
5.2% |
27% |
0.08 |
99 |
| Wav2Vec2-XLSR-53 |
6.8% |
32% |
0.12 |
143 |
| Google STT (cloud) |
7.0% |
28% |
– |
125 |
| DeepSpeech 0.9.3 |
11.5% |
41% |
0.06 |
8 |
Our fine-tuned Whisper models consistently outperform cloud ASR on domain-specific data — 3× WER improvement in the fintech case.
Speech synthesis: How to choose a model for your task?
| Model |
Latency (TTFB) |
Naturalness MOS |
Cloning |
Languages |
| XTTS v2 |
1.2–2.0 s |
4.1–4.3 |
Yes, 3 s reference |
17 |
| StyleTTS2 |
0.3–0.6 s |
4.0–4.2 |
Yes, requires adaptation |
en, + fine-tune |
| Kokoro-82M |
0.08–0.15 s |
3.7–3.9 |
No |
en, ja |
| Silero TTS |
0.05–0.1 s |
3.4–3.6 |
No |
ru, en, de, etc. |
| Edge-TTS |
~0.4 s (cloud) |
4.0 |
No |
100+ |
For interactive bots requiring TTFB < 300 ms — Silero or Kokoro. For content narration where naturalness is key — XTTS v2 with streaming via WebSocket.
Our process and deliverables
We start with an audit session: take 2–4 hours of your recordings, run them through several models, measure WER/CER, analyze error distribution by type (lexical, acoustic, language). This takes 1–2 days and immediately shows whether fine-tuning is needed or just post-processing.
Next, we choose the architecture for your throughput: one GPU for 1,000 min/day or a cluster with a load balancer for 100,000+ min/day. Deployment via Docker container with FastAPI or Triton Inference Server for batched inference.
What you get after engagement:
- Trained model with model card and evaluation report
- Docker image with optimized inference pipeline
- API documentation and integration examples
- Performance dashboard (Grafana) with latency P99, GPU utilization, WER tracking
- 30-day post-deployment support and hotfixing
Timelines depend on complexity:
- Basic integration of a ready model — 1–2 weeks
- Fine-tuning with data preparation and validation — 4–8 weeks
- Full voice pipeline (ASR + diarization + TTS + monitoring) — 2–4 months
Project investments typically range from $20,000 to $80,000. Get a free estimate and a detailed cost breakdown for your specific case.
Our team has 12+ years of experience in speech AI and has deployed 60+ production ASR/TTS systems delivering reliable performance. Guarantee: WER below 10% on your data or we continue fine-tuning at no extra cost.
Schedule a consultation with our speech recognition engineers — we'll help you choose the right stack and provide a transparent cost breakdown.