Real-Time Live Captions: Architecture, Latency, Integration
Real-time captions (live captions) are a technical rehabilitation tool per WCAG 2.1 (criterion 1.2.4) and the equivalent Russian standard GOST R 52872-2019. We build captioning systems that operate with less than 2 seconds of delay. This is critical for broadcasts, conferences, television, and educational platforms. Our team — 12 engineers with a combined experience of over 25 years in STT and NLP. We have implemented 10+ installations for events with audiences up to 5,000 people and provided information access for thousands of hearing-impaired users. Deploying AI captions can save up to 40% budget compared to manual captioning, with payback within 2–3 months for regular broadcasts. To assess your project, contact us — we will offer a turnkey solution.
Problems We Solve
Standard captions often lag 5–10 seconds behind speech — unacceptable for the hard of hearing. Typical challenges:
- Text-audio synchronization suffers when using batch audio processing.
- Cloud STT services don't always handle domain-specific vocabulary (medical, legal terminology).
- Integration with platforms like Zoom and Teams requires a separate bot and API setup.
We solve these by choosing streaming models, optimizing buffering, and customizing vocabulary. For instance, in one telemedicine conference project we fine-tuned Whisper on a medical term corpus — recognition accuracy rose from 82% to 95%.
How We Achieve Under 2 Seconds Latency
The key is choosing a streaming model and audio transmission architecture. Deepgram Nova-2 delivers partial results every 200ms, giving end-to-end latency around 1 second — 2–3 times faster than faster-whisper large-v3 in batch mode. For on-premise scenarios we use faster-whisper with a VAD filter and 3-second buffer, delivering 2.5–4 seconds. But if under 2 seconds is required — only cloud streaming.
Real-Time STT Stack
For captions with <2s latency from speech onset we use local faster-whisper or cloud Deepgram Nova-2. Core engine example in Python:
import asyncio
import websockets
from faster_whisper import WhisperModel
import numpy as np
import sounddevice as sd
class RealTimeCaptioner:
def __init__(self):
self.model = WhisperModel(
"large-v3",
device="cuda",
compute_type="float16"
)
self.buffer = []
self.chunk_duration = 3.0 # seconds of buffering
self.sample_rate = 16000
async def stream_captions(self, websocket, audio_queue: asyncio.Queue):
"""Stream captions via WebSocket"""
while True:
chunk = await audio_queue.get()
self.buffer.append(chunk)
buffer_duration = len(self.buffer) * len(chunk) / self.sample_rate
if buffer_duration >= self.chunk_duration:
audio_data = np.concatenate(self.buffer)
self.buffer = []
segments, _ = self.model.transcribe(
audio_data,
language="ru",
vad_filter=True,
vad_parameters={"min_silence_duration_ms": 500}
)
for segment in segments:
caption = {
"text": segment.text.strip(),
"start": segment.start,
"end": segment.end,
"confidence": segment.avg_logprob
}
await websocket.send(json.dumps(caption, ensure_ascii=False))
WebRTC Integration for Browser
The client side in JavaScript captures audio from the microphone and sends it to the server via WebSocket. The server returns captions, displayed with a rolling window.
// Client side: audio capture and streaming to server
class LiveCaptionClient {
constructor(wsUrl) {
this.ws = new WebSocket(wsUrl);
this.captionDiv = document.getElementById('captions');
}
async startCapturing() {
const stream = await navigator.mediaDevices.getUserMedia({
audio: { sampleRate: 16000, channelCount: 1, echoCancellation: true }
});
const audioContext = new AudioContext({ sampleRate: 16000 });
const processor = audioContext.createScriptProcessor(4096, 1, 1);
processor.onaudioprocess = (event) => {
const pcmData = event.inputBuffer.getChannelData(0);
const int16Array = new Int16Array(pcmData.length);
for (let i = 0; i < pcmData.length; i++) {
int16Array[i] = Math.max(-32768, Math.min(32767, pcmData[i] * 32768));
}
if (this.ws.readyState === WebSocket.OPEN) {
this.ws.send(int16Array.buffer);
}
};
this.ws.onmessage = (event) => {
const caption = JSON.parse(event.data);
this.displayCaption(caption.text);
};
const source = audioContext.createMediaStreamSource(stream);
source.connect(processor);
processor.connect(audioContext.destination);
}
displayCaption(text) {
// Rolling-window display (last 2-3 lines)
const line = document.createElement('p');
line.textContent = text;
line.className = 'caption-line';
this.captionDiv.appendChild(line);
// Remove old lines
while (this.captionDiv.children.length > 3) {
this.captionDiv.removeChild(this.captionDiv.firstChild);
}
// Auto-scroll
this.captionDiv.scrollTop = this.captionDiv.scrollHeight;
}
}
How to Choose an STT Model for Captions?
Comparison between on-premise and cloud solutions:
| Parameter | faster-whisper (on-premise) | Deepgram Nova-2 (cloud) |
|---|---|---|
| Latency | 0.3–0.8 sec (inference) | 0.1–0.3 sec (streaming) |
| Quality | high (large-v3) | high (specialized) |
| Privacy | full | data leaves to cloud |
| Cost | one GPU (~$0.5/hr) | $0.004/min audio |
| Language support | 99+ | 30+ |
For tasks requiring full data isolation (medical, government) — on-premise faster-whisper with Triton Inference Server. For typical broadcasts — cloud Deepgram or AssemblyAI. We will assess your project and suggest the optimal option. Request a preliminary audit — it is free and takes 30 minutes.
Display Requirements (WCAG 2.1)
/* Captions for hearing-impaired — WCAG 2.1 criterion 1.4.3 */
.caption-container {
background-color: rgba(0, 0, 0, 0.85);
color: #FFFFFF;
font-size: 1.5rem; /* minimum 24px */
line-height: 1.6;
padding: 12px 20px;
border-radius: 4px;
max-width: 80%;
font-family: Arial, sans-serif; /* high legibility */
}
/* High contrast (ratio 7:1 for AA+) */
.caption-line {
color: #FFFFFF;
text-shadow: 1px 1px 2px #000;
}
Integration with Zoom/Teams via Bot
# Zoom uses RTMP for streaming captions
import httpx
async def push_zoom_captions(meeting_id: str, caption_text: str, seq: int):
"""Send captions to Zoom via Closed Caption API"""
async with httpx.AsyncClient() as client:
await client.post(
f"https://api.zoom.us/v2/meetings/{meeting_id}/live_streaming/captions",
json={"text": caption_text, "seq": seq, "lang": "ru-RU"},
headers={"Authorization": f"Bearer {ZOOM_JWT_TOKEN}"}
)
What Is Streaming Transcription?
Streaming transcription is a technology where the model recognizes speech as audio fragments arrive, outputting partial results every 100–200ms. This allows captions to update smoothly, without pauses for the complete phrase. We use WebSocket to transmit audio fragments and receive text with timestamps.
Implementation Checklist
- [ ] Audit current audio channels and platform
- [ ] Choose STT model (on-premise/cloud)
- [ ] Calibrate for domain vocabulary
- [ ] Develop WebRTC/WebSocket server
- [ ] Integrate with Zoom, Teams, YouTube Live, RTMP
- [ ] Management interface and manual correction
- [ ] Documentation and operator training
- [ ] 6-month warranty
Estimated Timelines
Web captioning component — 1–2 weeks. Full platform integration — 2–3 weeks. To assess your project, contact us — describe the scenario and we will calculate the solution. Budget savings can reach 40% compared to manual captioning, with payback in 2–3 months. Get a consultation now!







