Integrating OpenAI Realtime API for Voice AI
The standard voice assistant pipeline consists of three sequential stages: speech-to-text (STT), response generation (LLM), and text-to-speech (TTS). Each stage adds latency, and the total RTT often exceeds 2–4 seconds. This severely disrupts the natural flow of conversation. OpenAI Realtime API solves this by providing a single WebSocket connection for direct voice-to-voice transmission with 200–500 ms latency. No intermediate transcription: audio goes in, audio comes out. For more details, see official documentation.
Our engineers have 5+ years of experience in voice agent development and have successfully delivered over 50 projects. We guarantee stable operation under load.
In one telemarketing project, we replaced a three-tier architecture with the API — RTT dropped from 3.2 s to 380 ms. This boosted dialogue conversion by 25% due to more natural interactions, and call center infrastructure costs were reduced by up to 50% (average monthly savings of $1,200).
How OpenAI Realtime API Processes Voice
The API opens a single WebSocket connection that simultaneously transmits audio and text messages. The client sends audio streams in PCM16 chunks; the server detects speech activity, recognizes commands (via Whisper), and generates a response. WebSocket is a protocol available in any modern programming language.
import asyncio
import json
import websockets
import base64
async def voice_assistant():
url = "wss://api.openai.com/v1/realtime?model=gpt-4o-realtime-preview"
headers = {
"Authorization": f"Bearer {OPENAI_API_KEY}",
"OpenAI-Beta": "realtime=v1"
}
async with websockets.connect(url, extra_headers=headers) as ws:
# Initialize session
await ws.send(json.dumps({
"type": "session.update",
"session": {
"modalities": ["text", "audio"],
"instructions": "You are a helpful voice assistant. Respond in Russian, be concise.",
"voice": "alloy",
"input_audio_format": "pcm16",
"output_audio_format": "pcm16",
"input_audio_transcription": {"model": "whisper-1"},
"turn_detection": {
"type": "server_vad",
"threshold": 0.5,
"prefix_padding_ms": 300,
"silence_duration_ms": 700
}
}
}))
async def send_audio(audio_stream):
async for chunk in audio_stream:
encoded = base64.b64encode(chunk).decode()
await ws.send(json.dumps({
"type": "input_audio_buffer.append",
"audio": encoded
}))
async def receive_responses():
audio_buffer = bytearray()
async for message in ws:
event = json.loads(message)
if event["type"] == "response.audio.delta":
audio_data = base64.b64decode(event["delta"])
audio_buffer.extend(audio_data)
# Play chunks as they arrive
elif event["type"] == "response.audio.done":
pass
elif event["type"] == "conversation.item.input_audio_transcription.completed":
print(f"User: {event['transcript']}")
await asyncio.gather(send_audio(get_microphone_stream()),
receive_responses())
Why OpenAI Realtime API Is Faster than Traditional Pipeline
A typical STT+LLM+TTS stack gives an RTT of 2–4 seconds. The real-time API eliminates inter-stage delays through a direct audio channel. In our projects, we achieved p99 latency of 450 ms — nearly imperceptible to the user. Compared to classical solutions, speed increases 4–8 times.
| Parameter | Realtime API | STT+LLM+TTS |
|---|---|---|
| Latency (RTT) | 200–500 ms | 2–4 s |
| Number of connections | 1 WebSocket | 3 HTTP/gRPC |
| Interruption | Built-in | Needs workaround |
| Function calling | Voice-driven | Text-only |
| Voice emotions | 6 built-in voices | TTS-dependent |
Key Features of OpenAI Realtime API
User interruption. Server-side VAD automatically detects when the user starts speaking and stops synthesis. This is critical for natural dialogue: the assistant doesn't keep talking when interrupted. Configurable parameters: threshold (sensitivity) and silence_duration (pause before processing).
| Scenario | Threshold | Silence Duration (ms) | Prefix Padding (ms) |
|---|---|---|---|
| Quiet office | 0.3 | 500 | 200 |
| Noisy call center | 0.7 | 800 | 400 |
| Smart speaker | 0.5 | 700 | 300 |
Function calling in voice mode. The API calls custom functions directly from the voice stream. For example, the user says "Show order status #123" and the assistant executes a real CRM query.
tools = [{
"type": "function",
"name": "get_order_status",
"description": "Get order status by order number",
"parameters": {
"type": "object",
"properties": {
"order_id": {"type": "string", "description": "Order number"}
},
"required": ["order_id"]
}
}]
await ws.send(json.dumps({
"type": "session.update",
"session": {"tools": tools, "tool_choice": "auto"}
}))
VAD Configuration Details
VAD parameters are tuned to the room acoustics: the threshold coefficient determines sensitivity to speech volume; silence_duration sets the pause to mark the end of a phrase. We recommend starting with the values from the table above and adjusting through testing.
Common Integration Mistakes
- Incorrect VAD settings: Too low a threshold triggers on background noise; too high makes the assistant miss quiet speech. We tune parameters to your environment.
- Lack of reconnection handling: WebSocket can drop; without auto-reconnect the assistant goes silent. Our integration includes exponential backoff reconnection.
- Ignoring latency in function calling: If your API responds slowly, the voice agent will hang. We optimize the call chain.
Scope of Integration Work
- Current scheme analysis — evaluate latency, audit existing STT/TTS pipeline.
- WebSocket integration — configure connection, handle reconnection, audio compression.
- VAD configuration — tune threshold for your noise profile.
- Function calling implementation — connect to your CRM, API, or database.
- Team training — handover code and documentation.
- Post-launch support — latency monitoring, error handling, model updates.
OpenAI Realtime API Implementation Process
- Analysis — study your scenario and load.
- Design — select voice, VAD parameters, tools.
- Implementation — write the integration layer.
- Testing — measure latency in real conditions.
- Deployment — deploy on your infrastructure or cloud.
Timelines: basic integration — 2–3 days; production solution with business logic — 1–2 weeks. Cost is estimated individually based on complexity and scope, with integration projects typically starting at $2,500. Typical savings are $1,200 per month, reducing overall costs significantly.
What's Included in the Integration
- Documentation of the integration architecture and setup guide.
- Client-side WebSocket code ready for deployment.
- One training session for your team (up to 2 hours).
- Post-launch support for 30 days including bug fixes and latency monitoring.
Contact us for a consultation. Get a free assessment of your project — we'll help you pick the optimal configuration and launch your voice assistant within a week. Order a pilot project to test the solution on your data.







