Voice Synthesis in Mobile Apps: Configuring OpenAI TTS
Voice synthesis (text-to-speech) — a technology that converts text into speech. A mobile app needs to voice long texts—news, audiobooks, voice prompts. Without optimization, users wait 3–5 seconds before playback starts. OpenAI TTS solves this, but only with proper streaming and caching. We’ll show you how to achieve sub-second latency on iOS and Android.
In this article we break down the integration architecture: from a simple REST request to streaming playback with ExoPlayer and AVAudioPlayer. We show how to cache synthesized audio and handle texts longer than 4096 characters. The result is a ready-to-deploy solution that can be integrated in 3–10 days.
What Problems Do We Solve?
Three main challenges arise: high latency, cost, and length limits. First, without streaming, you must wait for the entire file to load. Second, re-synthesizing identical text wastes API credits. Third, OpenAI TTS accepts up to 4096 characters per request, so long texts require splitting. Our solution tackles each through caching, streaming, and sentence-based splitting.
How the OpenAI TTS API Works
POST https://api.openai.com/v1/audio/speech
Authorization: Bearer {api_key}
Content-Type: application/json
{
"model": "tts-1-hd",
"input": "Your text here",
"voice": "nova",
"response_format": "mp3",
"speed": 1.0
}
According to OpenAI documentation, two models are available. tts-1 is faster, slightly lower quality, cheaper ($15/million characters). tts-1-hd is higher quality, about 30% slower, more expensive ($30/million characters). Voices: alloy (neutral), echo (male soft), fable (British), onyx (male deep), nova (female lively), shimmer (female calm). For Russian, nova and shimmer sound most natural. The speed parameter ranges from 0.25 to 4.0, default is 1.0; values above 1.3 start to break prosody.
| Characteristic | tts-1 | tts-1-hd |
|---|---|---|
| Quality | Standard | High |
| Latency | Minimal | Slight |
| Cost | Economical | Premium |
| Recommendation | Short phrases | Long texts |
| Voice | Gender | Style | Russian Recommendation |
|---|---|---|---|
| alloy | neutral | moderate | no |
| echo | male | soft | yes |
| fable | male | British | no |
| onyx | male | deep | yes |
| nova | female | lively | yes (best) |
| shimmer | female | calm | yes |
Why Caching Is Critical for UX
Every API call takes time and costs money. Caching avoids re-synthesizing the same text. For UI strings (greetings, hints), we pre-generate audio on first launch and cache it permanently. Estimated savings: with active caching, API costs can be reduced up to 40%, which translates to saving $200–$500 per month for active apps. Caching cuts API calls by up to 80% for repeated phrases, making it 5 times more cost-effective.
// iOS: cached synthesized audio
class TTSCache {
private let cacheURL: URL
init() {
cacheURL = FileManager.default.urls(for: .cachesDirectory, in: .userDomainMask)[0]
.appendingPathComponent("tts_cache")
try? FileManager.default.createDirectory(at: cacheURL, withIntermediateDirectories: true)
}
func key(text: String, voice: String) -> String {
let input = "\(text)|\(voice)"
return SHA256.hash(data: Data(input.utf8)).hexString
}
func get(_ key: String) -> Data? {
let url = cacheURL.appendingPathComponent(key + ".mp3")
return try? Data(contentsOf: url)
}
func set(_ key: String, data: Data) {
let url = cacheURL.appendingPathComponent(key + ".mp3")
try? data.write(to: url)
}
}
Before each TTS request, check the cache. A cache hit means instant playback.
Non-Streaming Implementation (for Short Texts)
// iOS: load and play
func speak(text: String, voice: String = "nova") async throws {
var request = URLRequest(url: URL(string: "https://api.openai.com/v1/audio/speech")!)
request.httpMethod = "POST"
request.setValue("Bearer \(apiKey)", forHTTPHeaderField: "Authorization")
request.setValue("application/json", forHTTPHeaderField: "Content-Type")
let body = TTSSpeechRequest(model: "tts-1", input: text, voice: voice, responseFormat: "mp3")
request.httpBody = try JSONEncoder().encode(body)
let (data, _) = try await URLSession.shared.data(for: request)
audioPlayer = try AVAudioPlayer(data: data)
audioPlayer?.play()
}
For short phrases (under 100 characters) on tts-1, latency is ~300–500 ms—acceptable without streaming. For long texts, streaming is needed.
Example of streaming playback on Android (ExoPlayer)
class OpenAITTSStreamer(private val apiKey: String, private val context: Context) {
private val exoPlayer = ExoPlayer.Builder(context).build()
fun speak(text: String, voice: String = "nova") {
val requestBody = JSONObject().apply {
put("model", "tts-1")
put("input", text)
put("voice", voice)
put("response_format", "mp3")
}.toString().toRequestBody("application/json".toMediaType())
// Use OkHttp as DataSource via a custom MediaSource
val call = OkHttpClient().newCall(
Request.Builder()
.url("https://api.openai.com/v1/audio/speech")
.header("Authorization", "Bearer $apiKey")
.post(requestBody)
.build()
)
call.enqueue(object : Callback {
override fun onResponse(call: Call, response: Response) {
// Write stream to temporary file, start playback simultaneously
val tempFile = File(context.cacheDir, "tts_${System.currentTimeMillis()}.mp3")
response.body!!.byteStream().use { input ->
tempFile.outputStream().use { output ->
val buffer = ByteArray(8192)
var bytes: Int
var firstChunk = true
while (input.read(buffer).also { bytes = it } != -1) {
output.write(buffer, 0, bytes)
if (firstChunk && tempFile.length() > 32768) {
firstChunk = false
// Start playback after first 32 KB
Handler(Looper.getMainLooper()).post {
exoPlayer.setMediaItem(MediaItem.fromUri(tempFile.toUri()))
exoPlayer.prepare()
exoPlayer.play()
}
}
}
}
}
}
override fun onFailure(call: Call, e: IOException) { /* handle error */ }
})
}
}
ExoPlayer supports playback from a file that is still being written—ProgressiveMediaSource reads data as it arrives. Latency to first audio is 400–700 ms. Streaming reduces the time to first audio by 3–5 times compared to waiting for the full download.
How to Handle Long Texts?
OpenAI TTS accepts up to 4096 characters per request. For long texts—split by sentences:
func splitBySentences(_ text: String, maxLength: Int = 1000) -> [String] {
var chunks: [String] = []
var current = ""
for sentence in text.components(separatedBy: CharacterSet(charactersIn: ".!?\n")) {
let trimmed = sentence.trimmingCharacters(in: .whitespaces)
if trimmed.isEmpty { continue }
if current.count + trimmed.count > maxLength {
if !current.isEmpty { chunks.append(current) }
current = trimmed
} else {
current += (current.isEmpty ? "" : ". ") + trimmed
}
}
if !current.isEmpty { chunks.append(current) }
return chunks
}
Synthesize chunks in parallel using TaskGroup, play sequentially—the overall latency is lower than sequential processing.
What's Included (Deliverables)
- Requirements analysis and architecture design
- Implementation of REST and streaming requests to OpenAI TTS
- Device-side caching (LRU cache with hashing)
- Long text handling (sentence-based splitting)
- Player integration (AVAudioPlayer / ExoPlayer) with background playback support
- Compliance with App Store Review Guidelines (Section 4.2) and AVAudioSession setup for iOS
- Testing on real devices (iOS 15+ / Android 10+)
- Operations documentation and API key setup guide
- Code repository access with full commit history
- Team training session (up to 2 hours) via video call
- 2 weeks of post-delivery support
Work Process
- Analysis — Study your app, identify TTS call points, measure current latencies.
- Prototyping — Build an MVP on a single screen, demonstrate streaming with cache.
- Integration — Embed the ready-made modules into your codebase.
- Testing — Load test, fine-tune for your use cases.
- Deployment — Publish to App Store / Google Play, monitor.
Timelines and Cost
Basic integration (REST + cache) — 3–4 days ($1,500). Extended integration (streaming + long text handling + UI) — 7–10 days ($3,500). You can reduce API costs by up to 40% with caching, saving an estimated $200–$500 per month for active apps. Exact cost is calculated after auditing your project—reach out to us for an estimate. For a precise quote, contact us.
Our Team’s Experience
We have over 5 years of mobile development experience. We’ve completed 15+ projects integrating AI services, including OpenAI, Google Cloud Speech, and Yandex SpeechKit. Our developers are Apple and Google certified. We guarantee stable operation and transparent collaboration.
Get a consultation on integrating voice synthesis into your app—contact us.







