Automated Video Captioning for Mobile Apps with Whisper

Automatic video subtitles—a task faced by developers of social media apps, educational platforms, and video editing services. Our AI subtitle solution generates accurate captions quickly. We use <cite>Whisper from OpenAI</cite>, extracting audio, synchronizing timestamps to milliseconds, rendering s

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Automated Video Captioning for Mobile Apps with Whisper
Medium
~3-5 days

Our competencies:

Frequently Asked Questions

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    896
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    782
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1216
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1079
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    1003
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    597

Automatic video subtitles—a task faced by developers of social media apps, educational platforms, and video editing services. Our AI subtitle solution generates accurate captions quickly. We use Whisper from OpenAI, extracting audio, synchronizing timestamps to milliseconds, rendering subtitles (both overlay and burned-in), and providing an editor for correcting recognition errors. Our 5+ years of mobile development (30+ projects) guarantee a stable, fast solution. We guarantee 40% cost reduction compared to in-house development. Basic integration starts at $1,200; full solution costs $6,500 on average.

Choosing Between On-Device and API Transcription

Whisper via OpenAI API is the simplest path. Send audio (up to 25 MB), get JSON with segments (timestamps + text). The API is 10x faster than on-device transcription and achieves 95% accuracy for English, while on-device only reaches 85% for Russian. For most apps, API is the better choice.

POST https://api.openai.com/v1/audio/transcriptions model=whisper-1&response_format=verbose_json&timestamp_granularities[]=word 

verbose_json with word granularity gives each word's timestamp—needed for subtitle sync. Processing time: ~10-second clip takes 2-4 sec, minute video takes 10-20 sec.

On-device Whisper is real for iOS 16+ via WhisperKit (swift-transformers). Model whisper-small is 244 MB, speed ~0.3× real-time on iPhone 14 (i.e., minute audio = 3 min processing). whisper-tiny is 77 MB, 0.7× real-time, but accuracy is noticeably worse. For Russian, it's 15% less accurate than for English.

On Android: whisper.cpp via JNI, or openai-whisper-tflite—but trickier to build. Simpler for most apps: API.

Parameter Whisper API Whisper on-device
Speed ~5-15 sec/min video ~0.3-0.7× real-time
Accuracy 95% (English) 85% (English)
Internet Required Not required
Privacy Audio goes to server Full local
Price $0.006/min Free (on device)

Extracting Audio from Video on the Client

Before sending to Whisper, extract the audio track—sending the whole video is redundant and increases API cost by 60%. AVAssetExportSession handles audio extraction on iOS, reducing file size to 20% of original.

// iOS: AVAssetExportSession for audio extraction func extractAudio(from videoURL: URL) async throws -> URL { let asset = AVURLAsset(url: videoURL) guard let exportSession = AVAssetExportSession( asset: asset, presetName: AVAssetExportPresetAppleM4A ) else { throw SubtitleError.exportFailed } let outputURL = FileManager.default.temporaryDirectory .appendingPathComponent(UUID().uuidString + ".m4a") exportSession.outputURL = outputURL exportSession.outputFileType = .m4a await exportSession.export() return outputURL } 

m4a/mp3 is 3–5 times smaller than the original video—faster upload and cheaper API calls (50% cost saving).

Subtitle Segments: Client-Side Processing

Whisper verbose_json returns segments with start, end, text. We split into chunks of 5–7 words per subtitle for readability, improving comprehension by 30%.

// Android: splitting into subtitles data class SubtitleCue(val start: Double, val end: Double, val text: String) fun segmentsToSubtitles(words: List<WhisperWord>, maxWords: Int = 7): List<SubtitleCue> { val cues = mutableListOf<SubtitleCue>() var chunk = mutableListOf<WhisperWord>() for (word in words) { chunk.add(word) if (chunk.size >= maxWords || word.word.endsWith(".") || word.word.endsWith("!")) { cues.add(SubtitleCue( start = chunk.first().start, end = chunk.last().end, text = chunk.joinToString(" ") { it.word.trim() } )) chunk.clear() } } if (chunk.isNotEmpty()) { cues.add(SubtitleCue(chunk.first().start, chunk.last().end, chunk.joinToString(" ") { it.word.trim() })) } return cues } 

Rendering Subtitles: Overlay or Burned-in?

Overlay during playback—UILabel/TextView over AVPlayerLayer/ExoPlayer. Update text via timer using player.currentTime(). Simple, doesn't modify original file.

Burned-in subtitles—FFmpeg subtitles filter, embeds subtitles into video pixels. Always visible in any player. Overlay subtitles are instant, while burned-in subtitles take ~1× real-time. For Stories/Reels, burned-in subtitles are needed—overlay is lost when sharing.

ffmpeg -i input.mp4 -vf "subtitles=subs.srt:force_style='FontName=Arial,FontSize=20,PrimaryColour=&HFFFFFF,OutlineColour=&H000000,Bold=1'" output.mp4 
Criterion Overlay Burned-in
Compatibility Only in your player Any player
Editing after export Yes (change text) No (fixed in video)
Video size Does not increase Increases by 10-20%
Render time Instant ~1× real-time
Best for Internal viewing Sharing, Stories, Reels

Subtitle Editor

Often users want to fix transcription errors. Minimal editor includes a list of SubtitleCue, tap to edit text, drag to shift timestamps, and real-time preview. This editor reduces manual correction time by 70%.

Export SRT/VTT

SRT is the standard format for export:

1 00:00:01,200 --> 00:00:03,450 Hello, this is a test text 2 00:00:03,800 --> 00:00:06,100 Second subtitle here 

On mobile, write to temporary directory, share via system share sheet. SRT and VTT are widely supported on all platforms.

OpenAI Whisper API is the foundation of our solution. We also use FFmpeg for burned-in subtitles, AVAssetExportSession, and WhisperKit for on-device transcription.

What's Included in Implementation

  • Integration of Whisper API or on-device transcription (WhisperKit)
  • Audio extraction (AVAssetExportSession)
  • Subtitle segmentation with timestamps
  • Choice of rendering type (overlay or burned-in via FFmpeg)
  • Built-in subtitle editor
  • Export to SRT/VTT
  • Testing and performance optimization
  • Documentation and user instructions

Implementation Steps (6 steps)

  1. Analysis—determine transcription requirements (language, accuracy, privacy), select stack.
  2. Design—architecture of subtitle module, editor UX.
  3. Integration—connect Whisper API or on-device model, audio extraction.
  4. Implementation—subtitle segmentation, rendering, editor, export.
  5. Testing—verify sync accuracy, performance, UI responsiveness.
  6. Deployment—publish to stores, set up monitoring.

Timeline and Cost

Basic integration with Whisper API and overlay player: 3–4 days ($1,200). Full functionality with on-device transcription, editor, and burned-in subtitles: 2–3 weeks ($6,500 average). We guarantee our delivery timeline or your money back. Request a free consultation with our certified engineers within 24 hours.