Implementing automatic video moderation in mobile apps
A user uploads a video — and you have seconds to decide whether to show it to others. Manual review doesn't scale: moderators get tired, miss violations, and a 10-minute queue kills retention. Our team of mobile engineers with experience in AI moderation has delivered 40+ projects for FinTech, EdTech, and Social platforms. We specialize in integrating on-device models and cloud APIs for automatic video content classification. For effective video moderation in mobile apps, combining on-device AI filtering with cloud-based analysis is key. Below, we break down how to avoid common mistakes and build a reliable system.
Where problems most often arise
Real-time moderation vs. post-processing
The most common architectural mistake is trying to run a model frame-by-frame on the client. CoreML on an iPhone 14 Pro can handle MobileNet v3 at 30 fps for short clips, but it kills battery and overheats the device. On Android, the situation is similar with MediaPipe: processing every frame in ImageAnalysis.Analyzer at 1080p leads to ImageProxy backlogs and crashes with java.lang.IllegalStateException: Image is already closed.
The right approach for video is not frame-by-frame analysis but selective sampling: every N frames or key scenes using AVAssetImageGenerator (iOS) / MediaMetadataRetriever.getFrameAtTime() (Android). For most moderation tasks, 1 frame per second is sufficient.
Server-side moderation via Video Intelligence API
For UGC video apps, we design the following scheme: the client uploads video to storage (S3/GCS), triggers a Cloud Function that calls Google Cloud Video Intelligence with EXPLICIT_CONTENT and OBJECT_TRACKING features. The response is JSON with timestamps and confidence scores per segment.
// Android: start upload and pass URI to backend
val uploadRef = storageRef.child("uploads/${UUID.randomUUID()}.mp4")
uploadRef.putFile(localUri)
.addOnSuccessListener { taskSnapshot ->
taskSnapshot.storage.downloadUrl.addOnSuccessListener { downloadUri ->
moderationApi.submitVideo(downloadUri.toString(), onComplete = { result ->
when (result.verdict) {
ModerationVerdict.SAFE -> publishVideo()
ModerationVerdict.UNSAFE -> rejectWithReason(result.reason)
ModerationVerdict.REVIEW -> sendToHumanReview()
}
})
}
}
AWS Rekognition Video is an alternative with a similar API: StartContentModeration + polling via GetContentModeration. For synchronous cases (short reels up to 30 sec), Rekognition Image applied to extracted frames works — response in 200–400 ms. Cost for processing a minute of video via Google Video Intelligence starts from $0.15 per minute.
On-device pre-filtering
Before sending to the server, it makes sense to run the first and last frames of the video through a local CoreML/TFLite model. This catches obvious NSFW on the client, saving traffic. NudeNet Lite in TFLite format is about 14 MB and delivers ~92% accuracy on NSFW benchmarks. False positives on medical content are a separate issue and require whitelist logic at the app category level.
How on-device pre-filtering works
On-device pre-filtering is applied to extracted key frames (first and last) via CoreML (iOS) or TFLite (Android). The model is NudeNet Lite, 14 MB, 92% accuracy on NSFW datasets. False positives (medical content) are handled by whitelist logic. On-device analysis is 3x faster than a server roundtrip for obvious violations and saves up to 30% of traffic. We use Swift Combine on iOS for reactive upload progress, and Kotlin Coroutines with Flow on Android for similar functionality.
Why server-side moderation is preferable for UGC
Server-side moderation via cloud APIs (Google Video Intelligence, AWS Rekognition) wins in accuracy and scalability. On-device models are limited by compute resources and support only basic classes (NSFW, violence). For complex moderation (weapon detection, contextual violations), a server-side model with larger context is required. In practice, we combine: fast on-device filter + deep cloud analysis for borderline cases.
Comparison of solutions
| Parameter | On-device (CoreML/TFLite) | Cloud API (Google/AWS) | Live (WebRTC + MobileViT) |
|---|---|---|---|
| Latency | 100–300 ms | 500–2000 ms | 2–4 sec per segment |
| Accuracy | 92% (NSFW) | 95–98% (all classes) | 90% (compromise) |
| Traffic | None | Upload video | Continuous stream |
| Cost | Model development from $5,000 | From $0.15/min | Higher due to real-time |
| Use case | Pre-filtering | Primary moderation | Streams |
Typical scenarios and recommended solutions
| Scenario | On-device filter | Cloud analysis | Live stack |
|---|---|---|---|
| UGC feed | Yes (first/last frame) | Yes (Video Intelligence) | No |
| Stories | Yes (every frame on capture) | Yes (after upload) | No |
| Live streaming | No | No | Yes (WebRTC + HLS) |
| Short reels (<30 sec) | Yes (all frames) | Yes (Rekognition Image) | Optional |
How we build the solution
The stack depends on latency requirements and budget. For startups with low traffic — Google Video Intelligence API: pay $0.15 per minute, no need to spin up infrastructure. For high-load platforms — custom inference service based on CLIP or a custom ONNX model behind a reverse proxy with caching of already-checked video hashes (perceptual hashing via pHash prevents re-moderation of the same clip).
On the client side (iOS/Android/Flutter), we implement:
- upload progress bar with
URLSession.uploadTask/okhttp3.MultipartBody - pending state for videos in the feed ("under review")
- push notification of result via FCM/APNs
A separate case is live streaming. Video Intelligence API is not suitable here due to latency. We use streaming via WebRTC + server-side analysis of HLS segments every 2–4 seconds with a speed-optimized model (MobileViT-S in TorchScript).
What is included in the work (deliverables)?
- Architectural decision records (ADR)
- SDK integration for upload and moderation
- UI states (pending, approved, rejected)
- Testing on 100+ edge cases
- Team lead training on administration
- 6-month code warranty
Process
- Requirements audit: content type (UGC, Stories, live), acceptable publication delay, compliance requirements (GDPR, COPPA).
- Stack selection: on-device pre-filter + cloud moderation vs. fully server-side.
- Development: integrate upload SDK, webhook/polling for results, status UI.
- Testing on edge-case dataset: multilingual subtitles in frame, medical content, animation.
Time estimates
Integration with Google Video Intelligence or AWS Rekognition Video — 3–5 days. Adding on-device pre-filter on CoreML/TFLite — another 2–3 days. Full solution with live streaming and human review system — 3–4 weeks.
For more information, we can prepare a proposal tailored to your stack and load. A consultation can help discuss details.







