We develop video calls in mobile apps. This is not just a WebRTC connection: camera and microphone management, system-level lifecycle, handling interruptions (incoming call, notification, screen lock), correct background and low-battery operation. Each of these details is a separate task with non-trivial solutions on iOS and Android. Our experience in this area exceeds 7 years, and we have delivered over 30 video call projects for clients from various industries.
Which Transport to Choose: WebRTC or a Ready SDK?
Pure WebRTC involves RTCPeerConnection, RTCSessionDescription, ICE/STUN/TURN servers, getUserMedia, signaling via WebSocket. A full implementation from scratch takes 3–6 weeks just for the transport layer, without the UI.
Most projects choose an SDK on top of WebRTC: Twilio Video, Daily.co, 100ms, Livekit, or Agora. They handle ICE negotiation, codec negotiation (VP8/VP9/H.264/AV1), adaptive bitrate, and provide ready native wrappers for iOS and Android.
If customization requirements are low and time-to-market is critical, we choose an SDK. If you need full control over codec, encryption, or minimal binary size, we opt for pure WebRTC via GoogleWebRTC (iOS) or org.webrtc (Android).
| Criterion |
WebRTC |
SDK (Twilio/Agora) |
| Development time |
3–6 weeks |
1–2 weeks |
| Control |
Full |
Limited |
| Binary size |
Minimal |
5–10 MB larger |
| Codec support |
Any |
Limited by SDK |
How to Implement Video Calls on iOS?
Video from the camera is captured via AVCaptureSession. The correct initialization:
let session = AVCaptureSession()
session.sessionPreset = .hd1280x720
let camera = AVCaptureDevice.default(.builtInWideAngleCamera, for: .video, position: .front)
let input = try AVCaptureDeviceInput(device: camera!)
let output = AVCaptureVideoDataOutput()
output.setSampleBufferDelegate(self, queue: DispatchQueue(label: "videoQueue"))
AVCaptureVideoDataOutput with a delegate provides raw CMSampleBuffer — we pass them to WebRTC via RTCVideoSource. Important: AVCaptureSession.startRunning() must be called strictly on a background thread, not the main thread — otherwise the UI freezes for 200–400 ms.
Switching between front and back camera: AVCaptureSession.beginConfiguration() + remove old input + add new + commitConfiguration(). Without begin/commit around these operations, video flickers during switching.
CallKit is mandatory for iOS apps with video calls. Without it, an incoming call appears as a push notification that can be missed. With CallKit, you get a full-screen system call screen, Bluetooth/AirPods integration, and correct audio interruption handling. We implement it via CXProvider and CXCallController. VoIP push via PKPushRegistry is the only way to wake the app for an incoming call.
How to Implement Video Calls on Android?
On Android, the Camera2 API (CameraManager.openCamera()) gives access to ImageReader with YUV_420_888 format — the standard format for WebRTC on Android. For most tasks, CameraDevice.StateCallback + CaptureRequest.Builder suffice.
ConnectionService is the Android equivalent of CallKit. We register a PhoneAccount and manage call state via TelecomManager. Without it on Android 10+, the app cannot stay in the foreground during a call without a persistent notification.
To keep the connection when the app is minimized, we use a ForegroundService with type mediaProjection or phoneCall. On Android 14, this requires explicit declaration in the manifest: android:foregroundServiceType="camera|microphone".
Camera and Microphone Management
Muting the microphone is not about disabling the device but replacing the audio track with silence. In WebRTC: RTCAudioTrack.isEnabled = false. Completely disabling the microphone at the AVAudioSession level breaks echo cancellation.
Flipping the camera during a call: call RTCCameraVideoCapturer.stopCapture(), change the device, and startCapture(with:fps:). Without a full stop, some Huawei and Xiaomi devices crash due to a race condition in Camera2.
Screen rotation: RTCVideoTrack itself does not account for orientation. You must pass RTCVideoRotation in the RTCVideoFrame during camera capture; otherwise, the remote party sees video rotated by 90° on devices with portrait-lock.
Connection Quality and Adaptation
Adaptive bitrate is a built-in feature of WebRTC (GCC algorithm). But you must explicitly set the range: RTCRtpEncodingParameters with minBitrateBps: 100_000 and maxBitrateBps: 1_500_000. Without an upper limit, on Wi-Fi WebRTC will try to use 4–8 Mbps — uncomfortable for other network users.
Connection quality indicator via RTCPeerConnection.getStats() — parse RTCInboundRtpStreamStats for framesPerSecond and packetsLost. Loss > 5% — show a warning to the user.
What Our Work Includes
- Requirements audit and optimal transport selection (WebRTC or SDK)
- Signaling development (WebSocket or third-party server)
- Camera and microphone integration with AVCaptureSession/Camera2
- CallKit/ConnectionService integration for system calls
- UI development: local video preview, remote video, control buttons
- Adaptive bitrate and connection quality monitoring
- Testing on real devices (iOS/Android)
- Integration documentation and source code handover
- Post-release support (2 weeks)
Process
- Requirements audit → 2. Transport selection → 3. Signaling implementation → 4. Camera/microphone integration → 5. CallKit/ConnectionService → 6. UI and control buttons → 7. Testing on real devices → 8. App store deployment.
We evaluate your project for free. Contact us, and we will propose the optimal solution considering your timeline and budget. A basic 1-on-1 video call using an SDK (Twilio/Agora/100ms) with camera control, mute, and flip takes 1–2 weeks. Pure WebRTC with signaling, CallKit, and background mode takes 3–5 weeks. The cost is determined after requirements analysis.
How to Choose a Camera Approach on Mobile Platforms?
Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.
What are the Differences Between CameraX, Camera2, and AVFoundation?
On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.
CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.
When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.
On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.
A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.
Why is AudioFocus Critical for Android Apps?
Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus
On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.
AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.
On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.
Video: Playback and Streaming
ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.
Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.
AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.
For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).
| Protocol |
Latency |
Application |
| RTMP |
2–5 sec |
Streaming to YouTube/Twitch |
| HLS |
6–30 sec |
VOD, broadcast |
| DASH |
6–30 sec |
VOD with adaptive bitrate |
| WebRTC |
< 500 ms |
Video calls, P2P |
| SRT |
1–4 sec |
Professional streaming |
WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.
RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.
Media Processing: Compression and Transcoding
ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.
On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.
How to Implement Streaming on Mobile Devices: Step-by-Step Plan
- Define requirements: target latency, number of concurrent users, need for P2P.
- Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
- Set up signaling (SIP, WebSocket, MQTT) and TURN server.
- Implement publishing/viewing via native API or cross-platform plugin.
- Test on real devices with different cameras and network conditions.
- Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
- Configuring AVFoundation session on the main thread.
- Missing AudioFocus Loss handling on Android.
- Ignoring
MediaCodec limitations on cheap devices.
- Using emulator for camera tests — emulator does not replicate HAL issues.
- Memory leaks when recreating media players without a pool.
What is Included in the Work
| Deliverable |
Description |
| Requirements analysis |
Stack selection, priorities, test devices |
| Design |
Architecture, data flow diagrams, API selection |
| Implementation |
Code using chosen tools |
| Backend integration |
GraphQL/REST, DRM, WebRTC signaling |
| Testing |
On real devices (at least 5 models) |
| Documentation |
API documentation, build instructions |
| Post-release support |
1 month incident support, team training |
Development Process for Media Functionality
Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).
Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.
Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.
Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.