Group Video Calls: SFU, Simulcast & Thermal Adaptation
Picture this: your app has 8 participants, after 10 minutes the iPhone heats to 42°C, the call drops. The issue is architecture: without an SFU (Selective Forwarding Unit), each mobile client sends video to all others — in a group of 8, each handles 7 incoming streams and 1 outgoing. The CPU and network load scales quadratically. The proper solution is an SFU with simulcast. We design the architecture for your scenario: 2 to 20 participants, with adaptation to device thermal state.
Why SFU over MCU?
MCU (Multipoint Control Unit) mixes all streams into one on the server and sends each participant a single video. Client load is minimal, but mixing requires powerful CPU and introduces 100–300 ms latency. Suitable for webinars where most watch one speaker. SFU (Selective Forwarding Unit) only routes streams without decoding. Each client receives N streams (one per participant) and chooses which to decode and display. Higher client load but lower latency and greater flexibility.
For mobile apps with conferences up to 20 participants, SFU is optimal. Ready solutions: Livekit (open-source, self-hosted), Mediasoup, Jitsi Videobridge, or managed services — Daily.co, 100ms, Twilio Video Rooms. Livekit is a solid choice for self-hosted: Go server, WebRTC SFU, support for simulcast and dynacast, native SDKs for iOS (LiveKit-iOS), Android (LiveKit-Android), Flutter, and React Native. MIT license.
How Simulcast Reduces CPU Load
Without simulcast, on a 10+ person conference a mobile device sends one 720p stream to all participants, even those with small tiles. With simulcast, the client sends three qualities simultaneously (e.g., 180p/360p/720p), and the SFU sends each recipient the quality matching their tile size. Resource savings are significant: CPU load can drop up to 40% with many participants.
On a recent telemedicine project, we reduced CPU load by 35% on iOS devices by implementing simulcast and thermal state adaptation, enabling stable 12-person conferences without overheating.
On iOS, simulcast is configured via RTCRtpEncodingParameters with three layers:
let encodings = [
RTCRtpEncodingParameters(rid: "q", scaleResolutionDownBy: 4, maxBitrateBps: 150_000),
RTCRtpEncodingParameters(rid: "h", scaleResolutionDownBy: 2, maxBitrateBps: 500_000),
RTCRtpEncodingParameters(rid: "f", scaleResolutionDownBy: 1, maxBitrateBps: 1_200_000),
]
On Android — similarly via RtpParameters.Encoding. This reduces network and CPU load on receivers with many participants.
| Layer |
Resolution |
Bitrate (bps) |
Use Case |
| q (quarter) |
180p |
150,000 |
Thumbnails, 9+ participants |
| h (half) |
360p |
500,000 |
Medium tile, 4–8 participants |
| f (full) |
720p |
1,200,000 |
Main speaker, 1–2 participants |
How to Display Participants: Grid and Dominant Speaker
Grid for 2–4 participants is a static layout. For 5–16 participants, use a dynamic grid that changes on join/leave. Rule: do not recreate RTCVideoRenderer on every grid update — only reassign the track to the existing renderer. Recreating causes flickering and re-render.
Dominant speaker detection — identify who is speaking and show them larger. Livekit and 100ms provide this out of the box via onActiveSpeakersChanged events. In raw WebRTC, analyze audioLevel from RTCPeerConnection.getStats().
On iOS, render video through RTCMTLVideoView (Metal) — mandatory for conferences. The old RTCEAGLVideoView (OpenGL ES) does not support multiple instances with good performance on A15+. With 6 participants, RTCEAGLVideoView gives 40 FPS, RTCMTLVideoView a stable 60 FPS.
Battery and Thermal State
Group conferencing is the most battery-intensive mode. Encoding 720p at 30 FPS consumes ~15–20% battery per hour on an iPhone 14. With 4+ participants, decoding multiple streams adds more.
We react to ProcessInfo.thermalState (iOS) — at .serious or .critical, we lower resolution to 360p and reduce FPS to 15. On Android, PowerManager.getThermalHeadroom() (Android 11+). This is not degradation, but adaptation: a stable 360p conference is better than overheating and forced CPU throttling.
When the app is backgrounded on iOS, we stop camera capture (AVCaptureSession.stopRunning()) while audio continues via AVAudioSession. On Android, a ForegroundService with android:foregroundServiceType="camera|microphone" holds permissions. This saves battery and reduces heat.
Common Mistakes
- Single
AVCaptureSession for the whole app — do not recreate per call. Initialization takes 200–500 ms; recreating every call is noticeable.
- Not handling
AVAudioSession.routeChangeNotification — connecting AirPods during a conference without handling this notification routes audio to the earpiece.
- Not releasing
RTCVideoTrack when a participant leaves — memory leak accumulates over long sessions and large rooms.
Process and Timelines
Requirements audit → SFU selection (self-hosted or managed) → SDK integration → UI (grid, dominant speaker, controls) → simulcast → thermal state adaptation → load testing. After delivery, we provide SDK documentation, management console access, and brief team training. Our team has 7+ years of mobile development experience and over 15 commercial WebRTC projects.
| Stage |
Timeline (weeks) |
Includes |
| Audit and architecture selection |
1–2 |
Technical analysis, SFU recommendation |
| SDK integration and basic UI |
2–3 |
Livekit/100ms integration, grid, controls |
| Simulcast and adaptive quality |
1–2 |
Encoding configuration, thermal state |
| Load testing and battery profiling |
1 |
Profiling, optimization |
Conference up to 8 participants via managed SFU (100ms, Daily) with ready SDK — 2–4 weeks. Self-hosted Livekit with custom UI and simulcast — 4–8 weeks. Cost is determined after requirements analysis. Contact us — we'll assess your project and propose a turnkey architecture.
What's Included
- Technical audit and architecture recommendations
- Selection and integration of SFU (self-hosted or managed)
- UI development: grid, dominant speaker, controls
- Simulcast and adaptive quality configuration
- Battery and thermal state optimization
- Load testing for up to 20 participants
- SDK integration documentation
- Management console access
- Brief team training
Step-by-Step Simulcast Setup on iOS
- Create
RTCRtpEncodingParameters objects for each layer with rid, scaleResolutionDownBy, maxBitrateBps.
- Apply these encodings to the video track via
RTCRtpSender.setParameters().
- Enable simulcast on the SFU side (in Livekit, the
SimulcastConfig flag).
- Verify that receivers correctly switch layers when tile size changes.
- Profile CPU and network load with 10+ participants.
According to WebRTC documentation, proper simulcast implementation can reduce bandwidth usage by up to 40%.
Get a consultation on integration — we'll help implement simulcast in your project.
How to Choose a Camera Approach on Mobile Platforms?
Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.
What are the Differences Between CameraX, Camera2, and AVFoundation?
On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.
CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.
When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.
On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.
A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.
Why is AudioFocus Critical for Android Apps?
Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus
On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.
AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.
On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.
Video: Playback and Streaming
ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.
Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.
AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.
For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).
| Protocol |
Latency |
Application |
| RTMP |
2–5 sec |
Streaming to YouTube/Twitch |
| HLS |
6–30 sec |
VOD, broadcast |
| DASH |
6–30 sec |
VOD with adaptive bitrate |
| WebRTC |
< 500 ms |
Video calls, P2P |
| SRT |
1–4 sec |
Professional streaming |
WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.
RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.
Media Processing: Compression and Transcoding
ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.
On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.
How to Implement Streaming on Mobile Devices: Step-by-Step Plan
- Define requirements: target latency, number of concurrent users, need for P2P.
- Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
- Set up signaling (SIP, WebSocket, MQTT) and TURN server.
- Implement publishing/viewing via native API or cross-platform plugin.
- Test on real devices with different cameras and network conditions.
- Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
- Configuring AVFoundation session on the main thread.
- Missing AudioFocus Loss handling on Android.
- Ignoring
MediaCodec limitations on cheap devices.
- Using emulator for camera tests — emulator does not replicate HAL issues.
- Memory leaks when recreating media players without a pool.
What is Included in the Work
| Deliverable |
Description |
| Requirements analysis |
Stack selection, priorities, test devices |
| Design |
Architecture, data flow diagrams, API selection |
| Implementation |
Code using chosen tools |
| Backend integration |
GraphQL/REST, DRM, WebRTC signaling |
| Testing |
On real devices (at least 5 models) |
| Documentation |
API documentation, build instructions |
| Post-release support |
1 month incident support, team training |
Development Process for Media Functionality
Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).
Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.
Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.
Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.