We frequently face a task in mobile apps: two streamers must broadcast simultaneously, viewers see both, and latency is minimal. Technically, this combines WebRTC and RTMP (broadcasting the result to viewers). On mobile devices with limited resources, merging them is nontrivial—especially when real-time audio and video synchronization is required. An MVP typically takes 5–7 weeks, but a complete solution supporting both platforms and AEC takes up to 12 weeks. Typical client-side mixing latency is 150–300 ms; we reduce it to 50–100 ms by optimizing the audio pipeline. Server-side approaches (MCU/SFU) give 100–150 ms latency but require server infrastructure. Our experience shows that for apps with 500–2000 DAU, client-side mixing saves up to 60% of infrastructure budget.
Co-stream Architecture
The standard scheme:
Streamer A: camera → WebRTC → Signaling Server ← WebRTC ← camera: Streamer B
↓
Mixing Server (SFU/MCU)
↓
RTMP → Twitch/YouTube/Custom
But on mobile, there is an option without MCU—client-side mixing. Streamer A receives streamer B's video via WebRTC, mixes both streams locally via Metal/OpenGL, and sends the mixed stream to RTMP. This is cheaper server-wise but requires a powerful device CPU and becomes unstable with poor network from the second participant.
In production for apps with 1000+ concurrent co-streams—only server-side MCU/SFU (LiveKit, mediasoup, Agora). For low-load MVPs, client-side mixing works.
Comparison of approaches:
| Criteria |
Client-side mixing |
Server-side MCU/SFU |
| Infrastructure cost |
Low (only signaling server) |
High (mixing server) |
| Latency |
Low (local mixing) |
Medium (depends on region) |
| Device requirements |
High (CPU/GPU) |
Low (only WebRTC) |
| Scalability |
Up to 2 participants |
10+ participants |
| Implementation complexity |
Medium |
High |
Additionally, compare mixing methods:
| Method |
Audio latency |
Implementation complexity |
| Client-side (Metal) |
150–200 ms |
Medium |
| Server-side (MCU) |
100–150 ms |
High |
| Server-side (SFU) |
50–100 ms |
High |
WebRTC on iOS and Android
iOS: GoogleWebRTC (CocoaPods) or WebRTC.xcframework from Google. The main object is RTCPeerConnection. Initialization:
let config = RTCConfiguration()
config.iceServers = [RTCIceServer(urlStrings: ["stun:stun.l.google.com:19302"],
username: nil, credential: nil)]
config.sdpSemantics = .unifiedPlan
let constraints = RTCMediaConstraints(
mandatoryConstraints: ["OfferToReceiveVideo": "true",
"OfferToReceiveAudio": "true"],
optionalConstraints: nil
)
let peerConnection = factory.peerConnection(with: config,
constraints: constraints,
delegate: self)
Android: org.webrtc:google-webrtc:1.0.+ or io.getstream:stream-webrtc-android. Logic is similar via PeerConnection.
The signaling server is a WebSocket that exchanges SDP offer/answer and ICE candidates between streamers. We usually write it in Node.js (ws) or use a ready-made one—LiveKit Server, Agora RTM.
Why Client-side Mixing Is Not a Panacea?
The main problems we encounter in real projects:
Audio Latency During Mixing
When streamer A hears streamer B via WebRTC with 150–200 ms latency, and the stream is built from A's local audio, viewers hear desync. Solution: compensate delay via AVAudioPlayerNode.scheduleBuffer with explicit AVAudioTime so that local audio in the final stream is delayed by the same amount as incoming audio from B.
Echo Cancellation
If the streamer is not using headphones, their microphone picks up sound from the speaker (WebRTC audio from the partner). Built-in AEC in WebRTC works only on the RTCPeerConnection audio track. With a custom audio pipeline, you need AVAudioEngine with AVAudioUnitEQ plus custom AEC or speex DSP.
Switching Between Co-stream and Solo
When a partner leaves the co-stream, you must smoothly remove their window from the composition and rearrange the layout without interrupting the RTMP stream. This means the Metal render pass must check for the presence of a second texture and correctly render full-frame mode if the second participant disconnects.
How We Do It: Implementation Experience
Our engineers with years of mobile development experience offer a turnkey co-streaming solution. We design the architecture, choose the optimal stack (client-side or server-side), implement WebRTC integration, video and audio composition, and the signaling server. For example, for a social platform with 10,000 DAU, we implemented a two-user co-stream on iOS with client-side mixing and audio latency compensation—the project took 6 weeks. The cost of such a solution varies based on complexity, which can be significantly cheaper than renting a server-side MCU. Contact us for a consultation to evaluate your project.
More about typical timelines and costing
For an MVP with client-side mixing on a single platform (iOS)—5–7 weeks. A full project with two platforms and server architecture—8–12 weeks. Exact cost is determined after requirement audit.
Process
- Analysis: discuss requirements, load, target audience. Determine client-side or server-side approach.
- Design: develop signaling server scheme, protocols, stream state API.
- Implementation: write code in Swift/Kotlin, integrate WebRTC, set up audio pipeline.
- Testing: check latency, quality under different network conditions, echo. Perform load testing with 1000 virtual users.
- Deployment: set up server part, CI/CD, publish to App Store or Google Play.
What Is Included
- Architecture and API documentation.
- Source code of the mobile app and signaling server.
- CI/CD setup (TestFlight, Firebase App Distribution).
- Integration with the chosen streaming platform.
- 30-day support after delivery.
Timelines
Client-side co-stream (iOS, two participants, Metal composition, basic signaling server): 5–7 weeks. Full implementation with MCU, Android support, AEC, state management: 8–12 weeks. Cost is calculated individually after requirement analysis and architecture selection.
Evaluate your project—contact us for a consultation. Get a consultation for your project—we will assess the requirements and propose the optimal solution.
How to Choose a Camera Approach on Mobile Platforms?
Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.
What are the Differences Between CameraX, Camera2, and AVFoundation?
On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.
CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.
When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.
On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.
A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.
Why is AudioFocus Critical for Android Apps?
Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus
On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.
AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.
On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.
Video: Playback and Streaming
ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.
Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.
AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.
For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).
| Protocol |
Latency |
Application |
| RTMP |
2–5 sec |
Streaming to YouTube/Twitch |
| HLS |
6–30 sec |
VOD, broadcast |
| DASH |
6–30 sec |
VOD with adaptive bitrate |
| WebRTC |
< 500 ms |
Video calls, P2P |
| SRT |
1–4 sec |
Professional streaming |
WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.
RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.
Media Processing: Compression and Transcoding
ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.
On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.
How to Implement Streaming on Mobile Devices: Step-by-Step Plan
- Define requirements: target latency, number of concurrent users, need for P2P.
- Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
- Set up signaling (SIP, WebSocket, MQTT) and TURN server.
- Implement publishing/viewing via native API or cross-platform plugin.
- Test on real devices with different cameras and network conditions.
- Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
- Configuring AVFoundation session on the main thread.
- Missing AudioFocus Loss handling on Android.
- Ignoring
MediaCodec limitations on cheap devices.
- Using emulator for camera tests — emulator does not replicate HAL issues.
- Memory leaks when recreating media players without a pool.
What is Included in the Work
| Deliverable |
Description |
| Requirements analysis |
Stack selection, priorities, test devices |
| Design |
Architecture, data flow diagrams, API selection |
| Implementation |
Code using chosen tools |
| Backend integration |
GraphQL/REST, DRM, WebRTC signaling |
| Testing |
On real devices (at least 5 models) |
| Documentation |
API documentation, build instructions |
| Post-release support |
1 month incident support, team training |
Development Process for Media Functionality
Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).
Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.
Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.
Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.