We frequently face a task in mobile apps: two streamers must broadcast simultaneously, viewers see both, and latency is minimal. Technically, this combines WebRTC and RTMP (broadcasting the result to viewers). On mobile devices with limited resources, merging them is nontrivial—especially when real-time audio and video synchronization is required. An MVP typically takes 5–7 weeks, but a complete solution supporting both platforms and AEC takes up to 12 weeks. Typical client-side mixing latency is 150–300 ms; we reduce it to 50–100 ms by optimizing the audio pipeline. Server-side approaches (MCU/SFU) give 100–150 ms latency but require server infrastructure. Our experience shows that for apps with 500–2000 DAU, client-side mixing saves up to 60% of infrastructure budget.
Co-stream Architecture
The standard scheme:
Streamer A: camera → WebRTC → Signaling Server ← WebRTC ← camera: Streamer B ↓ Mixing Server (SFU/MCU) ↓ RTMP → Twitch/YouTube/Custom But on mobile, there is an option without MCU—client-side mixing. Streamer A receives streamer B's video via WebRTC, mixes both streams locally via Metal/OpenGL, and sends the mixed stream to RTMP. This is cheaper server-wise but requires a powerful device CPU and becomes unstable with poor network from the second participant.
In production for apps with 1000+ concurrent co-streams—only server-side MCU/SFU (LiveKit, mediasoup, Agora). For low-load MVPs, client-side mixing works.
Comparison of approaches:
| Criteria | Client-side mixing | Server-side MCU/SFU |
|---|---|---|
| Infrastructure cost | Low (only signaling server) | High (mixing server) |
| Latency | Low (local mixing) | Medium (depends on region) |
| Device requirements | High (CPU/GPU) | Low (only WebRTC) |
| Scalability | Up to 2 participants | 10+ participants |
| Implementation complexity | Medium | High |
Additionally, compare mixing methods:
| Method | Audio latency | Implementation complexity |
|---|---|---|
| Client-side (Metal) | 150–200 ms | Medium |
| Server-side (MCU) | 100–150 ms | High |
| Server-side (SFU) | 50–100 ms | High |
WebRTC on iOS and Android
iOS: GoogleWebRTC (CocoaPods) or WebRTC.xcframework from Google. The main object is RTCPeerConnection. Initialization:
let config = RTCConfiguration() config.iceServers = [RTCIceServer(urlStrings: ["stun:stun.l.google.com:19302"], username: nil, credential: nil)] config.sdpSemantics = .unifiedPlan let constraints = RTCMediaConstraints( mandatoryConstraints: ["OfferToReceiveVideo": "true", "OfferToReceiveAudio": "true"], optionalConstraints: nil ) let peerConnection = factory.peerConnection(with: config, constraints: constraints, delegate: self) Android: org.webrtc:google-webrtc:1.0.+ or io.getstream:stream-webrtc-android. Logic is similar via PeerConnection.
The signaling server is a WebSocket that exchanges SDP offer/answer and ICE candidates between streamers. We usually write it in Node.js (ws) or use a ready-made one—LiveKit Server, Agora RTM.
Why Client-side Mixing Is Not a Panacea?
The main problems we encounter in real projects:
Audio Latency During Mixing
When streamer A hears streamer B via WebRTC with 150–200 ms latency, and the stream is built from A's local audio, viewers hear desync. Solution: compensate delay via AVAudioPlayerNode.scheduleBuffer with explicit AVAudioTime so that local audio in the final stream is delayed by the same amount as incoming audio from B.
Echo Cancellation
If the streamer is not using headphones, their microphone picks up sound from the speaker (WebRTC audio from the partner). Built-in AEC in WebRTC works only on the RTCPeerConnection audio track. With a custom audio pipeline, you need AVAudioEngine with AVAudioUnitEQ plus custom AEC or speex DSP.
Switching Between Co-stream and Solo
When a partner leaves the co-stream, you must smoothly remove their window from the composition and rearrange the layout without interrupting the RTMP stream. This means the Metal render pass must check for the presence of a second texture and correctly render full-frame mode if the second participant disconnects.
How We Do It: Implementation Experience
Our engineers with years of mobile development experience offer a turnkey co-streaming solution. We design the architecture, choose the optimal stack (client-side or server-side), implement WebRTC integration, video and audio composition, and the signaling server. For example, for a social platform with 10,000 DAU, we implemented a two-user co-stream on iOS with client-side mixing and audio latency compensation—the project took 6 weeks. The cost of such a solution varies based on complexity, which can be significantly cheaper than renting a server-side MCU. Contact us for a consultation to evaluate your project.
More about typical timelines and costing
For an MVP with client-side mixing on a single platform (iOS)—5–7 weeks. A full project with two platforms and server architecture—8–12 weeks. Exact cost is determined after requirement audit.Process
- Analysis: discuss requirements, load, target audience. Determine client-side or server-side approach.
- Design: develop signaling server scheme, protocols, stream state API.
- Implementation: write code in Swift/Kotlin, integrate WebRTC, set up audio pipeline.
- Testing: check latency, quality under different network conditions, echo. Perform load testing with 1000 virtual users.
- Deployment: set up server part, CI/CD, publish to App Store or Google Play.
What Is Included
- Architecture and API documentation.
- Source code of the mobile app and signaling server.
- CI/CD setup (TestFlight, Firebase App Distribution).
- Integration with the chosen streaming platform.
- 30-day support after delivery.
Timelines
Client-side co-stream (iOS, two participants, Metal composition, basic signaling server): 5–7 weeks. Full implementation with MCU, Android support, AEC, state management: 8–12 weeks. Cost is calculated individually after requirement analysis and architecture selection.
Evaluate your project—contact us for a consultation. Get a consultation for your project—we will assess the requirements and propose the optimal solution.







