Co-streaming with WebRTC and RTMP on Mobile: Architecture and Implementation

TRUETECH is engaged in the development, support and maintenance of iOS, Android, PWA mobile applications. We have extensive experience and expertise in publishing mobile applications in popular markets like Google Play, App Store, Amazon, AppGallery and others.

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Co-streaming with WebRTC and RTMP on Mobile: Architecture and Implementation
Complex
from 2 weeks to 3 months
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    858
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    744
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1160
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1034
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    968
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    562

We frequently face a task in mobile apps: two streamers must broadcast simultaneously, viewers see both, and latency is minimal. Technically, this combines WebRTC and RTMP (broadcasting the result to viewers). On mobile devices with limited resources, merging them is nontrivial—especially when real-time audio and video synchronization is required. An MVP typically takes 5–7 weeks, but a complete solution supporting both platforms and AEC takes up to 12 weeks. Typical client-side mixing latency is 150–300 ms; we reduce it to 50–100 ms by optimizing the audio pipeline. Server-side approaches (MCU/SFU) give 100–150 ms latency but require server infrastructure. Our experience shows that for apps with 500–2000 DAU, client-side mixing saves up to 60% of infrastructure budget.

Co-stream Architecture

The standard scheme:

Streamer A: camera → WebRTC → Signaling Server ← WebRTC ← camera: Streamer B
                                     ↓
                            Mixing Server (SFU/MCU)
                                     ↓
                          RTMP → Twitch/YouTube/Custom

But on mobile, there is an option without MCU—client-side mixing. Streamer A receives streamer B's video via WebRTC, mixes both streams locally via Metal/OpenGL, and sends the mixed stream to RTMP. This is cheaper server-wise but requires a powerful device CPU and becomes unstable with poor network from the second participant.

In production for apps with 1000+ concurrent co-streams—only server-side MCU/SFU (LiveKit, mediasoup, Agora). For low-load MVPs, client-side mixing works.

Comparison of approaches:

Criteria Client-side mixing Server-side MCU/SFU
Infrastructure cost Low (only signaling server) High (mixing server)
Latency Low (local mixing) Medium (depends on region)
Device requirements High (CPU/GPU) Low (only WebRTC)
Scalability Up to 2 participants 10+ participants
Implementation complexity Medium High

Additionally, compare mixing methods:

Method Audio latency Implementation complexity
Client-side (Metal) 150–200 ms Medium
Server-side (MCU) 100–150 ms High
Server-side (SFU) 50–100 ms High

WebRTC on iOS and Android

iOS: GoogleWebRTC (CocoaPods) or WebRTC.xcframework from Google. The main object is RTCPeerConnection. Initialization:

let config = RTCConfiguration()
config.iceServers = [RTCIceServer(urlStrings: ["stun:stun.l.google.com:19302"],
                                  username: nil, credential: nil)]
config.sdpSemantics = .unifiedPlan

let constraints = RTCMediaConstraints(
    mandatoryConstraints: ["OfferToReceiveVideo": "true",
                           "OfferToReceiveAudio": "true"],
    optionalConstraints: nil
)
let peerConnection = factory.peerConnection(with: config,
                                            constraints: constraints,
                                            delegate: self)

Android: org.webrtc:google-webrtc:1.0.+ or io.getstream:stream-webrtc-android. Logic is similar via PeerConnection.

The signaling server is a WebSocket that exchanges SDP offer/answer and ICE candidates between streamers. We usually write it in Node.js (ws) or use a ready-made one—LiveKit Server, Agora RTM.

Why Client-side Mixing Is Not a Panacea?

The main problems we encounter in real projects:

Audio Latency During Mixing

When streamer A hears streamer B via WebRTC with 150–200 ms latency, and the stream is built from A's local audio, viewers hear desync. Solution: compensate delay via AVAudioPlayerNode.scheduleBuffer with explicit AVAudioTime so that local audio in the final stream is delayed by the same amount as incoming audio from B.

Echo Cancellation

If the streamer is not using headphones, their microphone picks up sound from the speaker (WebRTC audio from the partner). Built-in AEC in WebRTC works only on the RTCPeerConnection audio track. With a custom audio pipeline, you need AVAudioEngine with AVAudioUnitEQ plus custom AEC or speex DSP.

Switching Between Co-stream and Solo

When a partner leaves the co-stream, you must smoothly remove their window from the composition and rearrange the layout without interrupting the RTMP stream. This means the Metal render pass must check for the presence of a second texture and correctly render full-frame mode if the second participant disconnects.

How We Do It: Implementation Experience

Our engineers with years of mobile development experience offer a turnkey co-streaming solution. We design the architecture, choose the optimal stack (client-side or server-side), implement WebRTC integration, video and audio composition, and the signaling server. For example, for a social platform with 10,000 DAU, we implemented a two-user co-stream on iOS with client-side mixing and audio latency compensation—the project took 6 weeks. The cost of such a solution varies based on complexity, which can be significantly cheaper than renting a server-side MCU. Contact us for a consultation to evaluate your project.

More about typical timelines and costing For an MVP with client-side mixing on a single platform (iOS)—5–7 weeks. A full project with two platforms and server architecture—8–12 weeks. Exact cost is determined after requirement audit.

Process

  1. Analysis: discuss requirements, load, target audience. Determine client-side or server-side approach.
  2. Design: develop signaling server scheme, protocols, stream state API.
  3. Implementation: write code in Swift/Kotlin, integrate WebRTC, set up audio pipeline.
  4. Testing: check latency, quality under different network conditions, echo. Perform load testing with 1000 virtual users.
  5. Deployment: set up server part, CI/CD, publish to App Store or Google Play.

What Is Included

  • Architecture and API documentation.
  • Source code of the mobile app and signaling server.
  • CI/CD setup (TestFlight, Firebase App Distribution).
  • Integration with the chosen streaming platform.
  • 30-day support after delivery.

Timelines

Client-side co-stream (iOS, two participants, Metal composition, basic signaling server): 5–7 weeks. Full implementation with MCU, Android support, AEC, state management: 8–12 weeks. Cost is calculated individually after requirement analysis and architecture selection.

Evaluate your project—contact us for a consultation. Get a consultation for your project—we will assess the requirements and propose the optimal solution.

How to Choose a Camera Approach on Mobile Platforms?

Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.

What are the Differences Between CameraX, Camera2, and AVFoundation?

On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.

CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.

When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.

On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.

A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.

Why is AudioFocus Critical for Android Apps?

Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus

On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.

AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.

On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.

Video: Playback and Streaming

ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.

Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.

AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.

For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).

Protocol Latency Application
RTMP 2–5 sec Streaming to YouTube/Twitch
HLS 6–30 sec VOD, broadcast
DASH 6–30 sec VOD with adaptive bitrate
WebRTC < 500 ms Video calls, P2P
SRT 1–4 sec Professional streaming

WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.

RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.

Media Processing: Compression and Transcoding

ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.

On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.

How to Implement Streaming on Mobile Devices: Step-by-Step Plan

  1. Define requirements: target latency, number of concurrent users, need for P2P.
  2. Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
  3. Set up signaling (SIP, WebSocket, MQTT) and TURN server.
  4. Implement publishing/viewing via native API or cross-platform plugin.
  5. Test on real devices with different cameras and network conditions.
  6. Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
  • Configuring AVFoundation session on the main thread.
  • Missing AudioFocus Loss handling on Android.
  • Ignoring MediaCodec limitations on cheap devices.
  • Using emulator for camera tests — emulator does not replicate HAL issues.
  • Memory leaks when recreating media players without a pool.

What is Included in the Work

Deliverable Description
Requirements analysis Stack selection, priorities, test devices
Design Architecture, data flow diagrams, API selection
Implementation Code using chosen tools
Backend integration GraphQL/REST, DRM, WebRTC signaling
Testing On real devices (at least 5 models)
Documentation API documentation, build instructions
Post-release support 1 month incident support, team training

Development Process for Media Functionality

Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).

Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.

Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.

Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.