We integrate screen sharing into iOS and Android apps using ReplayKit and MediaProjection APIs. With over 50 projects delivered and certified Apple and Google engineers, we ensure stable production-ready solutions. This article covers architecture, frame transmission, and common pitfalls.
How Does ReplayKit Work on iOS?
On iOS, screen capture is only possible via ReplayKit. You cannot get the screen content directly from the main app process—this is a sandbox restriction. For real-time broadcasting, we use RPSystemBroadcastPickerView, which shows a system picker to launch the broadcast. The broadcast runs via a Broadcast Upload Extension—a separate target in the project that runs in its own process (com.apple.broadcast-services-upload). The extension and the main app have no shared memory—communication is only through App Group (shared UserDefaults, shared FileManager, or Darwin notify for signals).
Structure of the extension:
class SampleHandler: RPBroadcastSampleHandler {
override func processSampleBuffer(_ sampleBuffer: CMSampleBuffer,
with sampleBufferType: RPSampleBufferType) {
switch sampleBufferType {
case .video:
// Pass CMSampleBuffer to WebRTC/Agora/Livekit
videoSource?.capturer(capturer, didCapture: toRTCVideoFrame(sampleBuffer))
case .audioApp:
// System audio
case .audioMic:
// Microphone (only available with explicit permission in Info.plist)
}
}
}
The extension has a memory limit of 50 MB—in 90% of cases, crashes are due to exceeding this limit. Agora and Livekit provide lightweight versions of their SDK specifically for Broadcast Extension. Communication from extension to main app is via CFNotificationCenter (Darwin notifications). Resolution and FPS are limited by ReplayKit: max 1080p, up to 60 FPS. In practice, 720p at 15 FPS is sufficient for screen sharing—mobile device screens are small, and bitrate drops by 40% without loss of text readability.
Apple Developer Documentation: ReplayKit
What is the MediaProjection API on Android?
On Android, screen capture uses the MediaProjection API. The user must explicitly allow recording:
val mediaProjectionManager = getSystemService(MEDIA_PROJECTION_SERVICE) as MediaProjectionManager
startActivityForResult(
mediaProjectionManager.createScreenCaptureIntent(),
REQUEST_MEDIA_PROJECTION
)
In onActivityResult, you get an Intent with permission—from it you create a MediaProjection. VirtualDisplay with MediaProjection.createVirtualDisplay() provides a Surface onto which the system draws the screen content.
For transmission via WebRTC, we use ScreenCapturerAndroid (from org.webrtc): it wraps MediaProjection and delivers frames to VideoSource. Since Android 10, when starting screen capture via ForegroundService, you need android:foregroundServiceType="mediaProjection" in the manifest. On Android 14, additionally you must explicitly specify ServiceInfo.FOREGROUND_SERVICE_TYPE_MEDIA_PROJECTION when starting the service via startForeground(). Without this—SecurityException. When the connection is broken, MediaProjection.Callback.onStop() is called—must be handled and the user notified.
Android Developers: MediaProjection
How Are Frames Transmitted in Real Time?
On both iOS and Android, captured screen frames are CMSampleBuffer (iOS) or Bitmap/Image via ImageReader (Android). For network transmission over WebRTC, we convert to RTCVideoFrame with YUV420PlanarBuffer (iOS) or JavaI420Buffer (Android). YUV conversion from BGRA is an expensive operation—we perform it in native code (C++) or via a hardware converter. For Agora, AgoraRtcKit.startScreenCapture() accepts configuration with contentHint: .text to optimize the codec for textual content (less motion, high sharpness). WebRTC is more flexible than Agora in bitrate configuration, but Agora is easier to integrate with Broadcast Extension. For high-load projects (1000+ simultaneous streams), WebRTC offers 30% lower infrastructure costs compared to Agora.
Comparison of iOS and Android
| Parameter |
iOS (ReplayKit) |
Android (MediaProjection) |
| Minimum version |
iOS 12+ |
Android 5.0 (API 21) |
| Screen capture |
Broadcast Extension (separate process) |
Direct access via Intent |
| Memory limit |
50 MB on extension |
No hard limit |
| Max resolution |
1080p |
Depends on device, often 1080p+ |
| FPS |
Up to 60 |
Up to 30 (typically) |
| Audio capture |
System + microphone |
Microphone via AudioRecord |
| Foreground service |
Not required |
Mandatory (Android 10+) |
Typical Screen Sharing Issues
| Issue |
Solution |
| Broadcast Extension crash on iOS |
Use lightweight SDK (Agora/Livekit) and monitor the 50 MB limit |
| SecurityException on Android 14 |
Explicitly set foregroundServiceType="mediaProjection" and call startForeground() with the correct type |
| User stops capture via Android notification shade |
Handle onStop() and correctly switch UI without errors |
| Delay during YUV conversion |
Use hardware converter or native C++ code |
Technical Details of Broadcast Extension
Broadcast Extension lives in a separate process with a 50 MB limit. If the SDK weighs more—the extension crashes without a clear message. Agora and Livekit provide "lite" versions for extensions. For data exchange between the extension and the main app, we use Darwin notifications via CFNotificationCenter. The extension cannot establish network connections directly—it only delivers frames, while the main app manages the WebRTC peer.
What Is Included in Screen Sharing Implementation?
Our deliverables include:
- Architecture and design: stack selection (WebRTC/Agora/Livekit), frame exchange scheme, lifecycle handling.
- Development of Broadcast Extension (iOS) and MediaProjection Service (Android): full capture and broadcasting cycle.
- Integration with signaling and WebRTC: peer connection configuration, ICE, relay setup.
- Permission management: screen recording requests, ATT, foreground service.
- Error handling and edge cases: extension crash, user cancels capture, network loss.
- Documentation and training: integration description, commented source code transfer, one training session.
- Testing: on real devices with 10+ different OS versions.
- Support: 3 months of post-deployment support and maintenance.
How to Choose the Stack for Screen Sharing?
Choosing between WebRTC, Agora, and Livekit depends on customization needs and time to market. WebRTC gives full control over bitrate and codec but requires more integration time—on average 1.5 weeks for both platforms. Agora and Livekit reduce this to 5 days thanks to ready modules, but limit configuration flexibility. For high-load projects (1000+ simultaneous broadcasts), WebRTC is 30% cheaper in infrastructure costs.
Estimated Timelines and Costs
iOS (ReplayKit + Broadcast Extension + WebRTC) — from 3 to 5 days if signaling is ready. Android (MediaProjection + WebRTC) — from 2 to 3 days. Both platforms with correct lifecycle, background mode, and UX notifications — from 1 to 1.5 weeks. Typical project cost ranges from $5,000 to $15,000 depending on complexity and stack. Contact us for a detailed quote—our certified engineers guarantee seamless integration and 100% pass through app store reviews.
How to Choose a Camera Approach on Mobile Platforms?
Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.
What are the Differences Between CameraX, Camera2, and AVFoundation?
On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.
CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.
When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.
On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.
A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.
Why is AudioFocus Critical for Android Apps?
Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus
On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.
AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.
On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.
Video: Playback and Streaming
ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.
Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.
AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.
For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).
| Protocol |
Latency |
Application |
| RTMP |
2–5 sec |
Streaming to YouTube/Twitch |
| HLS |
6–30 sec |
VOD, broadcast |
| DASH |
6–30 sec |
VOD with adaptive bitrate |
| WebRTC |
< 500 ms |
Video calls, P2P |
| SRT |
1–4 sec |
Professional streaming |
WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.
RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.
Media Processing: Compression and Transcoding
ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.
On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.
How to Implement Streaming on Mobile Devices: Step-by-Step Plan
- Define requirements: target latency, number of concurrent users, need for P2P.
- Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
- Set up signaling (SIP, WebSocket, MQTT) and TURN server.
- Implement publishing/viewing via native API or cross-platform plugin.
- Test on real devices with different cameras and network conditions.
- Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
- Configuring AVFoundation session on the main thread.
- Missing AudioFocus Loss handling on Android.
- Ignoring
MediaCodec limitations on cheap devices.
- Using emulator for camera tests — emulator does not replicate HAL issues.
- Memory leaks when recreating media players without a pool.
What is Included in the Work
| Deliverable |
Description |
| Requirements analysis |
Stack selection, priorities, test devices |
| Design |
Architecture, data flow diagrams, API selection |
| Implementation |
Code using chosen tools |
| Backend integration |
GraphQL/REST, DRM, WebRTC signaling |
| Testing |
On real devices (at least 5 models) |
| Documentation |
API documentation, build instructions |
| Post-release support |
1 month incident support, team training |
Development Process for Media Functionality
Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).
Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.
Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.
Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.