Imagine: a user starts a stream, but within a minute viewers complain about delays and artifacts. Typical scenario: the mobile app uses software encoding, and the network can't handle the peak bitrate. Result—loss of audience and negative reviews. Overheating and rapid battery drain are another problem of software encoding. We know how to avoid this. Our team has implemented over 20 live streaming projects for iOS and Android, ensuring stability even on unstable connections. Resource savings: hardware encoding reduces power consumption by 90% compared to software. Behind every stream is a complex pipeline: camera capture, hardware encoding H.264/H.265, container packing, sending to a media server, and delivery to CDN. Each stage has its own API and pitfalls. Let's break them down in detail.
Video and Audio Capture
AVCaptureSession is the entry point on iOS. We configure the quality preset (sessionPreset = .hd1280x720), add AVCaptureDeviceInput for camera and microphone, add AVCaptureVideoDataOutput and AVCaptureAudioDataOutput with delegates. captureOutput(_:didOutput:from:) delivers CMSampleBuffer—raw camera frames in real time.
Orientation: AVCaptureConnection.videoOrientation must be updated on device rotation. Otherwise, viewers see video sideways if the stream starts in portrait.
On Android, we use Camera2 API or CameraX (recommended). CameraX.bindToLifecycle() with Preview + VideoCapture use case. VideoCapture.output.prepareRecording() is native recording. For streaming we need a raw stream: ImageAnalysis use case with setOutputImageFormat(OUTPUT_IMAGE_FORMAT_YUV_420_888), frames processed manually via MediaCodec.
| Platform |
Capture API |
Frame Stream |
| iOS |
AVCaptureSession |
CMSampleBuffer |
| Android |
CameraX/Camera2 |
YUV_420_888 (Image) |
Why Hardware Encoding?
Software FFmpeg on a phone means overheating and a dead battery in 20 minutes. We use only hardware encoders: VideoToolbox on iOS and MediaCodec on Android. This is 10 times more energy-efficient and delivers stable 30 fps at 720p. Our certified developers guarantee reliable hardware encoding integration.
On iOS: VideoToolbox—VTCompressionSession. We create a session with kVTVideoEncoderSpecification_RequireHardwareAcceleratedVideoEncoder: kCFBooleanTrue. The outputCallback receives CMSampleBuffer with encoded H.264/H.265.
Key parameters:
-
kVTCompressionPropertyKey_RealTime: kCFBooleanTrue — real-time mode
-
kVTCompressionPropertyKey_ProfileLevel: kVTProfileLevel_H264_High_AutoLevel
-
kVTCompressionPropertyKey_AverageBitRate: 2_000_000 (2 Mbps for 720p)
-
kVTCompressionPropertyKey_MaxKeyFrameInterval: 60 (keyframe every 2 sec at 30 fps)
On Android: MediaCodec. MediaFormat.createVideoFormat("video/avc", width, height). configure(format, null, null, MediaCodec.CONFIGURE_FLAG_ENCODE). In dequeueOutputBuffer we get encoded NAL units.
For audio: AudioRecord → MediaCodec with audio/mp4a-latm (AAC-LC). Sample rate 44100 Hz, bitrate 128 kbps.
How to Ensure Stream Stability?
Adaptive bitrate (ABR) is the key mechanism for camera live streaming. We monitor upload speed: RTMPStream.info.byteCount in HaishinKit, NetworkInfo callback in rtmp-rtsp-stream-client-java. If bandwidth drops, we sequentially reduce bitrate, framerate, and resolution. Optimal check interval is 5-10 seconds to avoid sharp quality swings. This approach reduces interruptions by 70%.
HaishinKit is a popular library for iOS handling RTMP and SRT streams. For Android, we similarly use rtmp-rtsp-stream-client-java.
Packaging into RTMP or SRT
From CMSampleBuffer/MediaCodec output we get H.264 NAL units and AAC frames. They need to be packed into a transport protocol. We use SRT for its network resilience, or RTMP for broad CDN compatibility.
Ready libraries:
- iOS: HaishinKit (Swift)—RTMP, SRT, HLS. Native Swift without FFmpeg dependency.
RTMPStream, SRTStream.
- Android:
rtmp-rtsp-stream-client-java (Pedro Vicente)—RTMP and RTSP push. RtmpCamera2 with CameraX. SRTStream for SRT.
- Flutter:
rtmp_streaming (wrapper over native libraries) or native platform channel.
- Cross-platform with FFmpeg:
ffmpeg-kit-ios / ffmpeg-kit-android—FFmpegKit.executeAsync("-f avfoundation -i 0:0 -c:v libx264 -preset ultrafast -f flv rtmp://..."). Works but loads CPU more than hardware encoder.
Comparison of RTMP and SRT
SRT is 2-3 times better than RTMP in latency and loss resilience.
| Characteristic |
RTMP |
SRT |
| Latency |
2-5 sec |
0.5-2 sec |
| Loss resilience |
Low |
High (FEC/ARQ) |
| CDN support |
Wide |
Growing |
| Use case |
Traditional streams |
Weak networks, webinars |
For mobile streaming, we recommend SRT because of its network resilience.
Preview and UI
Camera preview—AVCaptureVideoPreviewLayer (iOS) / PreviewView CameraX (Android). Overlay on top: "Live" indicator, viewer count, bitrate, signal level.
Switching front/rear camera without interrupting the stream: iOS—AVCaptureSession.beginConfiguration() → remove old input → add new → commitConfiguration(). HaishinKit supports this in one line: stream.captureSettings.isVideoMirrored.
What's Included in Live Streaming Work
- Video/audio capture from camera
- Hardware encoding with parameter tuning
- Integration of chosen protocol (RTMP/SRT/HLS)
- Adaptive bitrate and network monitoring
- Preview UI with overlay indicators
- Testing on real devices in various networks
- Integration documentation and post-launch support
Contact us to discuss your project. Order live streaming development—get a consultation from an experienced engineer. You can also request examples of our projects. We guarantee quality and proven experience with over 20 successful implementations.
Timelines
Single-platform streaming from camera (RTMP or SRT) with preview and basic UI: 3-5 days. Cross-platform with adaptive bitrate, camera switching, and specific media server integration: 1-2 weeks. Development cost starts from $3,000, with detailed quote after requirements analysis.
How to Choose a Camera Approach on Mobile Platforms?
Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.
What are the Differences Between CameraX, Camera2, and AVFoundation?
On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.
CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.
When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.
On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.
A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.
Why is AudioFocus Critical for Android Apps?
Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus
On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.
AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.
On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.
Video: Playback and Streaming
ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.
Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.
AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.
For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).
| Protocol |
Latency |
Application |
| RTMP |
2–5 sec |
Streaming to YouTube/Twitch |
| HLS |
6–30 sec |
VOD, broadcast |
| DASH |
6–30 sec |
VOD with adaptive bitrate |
| WebRTC |
< 500 ms |
Video calls, P2P |
| SRT |
1–4 sec |
Professional streaming |
WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.
RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.
Media Processing: Compression and Transcoding
ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.
On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.
How to Implement Streaming on Mobile Devices: Step-by-Step Plan
- Define requirements: target latency, number of concurrent users, need for P2P.
- Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
- Set up signaling (SIP, WebSocket, MQTT) and TURN server.
- Implement publishing/viewing via native API or cross-platform plugin.
- Test on real devices with different cameras and network conditions.
- Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
- Configuring AVFoundation session on the main thread.
- Missing AudioFocus Loss handling on Android.
- Ignoring
MediaCodec limitations on cheap devices.
- Using emulator for camera tests — emulator does not replicate HAL issues.
- Memory leaks when recreating media players without a pool.
What is Included in the Work
| Deliverable |
Description |
| Requirements analysis |
Stack selection, priorities, test devices |
| Design |
Architecture, data flow diagrams, API selection |
| Implementation |
Code using chosen tools |
| Backend integration |
GraphQL/REST, DRM, WebRTC signaling |
| Testing |
On real devices (at least 5 models) |
| Documentation |
API documentation, build instructions |
| Post-release support |
1 month incident support, team training |
Development Process for Media Functionality
Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).
Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.
Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.
Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.