Ivan, owner of a smart intercom, complained: the app transmitted sound with a delay of about 900 ms. Conversations turned into chaos. We migrated him to WebRTC — RTT dropped to 200 ms, reducing latency by 78%. Our 5+ years of proven experience shows that choosing the right stack solves 80% of echo and latency issues.
Intercoms, baby monitors, walkie-talkies — common denominator: the phone hears the device and simultaneously speaks into it. Unlike a regular VoIP call between two phones, here one side is an embedded Linux microcomputer (ESP32, Raspberry Pi, NXP i.MX) that does not support SIP or WebRTC without additional software. This fundamentally changes the choice of architecture. For intelligible speech, RTT must not exceed 300–400 ms, otherwise dialogue becomes impossible. We have helped 30+ clients solve this problem.
How to Ensure Minimal Latency
Round-trip time (RTT) for intelligible speech should be no more than 300–400 ms. HLS and RTMP are not suitable. SIP is possible but has protocol overhead. WebRTC was designed specifically for this scenario. WebRTC establishes connections 3 times faster than SIP due to ICE + STUN. A typical WebRTC project starts at $2,000, saving up to 40% on infrastructure compared to SIP servers.
IoT device side: libwebrtc on Linux or specialized solutions: aiortc (Python), Pion (Go), GStreamer with webrtcbin plugin. Pion is minimal and easy to deploy on Raspberry Pi. GStreamer webrtcbin — if the device already uses GStreamer.
Mobile app side:
iOS: GoogleWebRTC (pod 'WebRTC-SDK') or native WebRTCFramework. Create RTCPeerConnection with audio track:
let audioConstraints = RTCMediaConstraints(mandatoryConstraints: nil, optionalConstraints: nil)
let audioSource = factory.audioSource(with: audioConstraints)
let audioTrack = factory.audioTrack(with: audioSource, trackId: "audio0")
peerConnection.add(audioTrack, streamIds: ["stream0"])
Configure RTCAudioSession with .voiceChat category — automatically enables echo suppression and noise suppression (AEC/NS) built into WebRTC.
Android: io.getstream:stream-webrtc-android or org.webrtc:google-webrtc. AudioManager.MODE_IN_COMMUNICATION is mandatory for correct audio routing (earpiece/speakerphone).
Flutter: flutter_webrtc. Configure mediaConstraints for audio:
final Map<String, dynamic> mediaConstraints = {
'audio': {
'echoCancellation': true,
'noiseSuppression': true,
'autoGainControl': true,
}
};
Echo Cancellation: The Main Pain of Two-Way Audio
Without AEC (Acoustic Echo Cancellation): the phone's microphone picks up sound from the speaker (or vice versa — the device hears itself) — the user hears echo with 200 ms delay. Unusable.
WebRTC includes built-in AEC3 (third generation). It works automatically when the audio session category is set correctly. Problems arise when:
- The IoT device does not support an echo reference path — then AEC on the device side is ineffective. Solution: offload AEC to the server side (media server with processing enabled).
- Bluetooth headset + WebRTC — on Android
AudioManager in COMMUNICATION mode switches BT profile to HFP (narrowband 8 kHz). For wideband audio, A2DP is needed, but it does not support recording. Compromise: either low quality with BT, or AirPods/wired headphones.
Why WebRTC is Better than SIP for IoT
| Parameter |
WebRTC |
SIP |
| Connection setup time |
<500 ms |
1–3 s |
| Server requirements |
STUN/TURN |
SIP server (Asterisk) |
| Echo cancellation |
Built-in AEC3 |
Depends on implementation |
| NAT Traversal |
ICE (automatic) |
Requires configuration |
WebRTC does not require server registration and establishes connections faster. SIP is indispensable when integrating with existing telephony (e.g., Grandstream IP intercoms). According to RFC 7874, WebRTC with Opus codec provides speech quality comparable to PSTN.
SIP as an Alternative
If the IoT device supports SIP (many IP intercoms: Grandstream, Panasonic, Commax), use a SIP client on mobile.
iOS: PJSIP (C library) with Swift wrapper or Linphone SDK. Android: MjSip or PJSIP via JNI, or the ready-made Linphone SDK for Android. Flutter: sip_ua (Dart SIP, works over WebSocket transport).
SIP on mobile requires registration on an Asterisk/FreeSWITCH server. Call from intercom → SIP INVITE → server → push notification to phone (via CallKit on iOS, ConnectionService/IncomingCallNotification on Android). Without push, the notification does not arrive when the app is closed.
CallKit (iOS): incoming call appears as a regular phone call — full-screen interface with the intercom name. CXProvider, CXCallUpdate — standard integration. Requires voip Background Mode in Info.plist + APNs VoIP certificate.
Android ConnectionService: analogous to CallKit. TelecomManager.addNewIncomingCall() — shows system incoming call interface. Works from Android 6+.
Environmental Noise and Aggressive Noise Suppression
Outside wind, construction nearby — the IoT device sends a noisy stream. Additional noise suppression: RTCRtpSender with RTCDefaultVideoEncoderFactory — audio only. WebRTC RNNoise is integrated into native WebRTC and enabled via AudioProcessing::Config::NoiseSuppression.
For heavy server-side processing: Janus with janus_audiobridge plugin applies noise suppression before mixing.
Step-by-Step WebRTC Setup on IoT Device
- Install Pion library on the device (Go):
go get github.com/pion/webrtc/v3.
- Configure ICE using a public STUN server (e.g.,
stun:stun.l.google.com:19302).
- Create an audio track with Opus 48000 Hz parameters.
- On the mobile app side, create
RTCPeerConnection and add an audio track.
- Exchange SDP offer and answer via a signaling server (WebSocket or MQTT).
- Test the connection: use Coturn for testing the TURN server.
Testing
The main challenge: simulating NAT traversal in a test environment. Use Coturn in Docker for local TURN testing. Test on symmetric NAT (corporate network with strict rules) — mandatory. Without a TURN server, approximately 15–20% of connections will not establish.
Timeline: two-way WebRTC audio with IoT device (Linux/Pion) + iOS or Android client — 5–7 business days. With SIP integration and CallKit — 8–12 days.
What's Included
- Documentation for WebRTC/SIP integration into your app.
- Source code repositories (iOS/Android/Flutter + IoT).
- Deployment instructions for the device.
- Support for 2 weeks after delivery.
With over 5 years of proven expertise in IoT audio communication, we guarantee reliable, low-latency connections. We have completed 30+ projects with two-way audio, from intercoms to industrial talkback devices. Infrastructure savings up to 40% compared to SIP servers. Get a consultation on your project — contact us.
How to Choose a Camera Approach on Mobile Platforms?
Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.
What are the Differences Between CameraX, Camera2, and AVFoundation?
On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.
CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.
When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.
On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.
A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.
Why is AudioFocus Critical for Android Apps?
Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus
On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.
AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.
On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.
Video: Playback and Streaming
ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.
Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.
AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.
For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).
| Protocol |
Latency |
Application |
| RTMP |
2–5 sec |
Streaming to YouTube/Twitch |
| HLS |
6–30 sec |
VOD, broadcast |
| DASH |
6–30 sec |
VOD with adaptive bitrate |
| WebRTC |
< 500 ms |
Video calls, P2P |
| SRT |
1–4 sec |
Professional streaming |
WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.
RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.
Media Processing: Compression and Transcoding
ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.
On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.
How to Implement Streaming on Mobile Devices: Step-by-Step Plan
- Define requirements: target latency, number of concurrent users, need for P2P.
- Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
- Set up signaling (SIP, WebSocket, MQTT) and TURN server.
- Implement publishing/viewing via native API or cross-platform plugin.
- Test on real devices with different cameras and network conditions.
- Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
- Configuring AVFoundation session on the main thread.
- Missing AudioFocus Loss handling on Android.
- Ignoring
MediaCodec limitations on cheap devices.
- Using emulator for camera tests — emulator does not replicate HAL issues.
- Memory leaks when recreating media players without a pool.
What is Included in the Work
| Deliverable |
Description |
| Requirements analysis |
Stack selection, priorities, test devices |
| Design |
Architecture, data flow diagrams, API selection |
| Implementation |
Code using chosen tools |
| Backend integration |
GraphQL/REST, DRM, WebRTC signaling |
| Testing |
On real devices (at least 5 models) |
| Documentation |
API documentation, build instructions |
| Post-release support |
1 month incident support, team training |
Development Process for Media Functionality
Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).
Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.
Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.
Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.