Implementing Document Scanning via Mobile Camera
We are a team of mobile engineers with 7+ years of experience in computer vision on iOS and Android. Over this time, we have implemented scanning for passports, contracts, receipts, and book spreads. The user holds the phone over the document, the app automatically finds the edges of the sheet, corrects perspective, and outputs a clean PDF. This is not "take a photo and crop" — inside there is a contour detector (Canny, Hough), homographic transformation, and post-processing for readability. Each step can be ruined if you do not account for lighting conditions and document types. Contact us to evaluate your project — we will help you choose the optimal solution.
Why Edge Detection Fails on Glare and Shadows
On iOS, VNDetectRectanglesRequest (Vision) returns VNRectangleObservation with four corner points in normalized coordinates. The problem is that on glossy paper under direct light, the algorithm confuses glare with the edge of the sheet. Solution: before detection, apply CIFilter with CIColorControls (reduce inputSaturation) and CIHighlightShadowAdjust. This removes glare as color artifacts. Additionally, you can increase contrast (level 1.2–1.5) to better separate edges.
On Android, ML Kit Document Scanner (com.google.android.gms:play-services-mlkit-document-scanner) handles shadows better but requires Google Play Services. An alternative without GMS dependency is OpenCV findContours + approxPolyDP with a filter by area and aspect ratio. A threshold of minArea = 30% of frame area filters out background objects. More about the algorithm in OpenCV documentation. For Flutter, we use a native channel via cunning_document_scanner, which delegates detection to the platform.
How to Choose Between Native SDK and OpenCV?
The choice depends on the ecosystem. If the app uses Google Play Services, ML Kit provides a ready-made UI and good accuracy. For devices without GMS (e.g., Huawei) — OpenCV. On iOS, Vision is the optimal choice since 2017, supports Live Photos and Metal acceleration. However, OpenCV requires licensing considerations (BSD) and more code. Performance: on iPhone 13, Vision detection takes ~80 ms, OpenCV (~120 ms with NEON optimization).
How to Correctly Perform Perspective Correction
After obtaining four points, apply perspective transform. iOS: CIPerspectiveCorrection with explicit passing of inputTopLeft, inputTopRight, inputBottomLeft, inputBottomRight in image coordinates (not preview). A common mistake is using preview-layer coordinates directly without recalculating via VNImagePointForNormalizedPoint. Android: getPerspectiveTransform + warpPerspective from OpenCV or matrix transform via android.graphics.Matrix.setPolyToPoly. The latter works without OpenCV but is limited to affine transformations — not suitable for strong perspective distortion. On Flutter — manual homography calculation using the image package or a native channel.
Technical Implementation of Perspective Correction
For iOS: after obtaining points from VNRectangleObservation, transform them to image coordinates via VNImagePointForNormalizedPoint. Then pass to CIPerspectiveCorrection. For debugging, draw the contour on AVCaptureVideoPreviewLayer via CAShapeLayer updating every 5 frames. On Android: use getPerspectiveTransform from OpenCV, but for non-OpenCV paths — setPolyToPoly with PST (perspective transform) via Matrix. Important: under strong distortion, affine transforms give up to 15% error at edges.
Post-Processing: Readability Over Beauty
After straightening, the document needs processing for readability when printing or OCR:
-
Adaptive binarization —
cv::adaptiveThreshold with Gaussian method works better than Otsu on documents with uneven lighting.
-
Deskew — if the document is rotated by 1–2° after transformation, Hough Lines find the slope of text lines and correct it.
-
Sharpness —
CISharpenLuminance (iOS) or Sharpness filter (Android) with a moderate value (0.4–0.6), no more.
Color modes should be given to the user: "Auto", "Document" (black & white), "Photo" (full color). In "Document" mode — binarization. In "Auto" — histogram analysis: if the document contains <5% saturated pixels, apply monochrome processing.
| Stage |
iOS |
Android |
Flutter |
| Detection |
Vision (VNDetectRectanglesRequest) |
ML Kit Document Scanner / OpenCV |
cunning_document_scanner / channel |
| Transformation |
CIPerspectiveCorrection |
OpenCV warpPerspective / Matrix.setPolyToPoly |
Dart manual (image package) |
| Post-processing |
CIFilters (Sharpen, Binarization) |
OpenCV adaptiveThreshold + deskew |
Platform channel / dart filters |
| PDF Export |
PDFKit (UIGraphicsPDFRenderer) |
android.graphics.pdf.PdfDocument |
pdf package (pub.dev) |
Performance Overview on Different Platforms
| Parameter |
iOS (iPhone 13) |
Android (Pixel 6) |
Flutter (native channel) |
| Detection time |
~80 ms |
~110 ms |
~150 ms (with bridge overhead) |
| PDF size (A4) |
~200 KB |
~220 KB |
~230 KB |
| Preview frame rate |
30 FPS |
30 FPS |
24 FPS |
Multi-Page Scan and PDF
We collect UIImage[] / Bitmap[], export via PDFKit (iOS 11+) or android.graphics.pdf.PdfDocument. On Flutter — the pdf package (pub.dev). We optimize PDF size: JPEG compression 85% is sufficient for readability, with an A4 page taking ~150–250 KB versus 2–4 MB for PNG. Real-time preview: we show the contour on top of AVCaptureVideoPreviewLayer / PreviewView via CAShapeLayer / SurfaceView. We update the contour every 3–5 frames (not every frame) — otherwise the detector consumes CPU and the preview lags.
What Is Included in the Scanning Integration Work
- Requirements audit: analysis of document types, shooting conditions, target platforms.
- SDK selection: native Vision/ML Kit vs OpenCV vs ready-made solutions.
- Integration of preview with dynamic contour overlay.
- Implementation of detection, perspective correction, and post-processing.
- PDF export with compression and color mode settings.
- Testing on 10+ document types: passport, contract, receipt, book spread.
- API documentation and source code delivery.
Our experience: over 50 successful scanning projects, over 5 years in mobile development. We guarantee stable operation on modern devices. Order document scanning integration into your application — contact us for timeline and cost estimation. Timelines: from 3 to 5 business days per platform, plus 2–3 days with OCR.
Additional information: homographic transformation is a key element of perspective correction.
How to Choose a Camera Approach on Mobile Platforms?
Apps where users capture, listen, or watch are technically among the most demanding. We deal with this every day. Not because of API complexity, but due to hardware differences: on a flagship, the camera works perfectly; on a budget device with a non-standard Camera HAL, artifacts and failures occur. On iOS, stabilization differs between generations. Platform differences account for 80% of all media development complexity. Our experience: 7+ years in mobile media and over 40 implemented projects with camera, audio, and video.
What are the Differences Between CameraX, Camera2, and AVFoundation?
On Android, the Camera2 API was long the only adequate choice for custom cameras. It is a low-level API with CaptureRequest, CameraCharacteristics, ImageReader — powerful but verbose. Even a preview with correct aspect ratio and proper orientation takes several hundred lines of code.
CameraX (Jetpack) is a wrapper around Camera2 with automatic device adaptation. Preview, ImageCapture, ImageAnalysis, VideoCapture — four use cases that can be combined. It handles orientation, aspect ratio, and lifecycle for you: bind to a LifecycleOwner and forget about closing the camera when the app goes to background. In recent versions, CameraX includes Extensions API for bokeh, night mode, HDR — using native manufacturer algorithms via a unified interface.
When is Camera2 needed directly?: RAW capture via ImageFormat.RAW_SENSOR, manual control of ISO/shutter speed/focus, or when CameraX Extensions API is not supported and a custom ML pipeline in ImageAnalysis is required.
On iOS, AVFoundation is the only path for a custom camera. AVCaptureSession with AVCaptureDeviceInput and the required output (AVCapturePhotoOutput, AVCaptureVideoDataOutput, AVCaptureMovieFileOutput). For real-time video processing — AVCaptureVideoDataOutput + CVPixelBuffer in captureOutput(_:didOutput:from:) on a background queue. This is where CoreML models receive frames for inference.
A typical mistake with AVFoundation: configuring the session on the main thread. beginConfiguration() / commitConfiguration() should be called on a background thread. Otherwise, the preview freezes, and the user sees a frozen UI. This mistake appears in 70% of the projects we have audited.
Why is AudioFocus Critical for Android Apps?
Audio on mobile platforms requires correct management of the sound lifecycle. AudioFocus is a coordination mechanism between apps. AudioManager.requestAudioFocus() with OnAudioFocusChangeListener. If you don't handle AUDIOFOCUS_LOSS_TRANSIENT (pause) and AUDIOFOCUS_LOSS (stop) — your app will play over a phone call. That guarantees a bad review on Google Play. Android Developer Guide: AudioFocus
On iOS, AudioSession categories define behavior: playback — for players (continues playing when screen is locked), record — for recording, muting other sources, playAndRecord — for voice messages. Wrong category — the app mutes the user's background music on start.
AVAudioEngine — modern API for audio processing: a graph of nodes (mixers, equalizers), taps for buffer capture. For real-time speech — SFSpeechRecognizer + inputNode.installTap.
On Android for recording with noise suppression — NoiseSuppressor.isAvailable() + create(audioRecord.audioSessionId). Works not on all devices, need a fallback.
Video: Playback and Streaming
ExoPlayer (Media3) — standard for Android. Supports HLS, DASH, SmoothStreaming, progressive playback. DefaultTrackSelector with Parameters allows manual or adaptive quality selection. DRM via DefaultDrmSessionManager with Widevine L1/L3.
Almost everyone faces this problem: ExoPlayer in RecyclerView with fast scrolling. Need a PlayerPool — a pool of reusable players. Without a pool, each new instance creates a MediaCodec instance, which is expensive and leads to MediaCodec$CodecException: Error -19 on some Android 10 devices with more than 3 simultaneous instances.
AVPlayer / AVPlayerViewController on iOS — for playback. For custom UI — AVPlayerLayer + custom controls. HLS works natively via AVPlayer(url:) with m3u8. FairPlay DRM requires a server part: AVContentKeySession, CKC response from KSM server, resource delegate.
For Flutter — video_player as a base layer, chewie for UI. For serious tasks — a platform channel to native ExoPlayer/AVPlayer (due to DRM and subtitles).
| Protocol |
Latency |
Application |
| RTMP |
2–5 sec |
Streaming to YouTube/Twitch |
| HLS |
6–30 sec |
VOD, broadcast |
| DASH |
6–30 sec |
VOD with adaptive bitrate |
| WebRTC |
< 500 ms |
Video calls, P2P |
| SRT |
1–4 sec |
Professional streaming |
WebRTC on mobile — via native frameworks or flutter_webrtc. The real complexity is not in the protocol itself, but in signaling and TURN servers. Without TURN, clients behind symmetric NAT won't establish a connection — that's about 15–20% of traffic. Coturn is the standard open-source server.
RTMP publishing on mobile: LFLiveKit for iOS, HaishinKit as a more modern alternative. On Android — rtmp-rtsp-stream-client-java or via FFmpeg with JNI. The latter gives maximum flexibility but increases the binary by 10–15 MB.
Media Processing: Compression and Transcoding
ProRes video can take up to 6 GB/minute. Compression is needed before upload. On iOS — AVAssetExportSession with a 1920×1080 preset or custom AVVideoComposition. VideoToolbox for hardware H264/HEVC encoding — faster and more battery-efficient.
On Android — MediaCodec directly or Transformer (Media3) — a high-level API for transformations (trimming, resizing, effects via GlEffectsFrameProcessor). For images — BitmapFactory.Options.inSampleSize for downsampling, Glide / Coil for caching. Coil on Coroutines fits well with Compose. Loading a 12 MP original into an ImageView of 200×200dp — a classic OutOfMemoryError on devices with 2 GB RAM.
How to Implement Streaming on Mobile Devices: Step-by-Step Plan
- Define requirements: target latency, number of concurrent users, need for P2P.
- Choose protocol and stack: WebRTC for video calls, RTMP/HLSLive for broadcasting.
- Set up signaling (SIP, WebSocket, MQTT) and TURN server.
- Implement publishing/viewing via native API or cross-platform plugin.
- Test on real devices with different cameras and network conditions.
- Optimize bitrate and resolution based on bandwidth.
Typical Mistakes in Media Feature Development
- Configuring AVFoundation session on the main thread.
- Missing AudioFocus Loss handling on Android.
- Ignoring
MediaCodec limitations on cheap devices.
- Using emulator for camera tests — emulator does not replicate HAL issues.
- Memory leaks when recreating media players without a pool.
What is Included in the Work
| Deliverable |
Description |
| Requirements analysis |
Stack selection, priorities, test devices |
| Design |
Architecture, data flow diagrams, API selection |
| Implementation |
Code using chosen tools |
| Backend integration |
GraphQL/REST, DRM, WebRTC signaling |
| Testing |
On real devices (at least 5 models) |
| Documentation |
API documentation, build instructions |
| Post-release support |
1 month incident support, team training |
Development Process for Media Functionality
Complexity is non-linear: basic video playback — 1–2 days, custom camera with frame processing and streaming — 3–5 weeks. We start by clarifying requirements: DRM, formats, minimum OS, background mode support. Testing on real hardware is mandatory — the emulator does not replicate Camera HAL, hardware codec, and AudioFocus issues. Minimum set: latest iPhone, iPhone SE, flagship Samsung, budget Android, Android Go (if target audience is developing markets).
Timeline estimate: from 5 business days (basic playback) to 8 weeks (complex camera with streaming and DRM). Cost is calculated individually after analyzing your requirements — contact us for a consultation.
Our service: "Mobile Media Integration" — this is our expertise. Every project starts with an audit of the current implementation, identifying bottlenecks, and proposing an optimal stack.
Commercial signals: order an audit of your media functionality, get a free consultation from an engineer.