Implementing Document Scanning via Mobile Camera
We are a team of mobile engineers with 7+ years of experience in computer vision on iOS and Android. Over this time, we have implemented scanning for passports, contracts, receipts, and book spreads. The user holds the phone over the document, the app automatically finds the edges of the sheet, corrects perspective, and outputs a clean PDF. This is not "take a photo and crop" — inside there is a contour detector (Canny, Hough), homographic transformation, and post-processing for readability. Each step can be ruined if you do not account for lighting conditions and document types. Contact us to evaluate your project — we will help you choose the optimal solution.
Why Edge Detection Fails on Glare and Shadows
On iOS, VNDetectRectanglesRequest (Vision) returns VNRectangleObservation with four corner points in normalized coordinates. The problem is that on glossy paper under direct light, the algorithm confuses glare with the edge of the sheet. Solution: before detection, apply CIFilter with CIColorControls (reduce inputSaturation) and CIHighlightShadowAdjust. This removes glare as color artifacts. Additionally, you can increase contrast (level 1.2–1.5) to better separate edges.
On Android, ML Kit Document Scanner (com.google.android.gms:play-services-mlkit-document-scanner) handles shadows better but requires Google Play Services. An alternative without GMS dependency is OpenCV findContours + approxPolyDP with a filter by area and aspect ratio. A threshold of minArea = 30% of frame area filters out background objects. More about the algorithm in OpenCV documentation. For Flutter, we use a native channel via cunning_document_scanner, which delegates detection to the platform.
How to Choose Between Native SDK and OpenCV?
The choice depends on the ecosystem. If the app uses Google Play Services, ML Kit provides a ready-made UI and good accuracy. For devices without GMS (e.g., Huawei) — OpenCV. On iOS, Vision is the optimal choice since 2017, supports Live Photos and Metal acceleration. However, OpenCV requires licensing considerations (BSD) and more code. Performance: on iPhone 13, Vision detection takes ~80 ms, OpenCV (~120 ms with NEON optimization).
How to Correctly Perform Perspective Correction
After obtaining four points, apply perspective transform. iOS: CIPerspectiveCorrection with explicit passing of inputTopLeft, inputTopRight, inputBottomLeft, inputBottomRight in image coordinates (not preview). A common mistake is using preview-layer coordinates directly without recalculating via VNImagePointForNormalizedPoint. Android: getPerspectiveTransform + warpPerspective from OpenCV or matrix transform via android.graphics.Matrix.setPolyToPoly. The latter works without OpenCV but is limited to affine transformations — not suitable for strong perspective distortion. On Flutter — manual homography calculation using the image package or a native channel.
Technical Implementation of Perspective Correction
For iOS: after obtaining points from VNRectangleObservation, transform them to image coordinates via VNImagePointForNormalizedPoint. Then pass to CIPerspectiveCorrection. For debugging, draw the contour on AVCaptureVideoPreviewLayer via CAShapeLayer updating every 5 frames. On Android: use getPerspectiveTransform from OpenCV, but for non-OpenCV paths — setPolyToPoly with PST (perspective transform) via Matrix. Important: under strong distortion, affine transforms give up to 15% error at edges.
Post-Processing: Readability Over Beauty
After straightening, the document needs processing for readability when printing or OCR:
-
Adaptive binarization —
cv::adaptiveThresholdwith Gaussian method works better than Otsu on documents with uneven lighting. - Deskew — if the document is rotated by 1–2° after transformation, Hough Lines find the slope of text lines and correct it.
-
Sharpness —
CISharpenLuminance(iOS) or Sharpness filter (Android) with a moderate value (0.4–0.6), no more.
Color modes should be given to the user: "Auto", "Document" (black & white), "Photo" (full color). In "Document" mode — binarization. In "Auto" — histogram analysis: if the document contains <5% saturated pixels, apply monochrome processing.
| Stage | iOS | Android | Flutter |
|---|---|---|---|
| Detection | Vision (VNDetectRectanglesRequest) | ML Kit Document Scanner / OpenCV | cunning_document_scanner / channel |
| Transformation | CIPerspectiveCorrection | OpenCV warpPerspective / Matrix.setPolyToPoly | Dart manual (image package) |
| Post-processing | CIFilters (Sharpen, Binarization) | OpenCV adaptiveThreshold + deskew | Platform channel / dart filters |
| PDF Export | PDFKit (UIGraphicsPDFRenderer) | android.graphics.pdf.PdfDocument | pdf package (pub.dev) |
Performance Overview on Different Platforms
| Parameter | iOS (iPhone 13) | Android (Pixel 6) | Flutter (native channel) |
|---|---|---|---|
| Detection time | ~80 ms | ~110 ms | ~150 ms (with bridge overhead) |
| PDF size (A4) | ~200 KB | ~220 KB | ~230 KB |
| Preview frame rate | 30 FPS | 30 FPS | 24 FPS |
Multi-Page Scan and PDF
We collect UIImage[] / Bitmap[], export via PDFKit (iOS 11+) or android.graphics.pdf.PdfDocument. On Flutter — the pdf package (pub.dev). We optimize PDF size: JPEG compression 85% is sufficient for readability, with an A4 page taking ~150–250 KB versus 2–4 MB for PNG. Real-time preview: we show the contour on top of AVCaptureVideoPreviewLayer / PreviewView via CAShapeLayer / SurfaceView. We update the contour every 3–5 frames (not every frame) — otherwise the detector consumes CPU and the preview lags.
What Is Included in the Scanning Integration Work
- Requirements audit: analysis of document types, shooting conditions, target platforms.
- SDK selection: native Vision/ML Kit vs OpenCV vs ready-made solutions.
- Integration of preview with dynamic contour overlay.
- Implementation of detection, perspective correction, and post-processing.
- PDF export with compression and color mode settings.
- Testing on 10+ document types: passport, contract, receipt, book spread.
- API documentation and source code delivery.
Our experience: over 50 successful scanning projects, over 5 years in mobile development. We guarantee stable operation on modern devices. Order document scanning integration into your application — contact us for timeline and cost estimation. Timelines: from 3 to 5 business days per platform, plus 2–3 days with OCR.
Additional information: homographic transformation is a key element of perspective correction.







