Implementing Image Recognition in Mobile Apps
We implement turnkey image recognition: from camera capture to result display. With extensive experience, we've encountered all typical pitfalls — accuracy at 90% on test sets often drops to 70% on real photos. Below is how to avoid this and build a production-ready pipeline. Using ready-made pipelines (Core ML Vision, TFLite Task Library) cuts development time by 40% compared to custom implementation. We guarantee stable operation on all devices.
Image Sources and Their Peculiarities
Camera via AVCaptureSession (iOS) or CameraX (Android) is the most challenging case. Data arrives as CMSampleBuffer / ImageProxy in YUV_420_888 or BGRA format. Models expect RGB float32 or uint8. YUV → RGB conversion without native code introduces up to 40 ms delay. On Android we use ImageAnalysis.Builder().setOutputImageFormat(ImageAnalysis.OUTPUT_IMAGE_FORMAT_RGBA_8888) — this directly gives the needed format without manual conversion.
Gallery is simpler but has an EXIF orientation pitfall. UIImage on iOS correctly accounts for orientation when displaying, but the underlying CGImage may be rotated. Passing CGImage directly to the model degrades accuracy for vertically shot photos. The correct approach: CIImage(image: uiImage) → CIContext.createCGImage with applied orientation transformation.
On Android, BitmapFactory.decodeFile does not respect EXIF. Use ExifInterface with subsequent Matrix.postRotate. Otherwise the model receives a rotated image, reducing accuracy by 25%.
Why Accuracy Drops on Real Data?
Models are trained on datasets with specific distributions — user photos always have more variation (lighting, angle, background). Case study: a mushroom identification app. The EfficientNetV2-S model in Core ML achieved 91% on the test set but only 73% on user photos. Reason: the dataset was top-down shots; users shoot from below at an angle. Solution: added VNClassifyImageRequest with a confidence threshold of 0.6; when confidence is low, we prompt to reshoot with instructions. Accuracy rose to 84%.
How to Avoid Accuracy Loss During Format Conversion?
The key is to match preprocessing with the training pipeline. If the model was trained on ImageNet with normalization mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225], any deviation causes a 15–30% accuracy drop. Resize strictly as in training: if center_crop was used, do not use fit with padding. The model will see padding as part of the object and misclassify.
What Preprocessing Does an ML Model Need?
| Parameter | Requirement | Impact on Accuracy |
|---|---|---|
| Normalization | mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225] | Deviation reduces accuracy by 15–30% |
| Input size | 224×224 for most classifiers | Mismatch causes resize artifacts |
| Color order | RGB (BGR → swap channels) | Model may output random results |
| EXIF correction | Apply rotation from metadata | Ignoring yields up to 25% loss |
How We Build the Inference Pipeline
| Platform | Framework | Notes |
|---|---|---|
| iOS | Core ML + Vision | Automatic orientation correction and resize. For heavy models: computeUnits = .cpuAndNeuralEngine |
| Android (ML Kit) | ImageLabeler | Use InputImage.fromMediaImage with rotationDegrees from ImageProxy |
| Android (TFLite) | Task Library ImageClassifier | Handles normalization and resize if specified in metadata |
TFLite Task Library speeds integration by 2x compared to manual pipeline thanks to built-in preprocessing and model loading utilities.
Results arrive asynchronously in a callback — update UI on main thread: LiveData on Android, @MainActor on iOS. Inference takes 30–50 ms.
What's Included
- Requirement audit: source, platform, target accuracy, latency.
- Model and framework selection (Core ML / TFLite / ML Kit).
- Implementation of preprocessing pipeline considering format and orientation.
- Inference integration with asynchronous processing.
- Testing on real user data — at least 100 samples.
- Confidence threshold tuning.
- Documentation and commented code.
- Handover to CI/CD.
Timelines and Cost
Timelines: 1–2 weeks depending on model complexity and preprocessing readiness. Cost is calculated individually. We provide post-delivery support — bug fixes and consultations for up to 3 months free of charge.
Example Swift preprocessing for Core ML
let model = try VNCoreMLModel(for: EfficientNet().model) let request = VNCoreMLRequest(model: model) { request, error in guard let results = request.results as? [VNClassificationObservation] else { return } // handle results } let handler = VNImageRequestHandler(cgImage: cgImage, orientation: .up) try! handler.perform([request]) Contact us to evaluate your project — we'll analyze your case for free within a day. You can get a consultation through the form on our website. Order image recognition integration for your app today.







