Mobile OCR: Real-Time Text Recognition via Camera

TRUETECH is engaged in the development, support and maintenance of iOS, Android, PWA mobile applications. We have extensive experience and expertise in publishing mobile applications in popular markets like Google Play, App Store, Amazon, AppGallery and others.

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Mobile OCR: Real-Time Text Recognition via Camera
Medium
~2-3 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    858
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    743
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1159
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1034
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    968
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    562

A user points a camera at a faded receipt in dim light. Without image preprocessing, standard OCR libraries produce up to 50% errors. The solution is to apply contrast stretching (vImageContrastStretch on iOS, OpenCV on Android) and binarization at the capture stage. Only then feed the frame to VNRecognizeTextRequest or ML Kit Text Recognition. In our practice, this chain raises accuracy from 70% to 98-99% on printed text.

Over five years, we have delivered more than 20 OCR projects for retail (price tags and receipts), logistics (waybill tracking), and fintech (passport data verification). Every project had its own set of filters and post-processing steps. A common mistake is expecting the OCR library to produce perfect results out of the box. Without image preprocessing and result post-processing, accuracy drops to 50–70%. Therefore, the first stage of any project is collecting real document samples and testing on them.

Why Preprocessing Is Critical for Accurate OCR

Preprocessing is a key stage that determines final accuracy. In poor lighting or blur, we use:

  • Contrast enhancement via vImageContrastStretch (iOS) or OpenCV (Android).
  • Grayscale conversion followed by AdaptiveThreshold.
  • Sharpen CIFilter before feeding into OCR.

For handwritten text, standard frameworks show 40–60% accuracy. In such cases, custom models based on TensorFlow Lite help — this is a separate task requiring labeled data and training.

How Native OCR Frameworks Work

iOS: Vision + VNRecognizeTextRequest

Since iOS 13, the Vision framework can recognize text offline. VNRecognizeTextRequest supports two modes: .fast (approximate, instant) and .accurate (slower but 15% more accurate for complex fonts). For complex fonts, the .accurate mode gives a 15% accuracy boost.

func recognizeText(in image: UIImage) {
    guard let cgImage = image.cgImage else { return }

    let request = VNRecognizeTextRequest { [weak self] request, error in
        guard let observations = request.results as? [VNRecognizedTextObservation] else { return }
        let text = observations.compactMap { $0.topCandidates(1).first?.string }.joined(separator: "\n")
        DispatchQueue.main.async { self?.handleRecognized(text: text) }
    }

    request.recognitionLevel = .accurate
    request.usesLanguageCorrection = true
    request.recognitionLanguages = ["ru-RU", "en-US"] // order = priority

    let handler = VNImageRequestHandler(cgImage: cgImage, options: [: ])
    try? handler.perform([request])
}

usesLanguageCorrection helps with typos, but sometimes "corrects" abbreviations and article numbers — for technical documents it is better to disable it.

Android: ML Kit Text Recognition v2

com.google.mlkit:text-recognition supports Latin, Cyrillic, Chinese, Japanese, Korean via separate modules. It downloads on first use (~5 MB for Latin).

val recognizer = TextRecognition.getClient(
    TextRecognizerOptions.DEFAULT_OPTIONS // or RussianTextRecognizerOptions
)

val image = InputImage.fromBitmap(bitmap, 0)
recognizer.process(image)
    .addOnSuccessListener { visionText ->
        val fullText = visionText.textBlocks
            .joinToString("\n") { block -> block.text }
        handleRecognized(fullText)
    }
    .addOnFailureListener { e -> handleError(e) }

ML Kit also returns bounding boxes for each text block — useful for highlighting recognized areas in the UI.

Live Mode: Real-Time Text from Video Stream

For live-overlay (text highlighted directly in the video stream) on iOS we use AVCaptureSession + CMSampleBuffer:

// Delegate method AVCaptureVideoDataOutput
func captureOutput(_ output: AVCaptureOutput,
                   didOutput sampleBuffer: CMSampleBuffer,
                   from connection: AVCaptureConnection) {
    guard let pixelBuffer = CMSampleBufferGetImageBuffer(sampleBuffer) else { return }

    // Do not start a new request if previous is still running
    guard !isProcessing else { return }
    isProcessing = true

    let request = VNRecognizeTextRequest { [weak self] request, _ in
        defer { self?.isProcessing = false }
        // handle results...
    }
    request.recognitionLevel = .fast // speed matters for live

    try? VNImageRequestHandler(cvPixelBuffer: pixelBuffer, options: [: ]).perform([request])
}

The isProcessing flag is mandatory — without it at 30 FPS, the request queue piles up and memory grows until crash.

On Android — CameraX + ImageAnalysis.Analyzer. ML Kit is optimized to work directly with ImageProxy without converting to Bitmap.

Parameter Static Recognition Live Recognition
Processing speed 100-200 ms up to 30 ms per frame
Accuracy up to 99% up to 95% (due to speed compromise)
Battery consumption low medium (continuous processing)
Application document scanning, receipts pointing at business cards, license plates

Improving Recognition Quality in Difficult Conditions

Contrast enhancement and binarization are standard techniques. For specific documents (e.g., faded receipts), we add custom filters.

Common integration mistakes:

  • Forgetting the isProcessing flag in live mode → memory leak.
  • Leaving usesLanguageCorrection enabled for technical texts → corrupts abbreviations.
  • Not checking bounding boxes for overlap with UI → text overlays on interface.

Post-Processing: From Raw Text to Structured Data

Raw OCR output is a stream of strings. Most tasks require structuring:

  • Receipts: extract lines with prices via regex, parse final amount.
  • Business cards: NSDataDetector (iOS) or Patterns (Android) for phones, emails, addresses.
  • Passports/documents: MRZ zone read by ICAO 9303 standard, parsers available.
  • License plates: separate task — better to use specialized models (OpenALPR, PlateRecognizer API).

For poor-quality Cyrillic text, preprocessing helps: increase contrast via vImageContrastStretch, grayscale, Sharpen CIFilter before OCR.

Comparison of Native Frameworks

Parameter Vision (iOS) ML Kit (Android)
Modes .fast, .accurate base model
Languages up to 15 in one request modules: Lat, Cyr, Chi, Jap, Kor
Offline yes yes (model ~5 MB)
Accuracy on printed ~98% ~97%
Bounding boxes yes yes
Speed (full HD) 100-200 ms 80-150 ms
Click for more details on preprocessing techniques

For extreme conditions, we use adaptive thresholding (e.g., cv2.adaptiveThreshold with block size 11) and morphological operations to remove noise. These steps can improve accuracy by an additional 10-15%.

Deliverables Included in Development

When ordering this service, you receive:

  • Integration of Vision or ML Kit into your app.
  • Tuning recognition parameters for your document type.
  • Live mode with text highlighting on camera (optional).
  • Data post-processing: parsing receipts, business cards, numbers.
  • Integration and testing documentation.
  • Support during App Store / Google Play review.

Work Process

  1. Define use cases: document types, languages, need for live mode or only static photo.
  2. Implement image capture (camera + gallery), preprocessing.
  3. Integrate OCR: native Vision/ML Kit or cloud (Google Vision API, AWS Textract) if higher accuracy is needed for complex documents.
  4. Post-processing for the specific task: data structuring, regex, NER.
  5. Test on real samples in different lighting conditions.

Timeline Estimates

Basic static text recognition via native framework — 2-3 days (starting from $2,000). Live mode with overlay + data structuring for a specific document type — 1-2 weeks (approx $5,000–$8,000). Complex scenarios with custom models — from one month (budget $15,000+). Contact us for an exact estimate for your project.

Our guaranteed accuracy on printed text is 98% with preprocessing; we provide a certified integration with documented results. Over 5 years of experience ensures reliable project delivery.

Get a consultation for your project — we will select the optimal solution. Contact us to discuss your task and estimate the cost of OCR development for your app. Our solution supports multilingual OCR offline including Russian, English, Chinese, and Japanese.

Machine Learning in Mobile Apps: CoreML, TFLite, and On-Device Models

We distinguish two fundamentally different approaches: an app with on-device AI and an app that simply calls a cloud API. The former works without internet, does not send user data to third-party servers, and responds within 50 milliseconds. The latter depends on network latency and pricing plans. Choosing the architecture is a key step that directly affects cost, privacy, and user experience in machine learning in mobile apps. Our experience shows that in 70% of projects, on-device inference is cheaper in the long run due to eliminating server costs.

How to Choose Between CoreML and TFLite for On-Device Inference?

CoreML — Apple's native framework for running ML models on device. Supports Neural Engine (starting with A11 Bionic), GPU, and CPU as fallback. Models are converted to .mlmodel format via coremltools from PyTorch, ONNX, or TensorFlow. Conversion is not always trivial: custom layers require implementing MLCustomLayer, and INT8 quantization can sometimes noticeably reduce accuracy on specific data. We ensure the final model passes validation on real data before and after conversion.

TensorFlow Lite — cross-platform alternative for Android and Flutter. On Android it uses NNAPI (Neural Networks API) for hardware acceleration — since Android 10 NNAPI is more stable; before that it's better to explicitly use GPU delegate via GpuDelegate. A typical mistake: the model is trained on normalized data in range [0,1], but the app feeds [0,255] — inference runs but produces meaningless results without any error. We include an automatic input data validation module in the SDK.

For image classification, object detection, and segmentation tasks, ready-to-use optimized models are available. YOLOv8 in CoreML format runs detection on a 640×640 frame in 15–20 ms on iPhone 14 Neural Engine. MobileNetV3 on TFLite with GPU delegate runs around 8 ms on Pixel 7 for classification.

Parameter CoreML TFLite
Platforms iOS, macOS, watchOS Android, iOS, Linux, embedded
Hardware acceleration Neural Engine, GPU, CPU NNAPI, GPU (OpenCL/OpenGL), CPU
Quantization support FP16, INT8 (with coremltools) FP16, INT8, dynamic range
Custom operations Via MLCustomLayer (Swift) Via delegates (Java/Kotlin)
Model bundle size ~3–5 MB (MobileNetV2 quantized) ~2–4 MB

What If You Need Text Generation On-Device?

Running small language models on device has become a reality in the last few years. Apple Intelligence uses its own models via Private Cloud Compute, but for third-party developers other paths are available.

llama.cpp with Metal backend on iOS is a working approach for phi-3-mini (3.8B parameters, 4-bit quantization, ~2.3 GB). Inference: 15–25 tokens/second on iPhone 15 Pro. For integration in Swift, use the Swift Package llama.swift or a wrapper via C interface llama.h. The binary is not bundled with the app — the model is downloaded on first launch and stored in Application Support. Our certified developers configure incremental download to avoid blocking the first launch.

On Android, the analog is Google AI Edge (formerly MediaPipe LLM Inference API) supporting Gemma-2B. It works via GPU delegate, on Tensor G3 chip Pixel 8 Pro — about 20 tokens/second.

Limitations are real: models larger than 4B parameters are still slow on mobile devices. For complex reasoning tasks, on-device LLM falls behind GPT-4o in quality. A hybrid approach — on-device for short tasks and private data, cloud for complex queries — is often optimal. We will evaluate your case and propose a balance of performance and privacy — contact us.

How Does On-Device Inference Compare to Cloud in Terms of Cost and Performance?

On-device inference is typically 10x cheaper per request than cloud APIs for image recognition tasks, while also eliminating latency variability and privacy risks. The table below summarizes the trade-offs.

Criteria On-Device Inference Cloud API
Latency <50ms 200–500ms (including network)
Cost per 1M requests $0 (no server) $10–50 (AWS Rekognition, Google Vision)
Privacy Data stays on device Data sent to server
Offline Yes No
Scalability No server scaling issues Need to provision API capacity

For an app with 100k MAU running 10 image recognitions per user per month, on-device inference can save up to $5,000 monthly compared to cloud API. Get a free consultation on your ML architecture today.

Integrating OpenAI API and Other Cloud Models

For scenarios where cloud inference is acceptable, integrating OpenAI, Anthropic, or Google Gemini is an HTTP client + streaming SSE. In Swift, AsyncThrowingStream is convenient for streaming responses. In Kotlin, use Flow.

Critically: API keys must never be stored in the app bundle. Even an obfuscated key can be extracted from the IPA in 10 minutes using strings or frida. Correct architecture: mobile app → your own backend → OpenAI API. The backend controls rate limiting, logs requests, and protects the key.

What Is Included in the Work (Deliverables)

  • Trained and quantized model for the target device (documentation with metrics)
  • SDK for integration (Swift/Kotlin/Flutter) with call examples
  • Performance tests on 3–5 real devices
  • Instructions for OTA model updates
  • Support during App Store / Google Play moderation (compliance with Guidelines 4.2, 5.1)
  • 2 weeks of technical support after release

Typical Project Pipeline

  1. Task analysis — measure latency, privacy, size, supported devices.
  2. Model prototyping — in Python, evaluate accuracy on target data.
  3. Conversion and quantization — for CoreML/TFLite with validation.
  4. Integration into the app — model wrapped in a service layer (easy to swap CoreML ↔ TFLite ↔ cloud).
  5. Testing — on real devices, measure FPS, RAM, battery.
  6. Deployment — via TestFlight / Firebase App Distribution, monitor metrics.

Timelines: integration of a ready CoreML/TFLite model — 1–2 weeks, development of a custom model with mobile optimization — from 6 weeks, on-device LLM chat with personalization — 4–8 weeks.

Why We Take on Complex Cases?

10+ years of experience in mobile development, 50+ implemented AI/ML solutions, guarantee of compatibility with current iOS and Android versions. All projects undergo code review and load testing. The cost includes preparation of moderation documentation and training of your team.

Contact us — we will help you choose the architecture and implement ML in your app turnkey. Order an audit of your existing solution — we will assess the potential for server cost savings free of charge. In some projects, savings can reach significant amounts per month.