On-Device ML Integration with ONNX Runtime for Mobile Apps

Imagine a mobile app that must process images, text, or sound without network access. Server latency is unacceptable, data privacy is critical. Every extra megabyte of traffic costs the user. On-device ML solves these issues, and [ONNX Runtime](https://github.com/microsoft/onnxruntime) is the key to

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
On-Device ML Integration with ONNX Runtime for Mobile Apps
Complex
~1-2 weeks

Our competencies:

Frequently Asked Questions

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    896
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    782
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1216
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1079
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    1003
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    597

Imagine a mobile app that must process images, text, or sound without network access. Server latency is unacceptable, data privacy is critical. Every extra megabyte of traffic costs the user. On-device ML solves these issues, and ONNX Runtime is the key tool for cross-platform deployment. We integrate it so that the model runs equally fast on iOS and Android. Switching to on-device can reduce server infrastructure costs by up to 90%, saving hundreds of thousands of rubles monthly under high loads.

ONNX Runtime Mobile appeals with one argument: one model, both platforms. Convert PyTorch or TensorFlow to ONNX, add onnxruntime-android and onnxruntime-objc, run the same .onnx file. In practice, the difference in execution providers between iOS and Android still requires platform-specific code, but the model itself is unified. Our experience: over 5 years in mobile ML, dozens of on-device inference projects. Contact us for an assessment of your model and a preliminary quote.

How to Prepare a Model for Mobile

Standard ONNX export from PyTorch:

import torch import onnx from onnxsim import simplify # onnx-simplifier for graph optimization model = MyModel(); model.eval() dummy = torch.zeros(1, 3, 224, 224) torch.onnx.export( model, dummy, "model.onnx", opset_version=17, input_names=["input"], output_names=["output"], dynamic_axes={"input": {0: "batch_size"}, "output": {0: "batch_size"}} ) # Simplify graph — removes redundant reshapes, transposes, makes graph cleaner model_onnx = onnx.load("model.onnx") model_simplified, check = simplify(model_onnx) onnx.save(model_simplified, "model_simplified.onnx") 

Additionally for mobile — quantization via onnxruntime.quantization:

from onnxruntime.quantization import quantize_dynamic, QuantType quantize_dynamic( "model_simplified.onnx", "model_int8.onnx", weight_type=QuantType.QInt8 ) # Model size reduces ~4× compared to FP32 
Quantization Type Model Size (FP32 → Int8) Accuracy Loss Speed on CPU
Dynamic ~75% smaller <1% ~30% faster
Static (calibration) ~75% smaller 0.5-2% ~40% faster

Which Execution Provider Delivers Maximum Performance?

Android: NNAPI vs XNNPACK

On Android, the choice of Execution Provider depends on hardware. NNAPI delegates operations to NPU/DSP, providing up to 2x acceleration on supported ops. XNNPACK is an optimized CPU backend using SIMD instructions, speeding up to 2x on CPU but without NPU access. On a project with object detection on MediaTek Dimensity, we got 45 ms on NNAPI vs 80 ms on XNNPACK. We recommend using NNAPI for devices with NPU, XNNPACK as fallback.

iOS: CoreML Execution Provider

appendCoreMLExecutionProvider on iOS 13+ delegates supported operations to Core ML, gaining access to ANE. Operations not supported by Core ML automatically run on CPU. In tests on iPhone 12, we got 35% speedup over CPU on ResNet-50. CoreML EP is convenient for fast cross-platform deployment, but for maximum performance, consider native Core ML.

When is ONNX Runtime Better Than Native Formats?

Use ONNX Runtime for prototyping, cross-platform projects, models with custom ops that coremltools cannot convert, and frequent model updates without rebuilding the conversion pipeline. If you need maximum performance on a single platform, choose the native format: on iOS — Core ML with full ANE acceleration (usually 20–40% faster than ORT+CoreML EP), on Android — TFLite + GPU Delegate (sometimes faster than ORT+NNAPI). For single-platform deployment and critical performance, native is preferable.

How to Integrate ONNX Runtime on Android and iOS

Android: Setup and Inference

// build.gradle implementation("com.microsoft.onnxruntime:onnxruntime-android:1.18.0") // Create session val sessionOptions = OrtSession.SessionOptions().apply { // NNAPI Execution Provider for Android NPU/DSP addNnapi(NNAPIFlags.USE_FP16) // FP16 mode in NNAPI // Or: addXnnpack(mapOf()) for XNNPACK (CPU SIMD) setOptimizationLevel(OrtSession.SessionOptions.OptLevel.ALL_OPT) setIntraOpNumThreads(4) } val env = OrtEnvironment.getEnvironment() val session = env.createSession( context.assets.open("model_simplified.onnx").readBytes(), sessionOptions ) // Inference val inputTensor = OnnxTensor.createTensor( env, FloatBuffer.wrap(preprocessedArray), longArrayOf(1, 3, 224, 224) ) val results = session.run(mapOf("input" to inputTensor)) val outputArray = (results["output"]?.value as Array<FloatArray>)[0] // Resource cleanup — mandatory inputTensor.close() results.close() 

Leaks from unclosed OnnxTensor and OrtSession.Result are common. In Kotlin, use use {} block: results.use { ... }.

iOS: ObjC/Swift Integration

// Package.swift or Podfile: pod 'onnxruntime-objc' import onnxruntime_objc // Setup let env = try ORTEnv(loggingLevel: ORTLoggingLevel.warning) let options = try ORTSessionOptions() try options.setIntraOpNumThreads(4) // On iOS — CoreML Execution Provider try options.appendCoreMLExecutionProvider(withFlags: [.enableOnSubgraphs]) let session = try ORTSession( env: env, modelPath: Bundle.main.path(forResource: "model_simplified", ofType: "onnx")!, sessionOptions: options ) // Prepare input let inputShape: [NSNumber] = [1, 3, 224, 224] let inputData = Data(bytes: preprocessedFloats, count: preprocessedFloats.count * MemoryLayout<Float>.size) let inputTensor = try ORTValue( tensorData: NSMutableData(data: inputData), elementType: .float, shape: inputShape ) let outputs = try session.run( withInputs: ["input": inputTensor], outputNames: ["output"], runOptions: nil ) let outputTensor = outputs["output"]! let outputData = try outputTensor.tensorData() as Data let floats = outputData.withUnsafeBytes { Array($0.bindMemory(to: Float.self)) } 

Why Quantization is Critical for Mobile Inference

Quantization reduces model size by 4× (50 MB → 12 MB), lowers memory consumption, and speeds up CPU inference by 30-40%. Dynamic quantization does not require calibration data but yields slightly less speed gain than static. In practice, we use static quantization with a representative dataset — it gives stable improvement without significant accuracy loss (0.5-2%).

What If an Operation is Not Supported?

# Check which ops NNAPI Execution Provider supports python -m onnxruntime.tools.check_nnapi_supported_ops --model model.onnx # If an op is not supported — it runs on CPU (fallback) # This is not a crash but can nullify all NNAPI acceleration 

To identify bottlenecks, use the ORT Profiling API. It records per-operator timing. Enable via options.enableProfiling("ort_profile") — generates JSON viewable in Chrome chrome://tracing. Profiling on target devices helps choose the optimal execution provider. For example, on one project we switched from NNAPI to XNNPACK for a model with 80% unsupported ops, reducing inference from 300 ms to 120 ms.

What Our Work Includes

  • Export and simplify ONNX graph, quantize to Int8.
  • Integrate ONNX Runtime on iOS and Android with optimal Execution Providers.
  • Profile performance on a fleet of 10+ real devices, including older ones.
  • Compare with native formats (Core ML, TFLite) and recommend the best solution.
  • Documentation for building and updating the model, integration source code.
  • Guarantee stable operation and lock in inference time.

Our Experience in Mobile ML

Over 5 years deploying on-device ML in commercial applications — from retail to healthcare. Completed 20+ projects with ONNX Runtime, Core ML, and TFLite. Our engineers hold Apple and Google certifications. We guarantee the model will work on all stated devices. Get a consultation on ONNX Runtime integration — we'll assess your project and propose the optimal turnkey solution. Order ONNX Runtime integration for your app — let's discuss the details.

Timeline Estimates

Basic cross-platform ONNX Runtime integration: 2–3 weeks. With EP optimization, profiling, testing on a device fleet: 4–6 weeks. Cost calculated individually after analyzing the model and performance requirements.