Imagine a mobile app that must process images, text, or sound without network access. Server latency is unacceptable, data privacy is critical. Every extra megabyte of traffic costs the user. On-device ML solves these issues, and ONNX Runtime is the key tool for cross-platform deployment. We integrate it so that the model runs equally fast on iOS and Android. Switching to on-device can reduce server infrastructure costs by up to 90%, saving hundreds of thousands of rubles monthly under high loads.
ONNX Runtime Mobile appeals with one argument: one model, both platforms. Convert PyTorch or TensorFlow to ONNX, add onnxruntime-android and onnxruntime-objc, run the same .onnx file. In practice, the difference in execution providers between iOS and Android still requires platform-specific code, but the model itself is unified. Our experience: over 5 years in mobile ML, dozens of on-device inference projects. Contact us for an assessment of your model and a preliminary quote.
How to Prepare a Model for Mobile
Standard ONNX export from PyTorch:
import torch
import onnx
from onnxsim import simplify # onnx-simplifier for graph optimization
model = MyModel(); model.eval()
dummy = torch.zeros(1, 3, 224, 224)
torch.onnx.export(
model, dummy, "model.onnx",
opset_version=17,
input_names=["input"],
output_names=["output"],
dynamic_axes={"input": {0: "batch_size"}, "output": {0: "batch_size"}}
)
# Simplify graph — removes redundant reshapes, transposes, makes graph cleaner
model_onnx = onnx.load("model.onnx")
model_simplified, check = simplify(model_onnx)
onnx.save(model_simplified, "model_simplified.onnx")
Additionally for mobile — quantization via onnxruntime.quantization:
from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic(
"model_simplified.onnx",
"model_int8.onnx",
weight_type=QuantType.QInt8
)
# Model size reduces ~4× compared to FP32
| Quantization Type | Model Size (FP32 → Int8) | Accuracy Loss | Speed on CPU |
|---|---|---|---|
| Dynamic | ~75% smaller | <1% | ~30% faster |
| Static (calibration) | ~75% smaller | 0.5-2% | ~40% faster |
Which Execution Provider Delivers Maximum Performance?
Android: NNAPI vs XNNPACK
On Android, the choice of Execution Provider depends on hardware. NNAPI delegates operations to NPU/DSP, providing up to 2x acceleration on supported ops. XNNPACK is an optimized CPU backend using SIMD instructions, speeding up to 2x on CPU but without NPU access. On a project with object detection on MediaTek Dimensity, we got 45 ms on NNAPI vs 80 ms on XNNPACK. We recommend using NNAPI for devices with NPU, XNNPACK as fallback.
iOS: CoreML Execution Provider
appendCoreMLExecutionProvider on iOS 13+ delegates supported operations to Core ML, gaining access to ANE. Operations not supported by Core ML automatically run on CPU. In tests on iPhone 12, we got 35% speedup over CPU on ResNet-50. CoreML EP is convenient for fast cross-platform deployment, but for maximum performance, consider native Core ML.
When is ONNX Runtime Better Than Native Formats?
Use ONNX Runtime for prototyping, cross-platform projects, models with custom ops that coremltools cannot convert, and frequent model updates without rebuilding the conversion pipeline. If you need maximum performance on a single platform, choose the native format: on iOS — Core ML with full ANE acceleration (usually 20–40% faster than ORT+CoreML EP), on Android — TFLite + GPU Delegate (sometimes faster than ORT+NNAPI). For single-platform deployment and critical performance, native is preferable.
How to Integrate ONNX Runtime on Android and iOS
Android: Setup and Inference
// build.gradle
implementation("com.microsoft.onnxruntime:onnxruntime-android:1.18.0")
// Create session
val sessionOptions = OrtSession.SessionOptions().apply {
// NNAPI Execution Provider for Android NPU/DSP
addNnapi(NNAPIFlags.USE_FP16) // FP16 mode in NNAPI
// Or: addXnnpack(mapOf()) for XNNPACK (CPU SIMD)
setOptimizationLevel(OrtSession.SessionOptions.OptLevel.ALL_OPT)
setIntraOpNumThreads(4)
}
val env = OrtEnvironment.getEnvironment()
val session = env.createSession(
context.assets.open("model_simplified.onnx").readBytes(),
sessionOptions
)
// Inference
val inputTensor = OnnxTensor.createTensor(
env,
FloatBuffer.wrap(preprocessedArray),
longArrayOf(1, 3, 224, 224)
)
val results = session.run(mapOf("input" to inputTensor))
val outputArray = (results["output"]?.value as Array<FloatArray>)[0]
// Resource cleanup — mandatory
inputTensor.close()
results.close()
Leaks from unclosed OnnxTensor and OrtSession.Result are common. In Kotlin, use use {} block: results.use { ... }.
iOS: ObjC/Swift Integration
// Package.swift or Podfile: pod 'onnxruntime-objc'
import onnxruntime_objc
// Setup
let env = try ORTEnv(loggingLevel: ORTLoggingLevel.warning)
let options = try ORTSessionOptions()
try options.setIntraOpNumThreads(4)
// On iOS — CoreML Execution Provider
try options.appendCoreMLExecutionProvider(withFlags: [.enableOnSubgraphs])
let session = try ORTSession(
env: env,
modelPath: Bundle.main.path(forResource: "model_simplified", ofType: "onnx")!,
sessionOptions: options
)
// Prepare input
let inputShape: [NSNumber] = [1, 3, 224, 224]
let inputData = Data(bytes: preprocessedFloats, count: preprocessedFloats.count * MemoryLayout<Float>.size)
let inputTensor = try ORTValue(
tensorData: NSMutableData(data: inputData),
elementType: .float,
shape: inputShape
)
let outputs = try session.run(
withInputs: ["input": inputTensor],
outputNames: ["output"],
runOptions: nil
)
let outputTensor = outputs["output"]!
let outputData = try outputTensor.tensorData() as Data
let floats = outputData.withUnsafeBytes { Array($0.bindMemory(to: Float.self)) }
Why Quantization is Critical for Mobile Inference
Quantization reduces model size by 4× (50 MB → 12 MB), lowers memory consumption, and speeds up CPU inference by 30-40%. Dynamic quantization does not require calibration data but yields slightly less speed gain than static. In practice, we use static quantization with a representative dataset — it gives stable improvement without significant accuracy loss (0.5-2%).
What If an Operation is Not Supported?
# Check which ops NNAPI Execution Provider supports
python -m onnxruntime.tools.check_nnapi_supported_ops --model model.onnx
# If an op is not supported — it runs on CPU (fallback)
# This is not a crash but can nullify all NNAPI acceleration
To identify bottlenecks, use the ORT Profiling API. It records per-operator timing. Enable via options.enableProfiling("ort_profile") — generates JSON viewable in Chrome chrome://tracing. Profiling on target devices helps choose the optimal execution provider. For example, on one project we switched from NNAPI to XNNPACK for a model with 80% unsupported ops, reducing inference from 300 ms to 120 ms.
What Our Work Includes
- Export and simplify ONNX graph, quantize to Int8.
- Integrate ONNX Runtime on iOS and Android with optimal Execution Providers.
- Profile performance on a fleet of 10+ real devices, including older ones.
- Compare with native formats (Core ML, TFLite) and recommend the best solution.
- Documentation for building and updating the model, integration source code.
- Guarantee stable operation and lock in inference time.
Our Experience in Mobile ML
Over 5 years deploying on-device ML in commercial applications — from retail to healthcare. Completed 20+ projects with ONNX Runtime, Core ML, and TFLite. Our engineers hold Apple and Google certifications. We guarantee the model will work on all stated devices. Get a consultation on ONNX Runtime integration — we'll assess your project and propose the optimal turnkey solution. Order ONNX Runtime integration for your app — let's discuss the details.
Timeline Estimates
Basic cross-platform ONNX Runtime integration: 2–3 weeks. With EP optimization, profiling, testing on a device fleet: 4–6 weeks. Cost calculated individually after analyzing the model and performance requirements.







