Integrate Core ML on iOS: Offline AI Without the Cloud

You train an image classification model on PyTorch—accuracy 95%, but how do you run it on iOS without latency and without sending data to the cloud? Core ML with on-device inference solves this. We integrate Core ML models into iOS apps for fully offline AI. Inference speed—single-digit milliseconds

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Integrate Core ML on iOS: Offline AI Without the Cloud
Complex
~1-2 weeks

Our competencies:

Frequently Asked Questions

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    896
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    782
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1216
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1079
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    1003
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    597

You train an image classification model on PyTorch—accuracy 95%, but how do you run it on iOS without latency and without sending data to the cloud? Core ML with on-device inference solves this. We integrate Core ML models into iOS apps for fully offline AI. Inference speed—single-digit milliseconds, data stays on the device, no network latency. Our team—7 years in mobile development, over 50 successful Core ML integrations. Savings on server inference can reach 60% (over 200,000 ₽ per year for an average project). We guarantee integration quality, confirmed by Apple Developer certifications.

How conversion works: from weights to .mlpackage

Most modern models arrive as PyTorch checkpoints or ONNX files. We convert via coremltools—Apple's official Python package:

import coremltools as ct import torch # Suppose we have a PyTorch image classification model model = MyModel() model.load_state_dict(torch.load("model.pth")) model.eval() # Tracing—need to pass example input example_input = torch.zeros(1, 3, 224, 224) traced = torch.jit.trace(model, example_input) # Conversion mlmodel = ct.convert( traced, inputs=[ct.ImageType( name="input_image", shape=(1, 3, 224, 224), color_layout=ct.colorlayout.RGB, bias=[-0.485/0.229, -0.456/0.224, -0.406/0.225], # ImageNet normalization scale=1/(255.0 * 0.229) # built into model, no need to do in Swift )], outputs=[ct.TensorType(name="class_probabilities")], compute_precision=ct.precision.FLOAT16, # for ANE minimum_deployment_target=ct.target.iOS16 ) mlmodel.save("MyClassifier.mlpackage") 

FLOAT16 + minimum_deployment_target=iOS16 activates the Apple Neural Engine. On iPhone 14, this is 4–8× faster than GPU for inference, with significantly lower battery consumption. According to Apple's Core ML documentation, ANE accelerates inference 4–8× compared to GPU. On older iOS versions, the same model runs via Metal GPU.

How to convert a model from PyTorch to Core ML?

Dynamic shapes—models with torch.Size([batch, seq_len, hidden]) where seq_len is not fixed break torch.jit.trace. Solution: ct.RangeDim for variable sizes or define multiple configurations via ct.EnumeratedShapes.

# Variable sequence length flexible_shape = ct.Shape(shape=(1, ct.RangeDim(1, 512), 768)) mlmodel = ct.convert(model, inputs=[ct.TensorType(shape=flexible_shape)]) 

Unsupported operations—for example, custom CUDA kernels. coremltools throws NotImplementedError. Path: either rewrite the operation using standard PyTorch primitives, or add a custom layer via C++/Swift extension.

Error Unsupported model format when loading .mlpackage on x86 simulator—the simulator uses CPU fallback, some FLOAT16 operations are not supported. Test accuracy only on a real device.

Loading and running on iOS

import CoreML import Vision // Load model (once at startup) let config = MLModelConfiguration() config.computeUnits = .all // ANE + GPU + CPU // .mlpackage loaded from bundle guard let modelURL = Bundle.main.url(forResource: "MyClassifier", withExtension: "mlpackage"), let model = try? MyClassifier(contentsOf: modelURL, configuration: config) else { fatalError("Failed to load model") } // Inference—on background thread DispatchQueue.global(qos: .userInitiated).async { do { let input = MyClassifierInput(input_image: cgImage) let output = try model.prediction(input: input) let probs = output.class_probabilities // probs — MLMultiArray, get value: probs[0].doubleValue } catch { print("Inference error: \(error)") } } 

Model loading takes ~100–300 ms (depends on size). Do not load it in viewDidLoad—load once at app startup or first use, keep in memory while needed.

Why on-device ML is faster and safer than cloud?

On-device ML eliminates network latency, preserves user data privacy, and works offline. You don't pay for server inference and don't depend on internet connection. For tasks where response speed is critical (e.g., real-time video processing), device is the only sensible option.

Criterion Core ML Cloud AI
Latency <10 ms 100–500 ms
Privacy Data on device Sent to server
Offline Yes No
Cost No inference cost Pay per API call

Performance on real devices:

Device Model computeUnits Inference time
iPhone 14 Pro MobileNetV3 (5 MB FP16) .all (ANE) 2–4 ms
iPhone 14 Pro ResNet-50 (48 MB FP16) .all (ANE) 8–15 ms
iPhone 12 BERT-base (350 MB FP16) .all 180–250 ms
iPhone SE 2nd gen MobileNetV3 (5 MB FP16) .cpuOnly 12–20 ms

For profiling, use Xcode Instruments → Core ML Instrument.

Vision Framework as a wrapper

For computer vision tasks, VNCoreMLRequest is more convenient—Vision handles input resizing, image orientation, coordinate transformations:

let coreMLModel = try VNCoreMLModel(for: model.model) // .model — MLModel from generated class let request = VNCoreMLRequest(model: coreMLModel) { request, error in guard let results = request.results as? [VNClassificationObservation] else { return } let topResult = results.sorted { $0.confidence > $1.confidence }.first print("\(topResult?.identifier ?? "?") — \(topResult?.confidence ?? 0)") } request.imageCropAndScaleOption = .centerCrop // or .scaleFit let handler = VNImageRequestHandler(cgImage: inputCGImage, options: [:]) try handler.perform([request]) 

VNCoreMLRequest automatically solves the input size mismatch problem—you pass an arbitrary image, Vision resizes it to the model's expected size. Without Vision, you'd have to do this manually via vImage or CIImage.

What's included in the work

  • Documentation on conversion and integration, including description of all steps and tools used.
  • Access to the repository with conversion source code and integration examples.
  • Training of the client's team on working with Core ML, profiling, and model updates.
  • Support for one month after integration to resolve any issues.

How to update a model without updating the app?

Core ML supports loading a model from an arbitrary URL, not only from bundle. This allows updating the model via server:

// Load mlpackage from documents directory let documentsURL = FileManager.default.urls(for: .documentDirectory, in: .userDomainMask)[0] let downloadedModelURL = documentsURL.appendingPathComponent("updated_model.mlpackage") if FileManager.default.fileExists(atPath: downloadedModelURL.path) { let model = try MyClassifier(contentsOf: downloadedModelURL, configuration: config) } else { // Fallback to bundle } 

Download model over network via URLSession, save to Documents, verify via SHA-256 hash before use.

Our approach: from analysis to deployment

  • Analysis of the original model (framework, weights, structure).
  • Conversion to Core ML with optimization of precision and compute units.
  • Optimization for target devices (profiling on real devices).
  • Integration into the app: loading, caching, fallback on errors.
  • Setting up remote model update (optional).
  • Documentation and team training.

Process stages:

  1. Analytics—receive weights, assess complexity, choose conversion strategy.
  2. Conversion—create .mlpackage, resolve operation and dimension issues.
  3. Profiling—measure speed and energy consumption on several iPhone generations.
  4. Integration—embed into SwiftUI/UIKit, add error handling.
  5. Deployment—publish via App Store, set up remote update.

Estimated timelines

Conversion of an existing model + basic iOS integration—1–2 weeks. Complex model with non-standard operations, multiple inputs/outputs, remote update—3–5 weeks. Cost is calculated individually for each project. We'll assess the task in one business day—just send the weights and a description of the task.

Contact us for a consultation on Core ML integration into your project. Request a preliminary analysis of your model—we'll find the optimal solution.

Learn more about conversion in the coremltools documentation.