You train an image classification model on PyTorch—accuracy 95%, but how do you run it on iOS without latency and without sending data to the cloud? Core ML with on-device inference solves this. We integrate Core ML models into iOS apps for fully offline AI. Inference speed—single-digit milliseconds, data stays on the device, no network latency. Our team—7 years in mobile development, over 50 successful Core ML integrations. Savings on server inference can reach 60% (over 200,000 ₽ per year for an average project). We guarantee integration quality, confirmed by Apple Developer certifications.
How conversion works: from weights to .mlpackage
Most modern models arrive as PyTorch checkpoints or ONNX files. We convert via coremltools—Apple's official Python package:
import coremltools as ct
import torch
# Suppose we have a PyTorch image classification model
model = MyModel()
model.load_state_dict(torch.load("model.pth"))
model.eval()
# Tracing—need to pass example input
example_input = torch.zeros(1, 3, 224, 224)
traced = torch.jit.trace(model, example_input)
# Conversion
mlmodel = ct.convert(
traced,
inputs=[ct.ImageType(
name="input_image",
shape=(1, 3, 224, 224),
color_layout=ct.colorlayout.RGB,
bias=[-0.485/0.229, -0.456/0.224, -0.406/0.225], # ImageNet normalization
scale=1/(255.0 * 0.229) # built into model, no need to do in Swift
)],
outputs=[ct.TensorType(name="class_probabilities")],
compute_precision=ct.precision.FLOAT16, # for ANE
minimum_deployment_target=ct.target.iOS16
)
mlmodel.save("MyClassifier.mlpackage")
FLOAT16 + minimum_deployment_target=iOS16 activates the Apple Neural Engine. On iPhone 14, this is 4–8× faster than GPU for inference, with significantly lower battery consumption. According to Apple's Core ML documentation, ANE accelerates inference 4–8× compared to GPU. On older iOS versions, the same model runs via Metal GPU.
How to convert a model from PyTorch to Core ML?
Dynamic shapes—models with torch.Size([batch, seq_len, hidden]) where seq_len is not fixed break torch.jit.trace. Solution: ct.RangeDim for variable sizes or define multiple configurations via ct.EnumeratedShapes.
# Variable sequence length
flexible_shape = ct.Shape(shape=(1, ct.RangeDim(1, 512), 768))
mlmodel = ct.convert(model, inputs=[ct.TensorType(shape=flexible_shape)])
Unsupported operations—for example, custom CUDA kernels. coremltools throws NotImplementedError. Path: either rewrite the operation using standard PyTorch primitives, or add a custom layer via C++/Swift extension.
Error Unsupported model format when loading .mlpackage on x86 simulator—the simulator uses CPU fallback, some FLOAT16 operations are not supported. Test accuracy only on a real device.
Loading and running on iOS
import CoreML
import Vision
// Load model (once at startup)
let config = MLModelConfiguration()
config.computeUnits = .all // ANE + GPU + CPU
// .mlpackage loaded from bundle
guard let modelURL = Bundle.main.url(forResource: "MyClassifier", withExtension: "mlpackage"),
let model = try? MyClassifier(contentsOf: modelURL, configuration: config) else {
fatalError("Failed to load model")
}
// Inference—on background thread
DispatchQueue.global(qos: .userInitiated).async {
do {
let input = MyClassifierInput(input_image: cgImage)
let output = try model.prediction(input: input)
let probs = output.class_probabilities
// probs — MLMultiArray, get value: probs[0].doubleValue
} catch {
print("Inference error: \(error)")
}
}
Model loading takes ~100–300 ms (depends on size). Do not load it in viewDidLoad—load once at app startup or first use, keep in memory while needed.
Why on-device ML is faster and safer than cloud?
On-device ML eliminates network latency, preserves user data privacy, and works offline. You don't pay for server inference and don't depend on internet connection. For tasks where response speed is critical (e.g., real-time video processing), device is the only sensible option.
| Criterion | Core ML | Cloud AI |
|---|---|---|
| Latency | <10 ms | 100–500 ms |
| Privacy | Data on device | Sent to server |
| Offline | Yes | No |
| Cost | No inference cost | Pay per API call |
Performance on real devices:
| Device | Model | computeUnits | Inference time |
|---|---|---|---|
| iPhone 14 Pro | MobileNetV3 (5 MB FP16) | .all (ANE) | 2–4 ms |
| iPhone 14 Pro | ResNet-50 (48 MB FP16) | .all (ANE) | 8–15 ms |
| iPhone 12 | BERT-base (350 MB FP16) | .all | 180–250 ms |
| iPhone SE 2nd gen | MobileNetV3 (5 MB FP16) | .cpuOnly | 12–20 ms |
For profiling, use Xcode Instruments → Core ML Instrument.
Vision Framework as a wrapper
For computer vision tasks, VNCoreMLRequest is more convenient—Vision handles input resizing, image orientation, coordinate transformations:
let coreMLModel = try VNCoreMLModel(for: model.model) // .model — MLModel from generated class
let request = VNCoreMLRequest(model: coreMLModel) { request, error in
guard let results = request.results as? [VNClassificationObservation] else { return }
let topResult = results.sorted { $0.confidence > $1.confidence }.first
print("\(topResult?.identifier ?? "?") — \(topResult?.confidence ?? 0)")
}
request.imageCropAndScaleOption = .centerCrop // or .scaleFit
let handler = VNImageRequestHandler(cgImage: inputCGImage, options: [:])
try handler.perform([request])
VNCoreMLRequest automatically solves the input size mismatch problem—you pass an arbitrary image, Vision resizes it to the model's expected size. Without Vision, you'd have to do this manually via vImage or CIImage.
What's included in the work
- Documentation on conversion and integration, including description of all steps and tools used.
- Access to the repository with conversion source code and integration examples.
- Training of the client's team on working with Core ML, profiling, and model updates.
- Support for one month after integration to resolve any issues.
How to update a model without updating the app?
Core ML supports loading a model from an arbitrary URL, not only from bundle. This allows updating the model via server:
// Load mlpackage from documents directory
let documentsURL = FileManager.default.urls(for: .documentDirectory, in: .userDomainMask)[0]
let downloadedModelURL = documentsURL.appendingPathComponent("updated_model.mlpackage")
if FileManager.default.fileExists(atPath: downloadedModelURL.path) {
let model = try MyClassifier(contentsOf: downloadedModelURL, configuration: config)
} else {
// Fallback to bundle
}
Download model over network via URLSession, save to Documents, verify via SHA-256 hash before use.
Our approach: from analysis to deployment
- Analysis of the original model (framework, weights, structure).
- Conversion to Core ML with optimization of precision and compute units.
- Optimization for target devices (profiling on real devices).
- Integration into the app: loading, caching, fallback on errors.
- Setting up remote model update (optional).
- Documentation and team training.
Process stages:
- Analytics—receive weights, assess complexity, choose conversion strategy.
- Conversion—create .mlpackage, resolve operation and dimension issues.
- Profiling—measure speed and energy consumption on several iPhone generations.
- Integration—embed into SwiftUI/UIKit, add error handling.
- Deployment—publish via App Store, set up remote update.
Estimated timelines
Conversion of an existing model + basic iOS integration—1–2 weeks. Complex model with non-standard operations, multiple inputs/outputs, remote update—3–5 weeks. Cost is calculated individually for each project. We'll assess the task in one business day—just send the weights and a description of the task.
Contact us for a consultation on Core ML integration into your project. Request a preliminary analysis of your model—we'll find the optimal solution.
Learn more about conversion in the coremltools documentation.







