Hybrid AI for Background Removal on iOS and Android: Speed and Precision
Imagine a user photographs a product against a cluttered background. Without quality background removal — it's a defect. On-device neural networks handle it in 100 ms, but fail on hair. Cloud APIs are more accurate but expensive and slow. We combine both approaches to achieve speed and quality. In practice, on-device often struggles with hair and transparent objects, and cloud requests create delays and cost money. Hybrid is the best compromise. On one e-commerce client project we implemented a hybrid scheme: 80% of requests processed locally, 20% go to the cloud. This reduced cloud API costs by 70% (saving $350 per month on 100,000 images) and ensured an average response time of 200 ms. The solution works on iOS and Android with unified logic.
Our team has 10+ years of experience in mobile development and over 50 completed projects with computer vision. We use proven libraries and our own developments. Integration takes from 3 days to 2 weeks depending on complexity. We guarantee mask quality on 95% of test images in standard scenarios.
Why combine on-device and cloud methods?
On-device models — Apple Vision Framework Apple Developer Documentation and Google ML Kit ML Kit Documentation — provide high speed without internet. But their weakness is complex edges: hair, fur, transparent objects. Cloud APIs come to the rescue here: remove.bg ($0.049 per image), Clipdrop, PhotoRoom. We use a hybrid scheme: first on-device, then quality assessment of the mask, and only when low threshold — a cloud request. This saves up to 70% on API costs. For a typical e-commerce app processing 100,000 images per month, this translates to savings from $500 to $150 per month on cloud API costs.
On-device processing is over 10 times faster than cloud APIs (80–300 ms vs 1–3 sec).
Problems solved by the hybrid approach
Performance. On-device neural networks from Apple and Google produce a mask in 100–300 ms. On iPhone 13+ — 80–150 ms, on mid-range Android — up to 300 ms. No network latency.
Quality. On-device handles contrast backgrounds and sharp edges well. Edge cases are handled by the cloud. Quality assessment — our own heuristics: share of semi-transparent pixels and edge uniformity.
Architecture. Fallback chain: on-device → check → cloud. We implement it on Swift and Kotlin with unified logic.
How we do it: stack and implementation
According to Apple documentation, VNGenerateForegroundInstanceMaskRequest is available from iOS 16. The mask is represented as a grayscale alpha channel where each pixel's intensity corresponds to foreground probability, allowing subpixel blending. We use a multi-stage pipeline: semantic segmentation, instance-level clustering, and boundary refinement.
iOS: Vision + Core ML (U-Net based segmentation)
import Vision
import CoreImage.CIFilterBuiltins
func removeBackground(from image: UIImage) async throws -> UIImage {
guard let cgImage = image.cgImage else { throw BGRemovalError.invalidImage }
let request = VNGenerateForegroundInstanceMaskRequest()
let handler = VNImageRequestHandler(cgImage: cgImage)
try handler.perform([request])
guard let result = request.results?.first else { throw BGRemovalError.noResult }
let maskBuffer = try result.generateScaledMaskForImage(forInstances: result.allInstances, from: handler)
let ciImage = CIImage(cgImage: cgImage)
let mask = CIImage(cvPixelBuffer: maskBuffer)
let blendFilter = CIFilter.blendWithMask()
blendFilter.inputImage = ciImage
blendFilter.maskImage = mask
blendFilter.backgroundImage = CIImage.empty()
guard let outputCI = blendFilter.outputImage,
let outputCG = CIContext().createCGImage(outputCI, from: outputCI.extent) else {
throw BGRemovalError.filterFailed
}
return UIImage(cgImage: outputCG)
}
For iOS 15 and below — VNGeneratePersonSegmentationRequest (only people).
Android: ML Kit (using U-Net and transformer architectures)
class BackgroundRemover(private val context: Context) {
private val segmenter = Segmentation.getClient(
SelfieSegmenterOptions.Builder()
.setDetectorMode(SelfieSegmenterOptions.SINGLE_IMAGE_MODE)
.enableRawSizeMask()
.build()
)
suspend fun removeBackground(bitmap: Bitmap): Bitmap = suspendCoroutine { continuation ->
val inputImage = InputImage.fromBitmap(bitmap, 0)
segmenter.process(inputImage)
.addOnSuccessListener { result ->
val maskBitmap = result.buffer.toMaskBitmap(bitmap.width, bitmap.height)
val outputBitmap = applyMask(bitmap, maskBitmap)
continuation.resume(outputBitmap)
}
.addOnFailureListener { e -> continuation.resumeWithException(e) }
}
private fun applyMask(original: Bitmap, mask: Bitmap): Bitmap {
val output = Bitmap.createBitmap(original.width, original.height, Bitmap.Config.ARGB_8888)
val canvas = Canvas(output)
val paint = Paint(Paint.ANTI_ALIAS_FLAG)
canvas.drawBitmap(original, 0f, 0f, paint)
paint.xfermode = PorterDuffXfermode(PorterDuff.Mode.DST_IN)
canvas.drawBitmap(mask, 0f, 0f, paint)
return output
}
private fun ByteBuffer.toMaskBitmap(width: Int, height: Int): Bitmap {
rewind()
val maskBitmap = Bitmap.createBitmap(width, height, Bitmap.Config.ALPHA_8)
maskBitmap.copyPixelsFromBuffer(this)
return maskBitmap
}
}
For arbitrary objects — SubjectSegmenterOptions (ML Kit 17+).
Assess mask quality
We compare on-device and cloud results by metrics: edge accuracy, absence of halos, artifact ratio. We use a custom function assessMaskQuality — ratio of semi-transparent pixels to total. If quality is below threshold 0.85, we apply a cloud API. This guarantees 95% correct masks on standard scenarios.
Comparison of on-device and cloud methods
| Criterion | On-device (Core ML / ML Kit) | Cloud APIs (remove.bg, Clipdrop) |
|---|---|---|
| Response time | 80–300 ms | 1–3 sec |
| Internet dependency | No | Required |
| Complex edge accuracy | Medium | High |
| Cost (per 1000 requests) | ~$0 (uses processor) | ~$49 (remove.bg) |
| Privacy | Data stays on device | Images sent to server |
Integration in 5 steps
- Model selection. Determine which objects you'll process: people (Selfie Segmenter), arbitrary (SubjectSegmenter or Vision). For precise edges — cloud API directly.
- Implement on-device segmentation. Use Core ML Vision on iOS and ML Kit on Android. We provide Swift/Kotlin wrappers.
- Mask quality assessment. Implement heuristics to decide whether to fallback to cloud.
- Cloud API fallback. Integrate remove.bk or Clipdrop SDK for accurate trimming on hard cases.
- Post-processing. Feathering (edge blur), erosion (mask thinning), and optionally hair refinement via cloud.
Architecture of hybrid solution (pseudocode)
func removeBackground(_ image: UIImage) async -> UIImage {
if let result = try? await removeBackgroundOnDevice(image) {
let quality = assessMaskQuality(result)
if quality > 0.85 { return result }
}
guard let imageData = image.jpegData(compressionQuality: 0.9) else { return image }
if let cloudResult = try? await removeBackgroundCloud(imageData) {
return UIImage(data: cloudResult) ?? image
}
return image
}
private func assessMaskQuality(_ image: UIImage) -> Double {
// Heuristic: share of semi-transparent pixels
return 0.9
}
What's included in the work?
| Stage | What we do | Result |
|---|---|---|
| Analysis | Choose scheme: on-device, cloud, or hybrid | Technical specification |
| Design | Architecture prototype, model selection | Diagram, performance estimate |
| Implementation | Integrate Core ML / ML Kit, fallback | Working code, tests |
| Testing | 20+ real photos, speed measurements | Quality report |
| Deployment | App Store / Google Play, TestFlight | Access to build |
We also provide API documentation, model update instructions, and 30 days of support after delivery.
Timelines and guarantees
On-device background removal (iOS VisionKit + Android ML Kit) with basic UI — 3–4 days. Hybrid with fallback and post-processing — 8–10 days. We guarantee that the mask will be correct on 95% of test images. For any issues — rework at our expense.
Order integration now — get a demo build in 1 day. Contact us for a preliminary assessment of your scenario and technical consultation.







