AI Face Grouping for Mobile Photo Galleries
Imagine: a user has 50,000 photos in their gallery. Finding all shots with a specific person manually takes hours. A typical gallery contains tens of thousands of photos, and each shot may include several faces — AI is indispensable. Automatic grouping by face solves this, but requires careful implementation: detection accuracy, embedding extraction speed, and correct clustering. We specialize in such pipelines: 8+ years in mobile development, 20+ projects with computer vision. Our solution runs entirely on-device: face detection (Vision/ML Kit) → embedding extraction with MobileFaceNet (128-dimensional vector) → DBSCAN clustering. This approach ensures privacy, speed, and offline functionality. On-device processing saves up to 70% on cloud infrastructure costs compared to server solutions — reducing the app's TCO.
How AI Grouping Works On-Device
The pipeline consists of three stages, each optimized for mobile processors.
Face Detection
On iOS, we use VNDetectFaceRectanglesRequest from the Vision framework:
let request = VNDetectFaceRectanglesRequest { req, _ in
guard let faces = req.results as? [VNFaceObservation], !faces.isEmpty else { return }
for face in faces {
let faceRect = VNImageRectForNormalizedRect(face.boundingBox, width, height)
self.extractEmbedding(from: originalImage.cropping(to: faceRect)!)
}
}
On Android — ML Kit FaceDetector (simple integration) or MediaPipe FaceDetector (more control). The choice depends on accuracy and model size requirements.
Embedding Extraction
Apple does not provide a built-in API for face recognition (only detection). We use MobileFaceNet — a compact model (1–3 MB) running via Core ML. After L2 normalization, the cosine distance between embeddings of the same person is <0.3, between different people >0.6. A threshold of 0.4–0.5 works well in practice.
func extractEmbedding(from faceImage: CGImage) -> [Float]? {
guard let input = try? MobileFaceNetInput(face_image: MLMultiArray(from: resize(faceImage, to: CGSize(width: 112, height: 112)))) else { return nil }
guard let output = try? facenetModel.prediction(input: input) else { return nil }
let embedding = (0..<128).map { output.embedding[$0].floatValue }
return l2Normalize(embedding)
}
func l2Normalize(_ v: [Float]) -> [Float] {
let norm = sqrt(v.reduce(0) { $0 + $1 * $1 })
return norm > 0 ? v.map { $0 / norm } : v
}
Clustering
For clustering without a predefined number of clusters, we use DBSCAN (density-based spatial clustering). In Swift, we implement it with Accelerate/BLAS for cosine distance computation:
func dbscan(embeddings: [[Float]], eps: Float = 0.45, minPoints: Int = 2) -> [Int] {
var labels = Array(repeating: -1, count: embeddings.count)
var clusterId = 0
for i in 0..<embeddings.count {
guard labels[i] == -1 else { continue }
let neighbours = rangeQuery(embeddings: embeddings, idx: i, eps: eps)
if neighbours.count < minPoints { continue }
labels[i] = clusterId
var seeds = neighbours
while !seeds.isEmpty {
let q = seeds.removeFirst()
if labels[q] == -1 { labels[q] = clusterId }
if labels[q] != clusterId { continue }
labels[q] = clusterId
let qNeighbours = rangeQuery(embeddings: embeddings, idx: q, eps: eps)
if qNeighbours.count >= minPoints { seeds.append(contentsOf: qNeighbours) }
}
clusterId += 1
}
return labels
}
For galleries with 20,000+ embeddings, DBSCAN with O(n²) can run 10–30 seconds. Speedup is achieved via Approximate Nearest Neighbor (ANN) using FAISS (with Swift bindings), reducing complexity to O(n log n).
More about the MobileFaceNet model
MobileFaceNet is a 4-layer convolutional network that extracts 128-dimensional embeddings. The model size is only 1–3 MB, ideal for mobile devices. It was trained on the MS-Celeb-1M dataset.Why On-Device is Better than Server-Side
| Criteria | On-Device | Server-Side |
|---|---|---|
| Privacy | Embeddings never leave the device | Requires consent to transmit biometrics |
| Speed | Instant, no network latency | Depends on connection and load |
| Offline Mode | Full functionality without internet | Requires constant connection |
| Infrastructure Cost | No server costs | High GPU and storage expenses |
On-device processing is 3–5 times faster for typical galleries (up to 10,000 photos) and fully complies with GDPR, CCPA, and other regulations. Cloud computing costs drop to zero, which is especially important for startups with limited budgets.
Embedding Model Comparison
| Model | Size | Dimensionality | Accuracy (LFW) |
|---|---|---|---|
| MobileFaceNet | 1-3 MB | 128 | 99.2% |
| FaceNet | 10-30 MB | 512 | 99.6% |
| ArcFace | 20-50 MB | 512 | 99.8% |
For mobile applications, MobileFaceNet is optimal — a balance of accuracy and performance.
Performance on Large Galleries
A real gallery has 5,000–50,000 photos. Faces appear in about 30–40% of them. For example, 10,000 photos with faces, 2 faces on average = 20,000 embeddings. DBSCAN processes them in 10–30 seconds on an iPhone 14. With FAISS — 2–3 seconds. We register the background task via BGProcessingTask (iOS 13+):
BGTaskScheduler.shared.register(forTaskWithIdentifier: "com.app.faceGrouping") { task in
let bgTask = task as! BGProcessingTask
self.runFaceGrouping(completion: { bgTask.setTaskCompleted(success: true) })
bgTask.expirationHandler = { /* save progress */ }
}
Storing Results
Embeddings are biometric data, so we store them locally in Core Data with Data Protection encryption (.complete). The mapping identifier is PHAsset.localIdentifier, not the photo itself. Embeddings are not synced to iCloud without explicit user consent.
What's Included in the Work
- Analysis — assessment of gallery size, selection of detection and clustering models.
- Design — pipeline architecture, integration with Core Data / Room.
- Implementation — coding in Swift/Kotlin, training/calibrating thresholds.
- Testing — A/B tests on real galleries, clustering accuracy verification.
- Documentation — API description, integration guide, configuration manual.
- Support — 3-month warranty, maintenance during OS updates.
Our team with 8 years of experience and Apple/Google certifications guarantees timely delivery. Order development of on-device AI grouping for your app. Get a consultation — we'll find the optimal solution for your gallery.
Timelines
A basic on-device pipeline (detection, embeddings, clustering) for medium galleries takes 2–3 weeks. A scalable version with FAISS, background processing, incremental updates, and UI takes 4–5 weeks. Contact us — we'll calculate cost and timelines individually.
FaceNet: A Unified Embedding for Face Recognition and Clustering, Schroff et al.







