AI Liveness Detection Implementation for Verification
Attacks with photos, videos, and masks are a reality for any service with remote verification. Liveness Detection distinguishes a live person from an artifact, but the choice between active and passive check is critical. Active requires actions (head turn, blinking) — this provides high protection against photo spoofing but reduces conversion by 15–20%. Passive analyzes a single frame or short video without requiring actions, but handles 2D attacks worse without depth analysis.
We integrate liveness solutions end-to-end: from strategy selection to app store publication. Over 5 years, we have delivered 20+ projects for fintech, EdTech, and government sectors. Our stack includes ARKit (iOS), ML Kit + ARCore (Android), CoreML/TFLite, and ready-made certified SDKs (Iproov, Jumio, Onfido, Sumsub). For each project, we build a threat model and select the optimal balance between security and UX.
One case — a banking app: after implementing a passive check with depth map, verification time dropped to 3 seconds (from 12 seconds for active), and conversion increased by 12%. Such results confirm that the right liveness choice directly impacts business metrics.
Choosing Between Active and Passive Liveness
Active liveness requires the user to perform specific actions: head turn, blink, or speak a code. This provides high resistance to photo spoofing, but at the cost of user experience — 15–20% of users fail on the first attempt, reducing conversion rates. Additionally, active liveness is vulnerable to deepfake videos that can synthesize the requested movements on command.
Passive liveness, on the other hand, analyzes a single frame or a short video sequence without any user instructions; the user simply looks at the camera. This offers a much better user experience with higher conversion, but it is less robust against 2D attacks like high-quality photos. Modern passive models that comply with ISO 30107-3 Level 2 can withstand attacks using printed photos and 2D screens, especially when combined with depth analysis.
Active liveness is easier to certify for ISO 30107 Level 2, while passive liveness often requires a depth sensor or a combination of techniques to achieve the same level. The choice depends on your threat model and UX priorities. If your primary concern is protection against photo spoofing and you can tolerate a lower conversion rate, active may be suitable. For high-traffic applications where user experience is critical, a well-designed passive solution with depth analysis can yield better business outcomes.
If you are unsure which strategy fits your scenario, get a consultation — we will analyze your threat model and propose the optimal solution.
Implementation on iOS via ARKit
ARKit on iPhone X+ and iPad Pro with TrueDepth camera provides real-time depth map access. ARFaceTrackingConfiguration creates ARFaceAnchor with 52 blend shapes — including eyeBlinkLeft, eyeBlinkRight, jawOpen. This already gives full active liveness without third-party SDKs.
let config = ARFaceTrackingConfiguration() config.maximumNumberOfTrackedFaces = 1 session.run(config) // In ARSCNViewDelegate func renderer(_ renderer: SCNSceneRenderer, didUpdate node: SCNNode, for anchor: ARAnchor) { guard let faceAnchor = anchor as? ARFaceAnchor else { return } let eyeBlink = faceAnchor.blendShapes[.eyeBlinkLeft]?.floatValue ?? 0 if eyeBlink > 0.7 { livenessDetector.recordBlink() } } Depth map from depthData on AVCapturePhotoOutput allows rejecting flat images: if stddev of face depth is <5mm — a flat surface is in front of the camera.
Limitation: ARKit FaceTracking works only on devices with TrueDepth (front True Depth camera). iPhone SE, iPad mini — not supported. For these devices, fallback is an RGB-only CoreML model.
Implementation on Android via ML Kit Face Mesh
com.google.mlkit:face-mesh-detection (ML Kit 18.0+) provides 468 3D mesh points from a monocular camera. This is not true depth but 3D reconstruction from 2D — better than nothing but weaker than ARKit TrueDepth.
val options = FaceMeshDetectorOptions.Builder() .setUseCase(FaceMeshDetectorOptions.FACE_MESH) .build() val detector = FaceMeshDetection.getClient(options) detector.process(inputImage) .addOnSuccessListener { faces -> faces.firstOrNull()?.let { face -> val zVariance = face.allPoints.map { it.position.z }.variance() if (zVariance < FLAT_THRESHOLD) rejectAsFlatImage() } } On Android 10+ with ARCore-compatible devices, it is better to use ArCoreApk + AugmentedFace — obtaining true depth via structured light or ToF (devices with such sensors: Pixel 6+, Samsung S21+).
| Platform | Technology | Devices |
|---|---|---|
| iOS | TrueDepth (ARKit) | iPhone X+, iPad Pro 2018+ |
| Android | ARCore (Depth API) | Pixel 6+, Samsung S21+ (with ToF) |
| Android (fallback) | ML Kit Face Mesh | All with camera (no depth) |
When to Use Third-Party SDKs?
Iproov, Jumio, Onfido, Sumsub — ready-made liveness SDKs with ISO 30107-3 Level 2 certification. Certification alone costs hundreds of thousands of dollars and takes months. If your product operates in a regulatory environment requiring a certified solution, a custom implementation is not advisable.
If regulation does not require certification and the goal is protection against basic attacks (photo spoofing, phone video) — a custom implementation on ARKit + CoreML / ML Kit + TFLite handles it and is significantly more cost-effective. Typical savings range from $50,000 to $150,000 per year compared to third-party SDK licensing, depending on scale.
More on standards and certification
The standard ISO/IEC 30107-3:2023 defines attack levels (Presentation Attack Detection) for biometric systems. Level 2 is the minimum for KYC, including protection against printed photos and 2D screens. Certification to this standard is mandatory for financial regulators in the EU, USA, and several other countries.Common Mistakes
Threshold without context. livenessScore > 0.85 in code with no explanation — after a month, no one remembers where the number came from or how it changed. Use a configurable threshold with A/B testing and FRR/FAR metrics.
Ignoring deepfake attacks. Passive liveness without depth is vulnerable to GAN-generated faces. If this is in your threat model, you need texture inconsistency analysis (GAN artifacts in frequency domain via FFT) or server-side inference on a heavier model.
What Our Work Includes
- Threat model analysis and selection of active/passive/hybrid liveness.
- SDK integration or custom development on ARKit/ML Kit/CoreML/TFLite.
- Threshold tuning with A/B testing and FAR/FRR metrics.
- Attack testing (photo, screen, mask, deepfake).
- Integration with the IDV pipeline and backend.
- Publication in App Store / Google Play with guideline compliance.
- Documentation of the system architecture, API contracts, and deployment procedures.
- Team training covering operation, maintenance, and response to evolution of attacks.
- Ongoing support for the first 3 months post-launch, with options for extended coverage.
Timelines: ready-made SDK integration — 2–4 weeks. Custom implementation with model training — 8–14 weeks. Cost is calculated individually.
Evaluate your project — contact us for a consultation. Get a preliminary estimate of timelines and cost within one business day.







