We integrate camera-based gesture recognition into mobile apps — from fitness trackers to AR interfaces. Under the hood is MediaPipe Hand Landmarker, which detects up to 2 hands and returns 21 landmarks per hand. This enables touchless control: flipping slides, recognizing static and dynamic camera gestures. Our track record: 5+ years in the market and 30+ delivered computer vision projects. We offer a free assessment — contact us for a consultation. All recognition runs on-device, reducing latency and preserving data privacy.
MediaPipe Hand Landmarker provides 21 landmarks per hand.
Why client-side gesture recognition is faster than cloud alternatives
Cloud services add latency for transmission and processing, which is unacceptable for real-time UI (needs ≤33 ms at 30 FPS). Client-side ML models like MediaPipe work locally, saving up to 40% of communication time. For fitness trackers, every missed frame reduces rep-count accuracy. We leverage on-device inference and edge computing to minimize overhead.
Problems we solve
Latency. On Pixel 7 (GPU) inference takes ~12 ms, on iPhone 14 (CPU) ~18 ms, on Snapdragon 665 ~45 ms. Real-time UI at 30 FPS needs to fit in 33 ms. We optimize by reducing frame rate to 15 FPS or limiting to one hand (numHands = 1), and applying model quantization for older devices. MediaPipe Gesture Recognizer is 2-5× faster than ML Kit on older devices.
Handedness. MediaPipe returns LEFT/RIGHT from the model's perspective, which is mirrored relative to the front camera. If logic depends on a specific hand, we apply a mirror correction after landmark normalization.
Distance from camera. Normalized coordinates carry no physical distance information — thresholds that work at 50 cm will be different at 150 cm. We account for this when tuning geometry.
Recognition approaches we use
| Approach | Inference time | Flexibility | When to use |
|---|---|---|---|
| MediaPipe Gesture Recognizer | 5-10 ms (additional) | 7 basic gestures | Quick start, standard camera gestures |
| Geometric | 0 ms | High for static gestures | Specific hand shape (e.g., Open_Palm) |
| ML classifier | 10-20 ms (additional) | Maximum | Complex or dynamic gestures (waves) |
How we do it
Tech stack and tools — gesture recognition integration
- iOS: Swift 5.9, SwiftUI,
MediaPipeTasksVisionvia SPM, actor class with@Published. - Android: Kotlin, Jetpack Compose,
com.google.mediapipe:tasks-visionwithRunningMode.LIVE_STREAM. - Backend: none required — everything runs on the client.
How to implement a custom gesture: step by step
- Collect reference videos of the gesture under different angles.
- Annotate landmarks using MediaPipe Model Maker.
- Train the classifier (10-15 minutes for 1000 samples), improving model accuracy via transfer learning.
- Integrate into the app with debouncing and confidence threshold.
Case: Gesture-controlled presentation
In our practice, a client wanted touchless slide control for stand-up presentations. We used MediaPipe Gesture Recognizer for static gestures (Open_Palm = stop) and geometric wrist tracking for dynamic swipes. The swipe threshold (delta_x > 0.3 over 333 ms) was tuned on 15 test users. Debouncing: a new gesture is registered only after >500 ms or when the type changes. We implemented gesture-to-action binding, mapping gestures to slide navigation. Result: 98% accuracy in real-world conditions.
Recognized gestures
MediaPipe Gesture Recognizer comes with 7 built-in gestures: Open_Palm, Closed_Fist, Pointing_Up, Thumb_Up, Thumb_Down, Victory, ILoveYou. For custom gestures (e.g., circle or pistol), we use geometry or retrain the classifier via MediaPipe Model Maker. Dynamic gestures (swipes, rotations) are implemented through landmark tracking over time.
Comparison of custom gesture approaches
| Characteristic | Geometric | ML classifier |
|---|---|---|
| Development time | From 2 days | From 5 days |
| Accuracy (test set) | 85-92% | 95-99% |
| Computational cost | 0 ms | 10-20 ms |
Avoiding false positives
Key techniques:
- Confidence threshold (usually 0.7).
- Debouncing: a gesture is not counted again if the hand hasn't left the frame. We store
lastGestureTimeandlastGestureType. - Accounting for handedness and distance. For front camera, apply mirror correction.
- The landmark coordinates are normalized to [0,1] and we use multiple frames to smooth detection.
What is included in the work (deliverables)
- Analysis of your scenario: reference capture, gesture selection.
- MediaPipe/ML Kit integration on the chosen platform.
- Development of custom gestures (geometry or ML).
- Binding gestures to actions with debouncing.
- Testing on devices (5+ models).
- Documentation and source code access.
- Team training (2 hours).
- Post-release support (2 weeks).
Timeline and cost
Basic integration — from 1 week (starting at $4,000). Custom gestures — from 2 weeks (starting at $7,000). Pricing is individual — contact us for an estimate. All projects are turnkey with a result guarantee. For example, one client reported saving $15,000 by using our pre-built gesture recognition modules instead of building from scratch. Request a consultation to discuss details.
Our expertise
Certified Apple and Google developers. 5+ years in mobile and computer vision. 30+ projects delivered with MediaPipe, ML Kit, and ARKit. We use official MediaPipe and ML Kit following all guidelines. This guarantees stability and performance.







