Integrating Gesture Recognition into Mobile Apps

TRUETECH is engaged in the development, support and maintenance of iOS, Android, PWA mobile applications. We have extensive experience and expertise in publishing mobile applications in popular markets like Google Play, App Store, Amazon, AppGallery and others.

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Integrating Gesture Recognition into Mobile Apps
Complex
~1-2 weeks
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    858
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    743
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1159
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1034
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    968
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    562

We integrate camera-based gesture recognition into mobile apps — from fitness trackers to AR interfaces. Under the hood is MediaPipe Hand Landmarker, which detects up to 2 hands and returns 21 landmarks per hand. This enables touchless control: flipping slides, recognizing static and dynamic camera gestures. Our track record: 5+ years in the market and 30+ delivered computer vision projects. We offer a free assessment — contact us for a consultation. All recognition runs on-device, reducing latency and preserving data privacy.

MediaPipe Hand Landmarker provides 21 landmarks per hand.

Why client-side gesture recognition is faster than cloud alternatives

Cloud services add latency for transmission and processing, which is unacceptable for real-time UI (needs ≤33 ms at 30 FPS). Client-side ML models like MediaPipe work locally, saving up to 40% of communication time. For fitness trackers, every missed frame reduces rep-count accuracy. We leverage on-device inference and edge computing to minimize overhead.

Problems we solve

Latency. On Pixel 7 (GPU) inference takes ~12 ms, on iPhone 14 (CPU) ~18 ms, on Snapdragon 665 ~45 ms. Real-time UI at 30 FPS needs to fit in 33 ms. We optimize by reducing frame rate to 15 FPS or limiting to one hand (numHands = 1), and applying model quantization for older devices. MediaPipe Gesture Recognizer is 2-5× faster than ML Kit on older devices.

Handedness. MediaPipe returns LEFT/RIGHT from the model's perspective, which is mirrored relative to the front camera. If logic depends on a specific hand, we apply a mirror correction after landmark normalization.

Distance from camera. Normalized coordinates carry no physical distance information — thresholds that work at 50 cm will be different at 150 cm. We account for this when tuning geometry.

Recognition approaches we use

Approach Inference time Flexibility When to use
MediaPipe Gesture Recognizer 5-10 ms (additional) 7 basic gestures Quick start, standard camera gestures
Geometric 0 ms High for static gestures Specific hand shape (e.g., Open_Palm)
ML classifier 10-20 ms (additional) Maximum Complex or dynamic gestures (waves)

How we do it

Tech stack and tools — gesture recognition integration

  • iOS: Swift 5.9, SwiftUI, MediaPipeTasksVision via SPM, actor class with @Published.
  • Android: Kotlin, Jetpack Compose, com.google.mediapipe:tasks-vision with RunningMode.LIVE_STREAM.
  • Backend: none required — everything runs on the client.

How to implement a custom gesture: step by step

  1. Collect reference videos of the gesture under different angles.
  2. Annotate landmarks using MediaPipe Model Maker.
  3. Train the classifier (10-15 minutes for 1000 samples), improving model accuracy via transfer learning.
  4. Integrate into the app with debouncing and confidence threshold.

Case: Gesture-controlled presentation

In our practice, a client wanted touchless slide control for stand-up presentations. We used MediaPipe Gesture Recognizer for static gestures (Open_Palm = stop) and geometric wrist tracking for dynamic swipes. The swipe threshold (delta_x > 0.3 over 333 ms) was tuned on 15 test users. Debouncing: a new gesture is registered only after >500 ms or when the type changes. We implemented gesture-to-action binding, mapping gestures to slide navigation. Result: 98% accuracy in real-world conditions.

Recognized gestures

MediaPipe Gesture Recognizer comes with 7 built-in gestures: Open_Palm, Closed_Fist, Pointing_Up, Thumb_Up, Thumb_Down, Victory, ILoveYou. For custom gestures (e.g., circle or pistol), we use geometry or retrain the classifier via MediaPipe Model Maker. Dynamic gestures (swipes, rotations) are implemented through landmark tracking over time.

Comparison of custom gesture approaches

Characteristic Geometric ML classifier
Development time From 2 days From 5 days
Accuracy (test set) 85-92% 95-99%
Computational cost 0 ms 10-20 ms

Avoiding false positives

Key techniques:

  • Confidence threshold (usually 0.7).
  • Debouncing: a gesture is not counted again if the hand hasn't left the frame. We store lastGestureTime and lastGestureType.
  • Accounting for handedness and distance. For front camera, apply mirror correction.
  • The landmark coordinates are normalized to [0,1] and we use multiple frames to smooth detection.

What is included in the work (deliverables)

  • Analysis of your scenario: reference capture, gesture selection.
  • MediaPipe/ML Kit integration on the chosen platform.
  • Development of custom gestures (geometry or ML).
  • Binding gestures to actions with debouncing.
  • Testing on devices (5+ models).
  • Documentation and source code access.
  • Team training (2 hours).
  • Post-release support (2 weeks).

Timeline and cost

Basic integration — from 1 week (starting at $4,000). Custom gestures — from 2 weeks (starting at $7,000). Pricing is individual — contact us for an estimate. All projects are turnkey with a result guarantee. For example, one client reported saving $15,000 by using our pre-built gesture recognition modules instead of building from scratch. Request a consultation to discuss details.

Our expertise

Certified Apple and Google developers. 5+ years in mobile and computer vision. 30+ projects delivered with MediaPipe, ML Kit, and ARKit. We use official MediaPipe and ML Kit following all guidelines. This guarantees stability and performance.

Machine Learning in Mobile Apps: CoreML, TFLite, and On-Device Models

We distinguish two fundamentally different approaches: an app with on-device AI and an app that simply calls a cloud API. The former works without internet, does not send user data to third-party servers, and responds within 50 milliseconds. The latter depends on network latency and pricing plans. Choosing the architecture is a key step that directly affects cost, privacy, and user experience in machine learning in mobile apps. Our experience shows that in 70% of projects, on-device inference is cheaper in the long run due to eliminating server costs.

How to Choose Between CoreML and TFLite for On-Device Inference?

CoreML — Apple's native framework for running ML models on device. Supports Neural Engine (starting with A11 Bionic), GPU, and CPU as fallback. Models are converted to .mlmodel format via coremltools from PyTorch, ONNX, or TensorFlow. Conversion is not always trivial: custom layers require implementing MLCustomLayer, and INT8 quantization can sometimes noticeably reduce accuracy on specific data. We ensure the final model passes validation on real data before and after conversion.

TensorFlow Lite — cross-platform alternative for Android and Flutter. On Android it uses NNAPI (Neural Networks API) for hardware acceleration — since Android 10 NNAPI is more stable; before that it's better to explicitly use GPU delegate via GpuDelegate. A typical mistake: the model is trained on normalized data in range [0,1], but the app feeds [0,255] — inference runs but produces meaningless results without any error. We include an automatic input data validation module in the SDK.

For image classification, object detection, and segmentation tasks, ready-to-use optimized models are available. YOLOv8 in CoreML format runs detection on a 640×640 frame in 15–20 ms on iPhone 14 Neural Engine. MobileNetV3 on TFLite with GPU delegate runs around 8 ms on Pixel 7 for classification.

Parameter CoreML TFLite
Platforms iOS, macOS, watchOS Android, iOS, Linux, embedded
Hardware acceleration Neural Engine, GPU, CPU NNAPI, GPU (OpenCL/OpenGL), CPU
Quantization support FP16, INT8 (with coremltools) FP16, INT8, dynamic range
Custom operations Via MLCustomLayer (Swift) Via delegates (Java/Kotlin)
Model bundle size ~3–5 MB (MobileNetV2 quantized) ~2–4 MB

What If You Need Text Generation On-Device?

Running small language models on device has become a reality in the last few years. Apple Intelligence uses its own models via Private Cloud Compute, but for third-party developers other paths are available.

llama.cpp with Metal backend on iOS is a working approach for phi-3-mini (3.8B parameters, 4-bit quantization, ~2.3 GB). Inference: 15–25 tokens/second on iPhone 15 Pro. For integration in Swift, use the Swift Package llama.swift or a wrapper via C interface llama.h. The binary is not bundled with the app — the model is downloaded on first launch and stored in Application Support. Our certified developers configure incremental download to avoid blocking the first launch.

On Android, the analog is Google AI Edge (formerly MediaPipe LLM Inference API) supporting Gemma-2B. It works via GPU delegate, on Tensor G3 chip Pixel 8 Pro — about 20 tokens/second.

Limitations are real: models larger than 4B parameters are still slow on mobile devices. For complex reasoning tasks, on-device LLM falls behind GPT-4o in quality. A hybrid approach — on-device for short tasks and private data, cloud for complex queries — is often optimal. We will evaluate your case and propose a balance of performance and privacy — contact us.

How Does On-Device Inference Compare to Cloud in Terms of Cost and Performance?

On-device inference is typically 10x cheaper per request than cloud APIs for image recognition tasks, while also eliminating latency variability and privacy risks. The table below summarizes the trade-offs.

Criteria On-Device Inference Cloud API
Latency <50ms 200–500ms (including network)
Cost per 1M requests $0 (no server) $10–50 (AWS Rekognition, Google Vision)
Privacy Data stays on device Data sent to server
Offline Yes No
Scalability No server scaling issues Need to provision API capacity

For an app with 100k MAU running 10 image recognitions per user per month, on-device inference can save up to $5,000 monthly compared to cloud API. Get a free consultation on your ML architecture today.

Integrating OpenAI API and Other Cloud Models

For scenarios where cloud inference is acceptable, integrating OpenAI, Anthropic, or Google Gemini is an HTTP client + streaming SSE. In Swift, AsyncThrowingStream is convenient for streaming responses. In Kotlin, use Flow.

Critically: API keys must never be stored in the app bundle. Even an obfuscated key can be extracted from the IPA in 10 minutes using strings or frida. Correct architecture: mobile app → your own backend → OpenAI API. The backend controls rate limiting, logs requests, and protects the key.

What Is Included in the Work (Deliverables)

  • Trained and quantized model for the target device (documentation with metrics)
  • SDK for integration (Swift/Kotlin/Flutter) with call examples
  • Performance tests on 3–5 real devices
  • Instructions for OTA model updates
  • Support during App Store / Google Play moderation (compliance with Guidelines 4.2, 5.1)
  • 2 weeks of technical support after release

Typical Project Pipeline

  1. Task analysis — measure latency, privacy, size, supported devices.
  2. Model prototyping — in Python, evaluate accuracy on target data.
  3. Conversion and quantization — for CoreML/TFLite with validation.
  4. Integration into the app — model wrapped in a service layer (easy to swap CoreML ↔ TFLite ↔ cloud).
  5. Testing — on real devices, measure FPS, RAM, battery.
  6. Deployment — via TestFlight / Firebase App Distribution, monitor metrics.

Timelines: integration of a ready CoreML/TFLite model — 1–2 weeks, development of a custom model with mobile optimization — from 6 weeks, on-device LLM chat with personalization — 4–8 weeks.

Why We Take on Complex Cases?

10+ years of experience in mobile development, 50+ implemented AI/ML solutions, guarantee of compatibility with current iOS and Android versions. All projects undergo code review and load testing. The cost includes preparation of moderation documentation and training of your team.

Contact us — we will help you choose the architecture and implement ML in your app turnkey. Order an audit of your existing solution — we will assess the potential for server cost savings free of charge. In some projects, savings can reach significant amounts per month.