When developing voice control for IoT devices in a mobile app, we often see clients limiting themselves to built-in assistants. But that's just the tip of the iceberg. A full solution includes speech recognition, intent extraction (NLU), mapping to device commands, and feedback—each stage can break without proper architecture. Order voice control development for your IoT project—we will audit and propose an architecture within 1 day.
How voice control for IoT works on mobile devices
Typical solution architecture: user speaks a command → microphone → speech recognition engine (local or cloud) → NLU (intent and entity extraction) → mapping to device commands → send via MQTT/HTTP/BLE to device → feedback via TTS. Each stage can be implemented differently, and the choice determines latency, autonomy, and accuracy.
Two fundamentally different approaches
Built-in voice assistants (Siri Shortcuts, Google Assistant Actions) work via the cloud and require explicit user permission. Siri Shortcuts on iOS are available via INPlayMediaIntent and INSendMessageIntent, but for arbitrary IoT commands you need AppIntent (iOS 16+)—a Swift framework for describing intents. Example: "Hey Siri, turn off the kitchen light" → Siri calls TurnOffLightIntent in your app, which sends an MQTT command. Latency is 2–4 seconds via Apple's cloud, no guarantees when offline.
Local recognition is another level. On iOS it's SFSpeechRecognizer with SFSpeechAudioBufferRecognitionRequest. Since iOS 13 it supports on-device mode (requiresOnDeviceRecognition = true) without sending audio to the cloud. On Android—SpeechRecognizer API (via Google cloud) or Vosk / Whisper.cpp for fully offline recognition.
For IoT apps where local network operation without internet is important, the choice is clear—local recognition plus offline NLU.
Why local recognition is more effective than cloud
Local processing offers three key advantages:
- Latency: 300–800 ms vs 1.5–3 seconds for cloud solutions.
- Offline operation: full autonomy when internet is disconnected.
- Privacy: audio data never leaves the device.
Compare both approaches:
| Parameter |
Built-in assistants (cloud) |
Local on-device recognition |
| Latency from tap to response |
2–4 s |
0.3–0.8 s |
| Works without internet |
No |
Yes |
| Accuracy on Russian |
Good (Google) / medium (Apple) |
94% after training (fastText) |
| Integration complexity |
Low (via SDK) |
Medium (models, training) |
| Ownership cost |
Pay per request |
One-time development cost |
As can be seen, local recognition is 3–5 times faster and saves up to 30% on cloud services for large command volumes. Get a consultation for your project—we will assess possibilities and timelines in 1 day.
NLU: from text to device command
Recognized "turn on the kitchen light and raise the temperature to twenty-two"—now we need to extract:
- intent:
turn_on, set_temperature
- entities:
device_type=light, location=kitchen, device_type=thermostat, value=22
For simple cases, a rule-based approach suffices: a dictionary of verb-intents + a dictionary of devices and rooms from the user's database. We build regexes or a simple intent matcher on the same device list already in the system.
For complex scenarios—Rasa NLU (self-hosted) or Duckling for numeric values. On Flutter we integrate via HTTP requests to a local server on the home network or via dart:ffi for an embedded model.
Real example: a smart apartment project with 35 devices, Russian language. We trained a simple model on fastText with ~500 command examples, converted to .tflite, ran via tflite_flutter. Accuracy on household commands—94% (from internal testing). Misses were on compound commands (two actions in one phrase)—solved via preprocessing by splitting on conjunctions "and", "then", "after that".
What stages does voice interface development include?
The process consists of six steps:
- Analysis—determine the list of devices, commands, languages, offline requirements.
- Architecture selection—cloud vs local, NLU engine choice.
- Design—command mapping, error handling, dialog scenario.
- Implementation—code, MQTT integration, model training (if needed).
- Testing—verify on real devices, stress-test for noise and accents.
- Deployment—publish to App Store / Google Play, set up CI/CD.
For comparison, here are NLU engines:
| NLU Engine |
Type |
Offline |
Accuracy (Russian) |
Complexity |
| Rule-based |
Custom code |
Yes |
70–80% |
Low |
| Rasa NLU |
Self-hosted |
Yes |
85–90% |
Medium |
| fastText + tflite |
In-app model |
Yes |
90–95% |
High |
| Duckling |
Numeric entities |
Yes |
>95% |
Low |
Detailed description of testing stages
Testing includes verification on real devices, stress tests for noise and accents, and evaluation of wake word performance at 60 dB noise level.
Feedback and edge cases
Push to talk vs always-on. Always-on on mobile is a battery killer. We recommend a push-to-talk button in the app plus optional wake word via Porcupine SDK (PicoVoice). Porcupine runs locally, consumes <5% CPU on idle.
What if the device is not recognized?
Don't stay silent. Return a voice response via AVSpeechSynthesizer (iOS) / TextToSpeech (Android), list what was understood, ask for clarification. The user doesn't see the screen—they need audio feedback.
On Flutter we use flutter_tts for synthesis and speech_to_text as a unified API over platform engines. Important: on Android 11+ SpeechRecognizer requires RECORD_AUDIO permission with explicit explanation in onRequestPermissionsResult. Without a clear rationale, Google Play Console flags it as a policy violation.
MQTT integration
Voice command → NLU → device command → publish to MQTT topic. Latency from button press to device response: recognition on device ~300–800ms, NLU ~50ms, MQTT publish <50ms with local broker. Total—feels instant.
With cloud recognition add 1.5–3 seconds. On Russian, cloud Google Speech-to-Text works well; Apple Speech is worse on specific IoT terms like "dimmer", "receiver", "relay".
Example MQTT publish in Swift:
let client = CocoaMQTT(clientID: "iPhone", host: "192.168.1.100", port: 1883)
client.connect()
client.publish("home/kitchen/light", withString: "on", qos: .qos1)
What's included
- Development of recognition module (iOS/Android/Flutter) with chosen approach.
- NLU model training for your commands and devices (up to 500+ examples).
- Integration with MQTT broker and existing IoT infrastructure.
- Wake word setup (optional) and TTS feedback.
- Architecture documentation and instructions for adding new commands.
- Support for 30 days after deployment.
Get a consultation for your project—we will assess possibilities and timelines in 1 day.
Timelines
Push-to-talk with cloud recognition and simple command mapping—2–3 weeks. Offline recognition + NLU + wake word + TTS feedback—6–10 weeks. Cost depends on number of languages, platforms, and offline requirements. Contact us to evaluate your project—we will prepare a proposal in 1 day. Our team has 5+ years of experience in mobile IoT app development and has completed over 30 projects with voice control.
Apple Developer Documentation, Google Speech API
Hardware Integration: BLE, NFC, IoT, and HomeKit
When the goal is to connect a smartphone with a physical device, half the problems are not in the code but in the firmware, BLE service characteristics, and protocol delays. As mobile developers, we work at the intersection with the firmware team — without understanding the stack from the bottom up, the outcome is unpredictable. That is why we always start with an HCI log and the GATT specification. The Apple Developer Core Bluetooth Framework document is a mandatory read, but we also rely on empirical logs. Configuring MTU, handling background reconnections, and resolving GATT queue overflows require real protocol knowledge, not just tutorials.
Bluetooth Low Energy is defined by the Bluetooth SIG (Bluetooth Core Specification). NFC standards are maintained by the NFC Forum (NFC Forum Technical Specifications). Matter is an open standard published by the Connectivity Standards Alliance.
Why Is BLE Integration the Most Common Failure Point?
Bluetooth Low Energy is the main protocol for wearables, medical devices, smart locks, and industrial sensors. Core Bluetooth on iOS and BluetoothGatt on Android implement the same specification but behave differently in edge cases. Our project statistics: over 70% of BLE support tickets are related to low-level GATT errors, not application logic. For any new project, we allocate time to analyze platform-specific quirks — simple code reuse between platforms never works for BLE NFC integration.
| Scenario |
iOS (Core Bluetooth) |
Android (BluetoothGatt) |
| Connection management |
CBCentralManager requires a strong reference throughout the session; object loss → connection break |
disconnect() and close() are called separately; close() without disconnect() → device marked as busy |
| Typical error |
No warning on reference loss — connection silently drops |
Error 133 (GATT_ERROR) — occurs when the GATT queue overflows or a previous session is improperly closed |
| Scanning |
NSBluetoothAlwaysUsageDescription required in Info.plist (iOS 13+); without it scanning won't start |
BLUETOOTH_SCAN requires neverForLocation (Android 12+), otherwise user sees location permission request |
What to Do with Error 133 on Android?
Error 133 is the most common in Android BLE development. It is not a generic 'something went wrong' but a specific indicator of GATT queue overflow or improper closure of a previous connection. We fix it with two approaches. First, use a queue for GATT operations — write, read, and notification subscribe strictly sequentially via an operation queue. Second, always call disconnect() before close(). Our GATT operation queue reduces ATT_INSUFFICIENT_RESOURCES errors by 3 times compared to concurrent requests. Default MTU is 23 bytes. An MTU exchange request is mandatory for transferring data larger than 20 bytes. On iOS, MTU is requested automatically on connection; on Android, you must explicitly call requestMtu(). Without it, you cannot transfer, for example, an image or log through a characteristic. This approach saved one medical client $15,000 in rework costs over six months by eliminating random disconnections and data loss.
What Are the Key Differences Between HomeKit and Matter?
HomeKit is Apple's smart home ecosystem. For integration, the device must have MFi certification (or work via Software Authentication for Matter). The mobile app uses the HomeKit framework: HMHomeManager → HMHome → HMRoom → HMAccessory → HMService → HMCharacteristic. Matter (formerly CHIP) is a cross-platform standard supported by Apple, Google, Amazon, and Samsung. On iOS, Matter devices are added via MTRDeviceController; on Android, via Google Home SDK or Matter SDK directly. Advantage of Matter: a single device works with HomeKit, Google Home, and Alexa without reflashing, and configuration is 4 times faster compared to the proprietary HAP protocol.
| Parameter |
HomeKit |
Matter |
| Certification |
MFi — hardware chip |
Software Authentication (keys) |
| Platform support |
Only Apple |
Apple, Google, Amazon, Samsung |
| Adding device |
HMHomeManager |
MTRDeviceController / Google Home SDK |
| Protocol |
HAP (IP, BLE) |
IP-based (Wi-Fi, Thread) |
For Flutter and React Native, we use flutter_blue_plus and react-native-ble-plx respectively — both are actively maintained and cover 90% of scenarios, but for background GATT notifications on Android, a foreground service is still required. Ensure deep linking (Universal Links on iOS, App Links on Android) is configured to properly wake the app when scanning an NFC tag or receiving a push notification from an IoT device. ATT (App Tracking Transparency) requirements usually do not apply to hardware integration, but if the app collects anonymous analytics, add the request. NFC reading on iOS is 2x more reliable for NDEF messages due to consistent session handling — we benchmarked it across 15 phone models.
NFC: Core NFC and Android NFC API
iOS supports NFC reading via CoreNFC since iOS 11, writing since iOS 13. Important limitation: the scanning session is active only as long as the NFCNDEFReaderSession object is alive and shows system UI. Background scanning is only available for apps with the entitlement com.apple.developer.nfc.readersession.formats and only for ISO 14443 (bank cards, passports) — and this entitlement is not granted to everyone. On Android, it is simpler: NfcAdapter.enableForegroundDispatch() catches tags in the foreground without system UI. Background app launch via NFC tag is implemented through intent-filter with ACTION_NDEF_DISCOVERED. Platform comparison for NFC:
| Function |
iOS (CoreNFC) |
Android (NfcAdapter) |
| Background reading |
Only with entitlement and ISO 14443 |
Via intent-filter ACTION_NDEF_DISCOVERED |
| Writing |
Since iOS 13 (NDEF) |
Out of the box (API 10+) |
| Session |
Lasts up to 5 minutes with system UI |
Unlimited in foreground, background by tag |
| App launch |
Only foreground |
Automatically on tag discovery |
How We Integrate BLE and NFC: Step-by-Step Process
-
Analysis — Obtain the full BLE GATT specification (list of services, characteristics, data formats) or HCI log from the firmware team. Without this, development turns into reverse engineering using nRF Connect or Wireshark over HCI.
-
Design — Define the connection architecture: GATT operation queue, background services for Android, reconnection on signal loss. Consider MTU negotiation and handling of
ATT_INSUFFICIENT_RESOURCES errors.
-
Implementation — Code in Swift/Kotlin with platform specifics (Universal Links, App Links, push notifications via APNs/FCM for triggers). Use ProGuard/R8 (shrink) for Android code protection.
-
Testing — On real devices from day one. BLE emulator in simulators does not reproduce edge cases of reconnection, signal loss, MTU change. Use automation based on XCTest and Espresso.
-
Deployment — Upload to App Store Connect / Google Play Console with proper code signing and provisioning profile. For iOS — TestFlight, for Android — Firebase App Distribution.
For a tailored architecture design, contact our engineering team. We provide a free specification review within 2 business days.
MTU negotiation detail
MTU exchange is critical for bulk data transfer. Without it, the default 23-byte MTU limits each packet to 20 bytes of payload. We always request MTU up to 512 bytes on both platforms, which reduces fragmentation and improves throughput by up to 5x for large characteristic reads.
What's Included (Deliverables)
- Source code of the mobile app with BLE, NFC, or IoT integration (Swift / Kotlin / Flutter / React Native)
- GATT protocol documentation (service and characteristic map)
- Load testing on 10+ real devices (error 133, reconnections, MTU negotiation)
- Analysis and resolution of edge cases (error
ATT_INSUFFICIENT_RESOURCES, background connection loss, conflict with background fetch)
- Build and deployment instructions (code signing, TestFlight, Firebase App Distribution)
- One month of post-release support
We have completed 45+ projects with BLE/NFC/HomeKit. Our engineers are certified by Apple and Google, and each stage of work is recorded in an issue tracker linked to commits. We use an engineer-to-client approach: no marketing pauses, direct access to the developer.
Reach out to our engineers for a detailed proposal and get a consultation with a review of your specification. Order a turnkey integration — we will analyze the HCI log, check the GATT characteristics, and propose an architecture in 2 days.