AI Scene Recognition for Smart Home Automation in a Mobile App
Typical scenario: you launch a smart home with a camera, and automation only triggers when a user enters a room but fails to respond to finer scenes—like dimming the lights during a movie. The problem is that cloud ML services add 200–500ms latency and risk privacy. Local classification on the device is the only way for real-time scenarios. This reduces operational costs for cloud computing, especially when working with multiple cameras. We have implemented over 10 such integrations and know all the pitfalls: from false triggers in poor lighting to conflicts with App Store Review Guidelines. Experience shows that the right architecture saves resources and time.
Why Local Processing Is Critical for Smart Home?
Sending camera frames to a server for classification is a bad idea for home automation. Latency is unacceptable, and the user loses control over data. Everything must work locally. For example, on iOS we use Apple Vision and CoreML, which perform classification on the A14+ chip in <30ms.
iOS: CoreML + Vision Framework
Apple Vision Scene Classification—built-in model VNClassifyImageRequest. Works offline, returns VNClassificationObservation with confidence score. For smart home, ~20 categories out of 3000+ built-in are enough. Apple documentation (Apple Vision Scene Classification) recommends this approach.
import Vision import CoreML class SceneClassifier { private lazy var request: VNClassifyImageRequest = { let r = VNClassifyImageRequest { [weak self] request, error in self?.handleResults(request.results as? [VNClassificationObservation]) } return r }() func classify(pixelBuffer: CVPixelBuffer) { let handler = VNImageRequestHandler(cvPixelBuffer: pixelBuffer, options: [:}) try? handler.perform([request]) } private func handleResults(_ results: [VNClassificationObservation]?) { guard let top = results?.filter({ $0.confidence > 0.6 }).first else { return } // top.identifier: "bedroom", "kitchen", "living_room", "bathroom" SmartHomeAutomation.shared.triggerScene(top.identifier) } } Filter by confidence > 0.6 and by a list of relevant identifiers. Do not process frames more often than every 2–3 seconds—this saves battery and CPU. For custom scenarios, we use Create ML with MobileNetV3, export to .mlpackage, size ~4 MB.
Android: ML Kit Scene Detection + TFLite
ML Kit Subject Segmentation and Scene Detection work offline on the device:
val image = InputImage.fromMediaImage(mediaImage, rotation) val labeler = ImageLabeling.getClient( ImageLabelerOptions.Builder() .setConfidenceThreshold(0.65f) .build() ) labeler.process(image) .addOnSuccessListener { labels -> val sceneLabel = labels.firstOrNull { it.text in SMART_HOME_SCENES } sceneLabel?.let { automationEngine.trigger(it.text, it.confidence) } } SMART_HOME_SCENES is a set of "bedroom", "kitchen", "living room", "bathroom", "office". For custom models—TFLite Interpreter with .tflite file, optimized via TensorFlow Model Maker. Personalized model on 500–1000 photos per class, fine-tuning MobileNetV2, export to INT8 quantized—model size ~2 MB, inference <50ms on Snapdragon 778G.
How to Implement Debounce for Scene Change?
Scene recognition is only a trigger. Next, automation logic without false triggers is needed. Pattern: scene change is counted only if one category dominates for 3 seconds with confidence > 0.7.
class SceneDebouncer(private val windowMs: Long = 3000) { private var currentScene: String? = null private var firstSeenAt: Long = 0 fun process(scene: String, confidence: Float): String? { if (confidence < 0.7f) return null val now = System.currentTimeMillis() if (scene != currentScene) { currentScene = scene firstSeenAt = now return null } return if (now - firstSeenAt >= windowMs) scene else null } } How to Control IoT via MQTT or Matter?
After scene confirmation, we publish a command to an MQTT broker or send through a Matter controller:
// MQTT mqttClient.publish( "home/automation/scene", MqttMessage("""{\"scene\":\"bedroom\",\"timestamp\":${System.currentTimeMillis()}}""".toByteArray()), qos = 1, retained = false ) // Matter SDK (via Google Home Mobile SDK) val deviceController = ChipDeviceController() deviceController.sendCommand( nodeId = lightbulbNodeId, endpointId = 1, clusterId = OnOffCluster.CLUSTER_ID, commandId = OnOffCluster.Commands.On.ID, tlvData = byteArrayOf() ) Schedule and Context
Scene-based automation should consider time of day: "bedroom" at 23:00 → dim lights, "bedroom" at 7:00 → open curtains. Context is added via TimeOfDay filter in rules at the application level.
Comparing CoreML vs TFLite
| Parameter | CoreML (iOS) | TFLite (Android) |
|---|---|---|
| Built-in model | VNClassifyImageRequest (3000+ classes) | ML Kit Scene Detection (5+ classes) |
| Fine-tuning | Create ML (MobileNetV3) | TensorFlow Model Maker (MobileNetV2) |
| Custom model size | ~4 MB | ~2 MB (INT8 quantized) |
| Inference time | <30 ms (Apple A14+) | <50 ms (Snapdragon 778G) |
| Privacy | Fully local | Fully local |
CoreML is 40% faster on comparable devices, but TFLite offers greater flexibility for cross-platform development.
Step-by-Step Guide for Integrating Scene Recognition
- Define target scenes (e.g., "bedroom", "kitchen") and IoT devices.
- Choose platform: CoreML for iOS, TFLite for Android, or cross-platform Flutter/RN.
- Integrate the ML model and configure the local classifier.
- Implement debounce logic to avoid false triggers.
- Connect MQTT or Matter for sending commands.
- Test under different lighting conditions and camera angles.
More on custom models
For fine-tuning the model, use Transfer Learning: freeze the first layers of MobileNetV2/V3, add a head for your classes. Each scene needs 200–500 labeled frames. Optimization via Quantization Aware Training reduces model size by half without loss of accuracy.
Why Privacy Matters When Using a Smart Home Camera?
An app with constant camera access is a red flag for users and App Store/Google Play moderators. Rules:
- Classification only when the user explicitly enabled "Scene Detection" mode.
- No frames are saved or leave the device.
- On iOS—
NSCameraUsageDescriptionwith explicit explanation of local processing. - Privacy manifest in iOS 17+ with declaration of
NSPrivacyAccessedAPICategoryCamera.
App Store rejections for 4.3 Spam or privacy violations are a real risk. The description in App Privacy Report must be honest.
What Is Included in Development
- Requirements audit: target devices, IoT protocols (MQTT, Matter, Zigbee via hub, HomeKit), set of trigger scenes.
- Classification model development: built-in or custom with fine-tuning.
- Integration with MQTT broker or Matter SDK.
- Implementation of debounce and automation logic.
- Testing in real conditions—different lighting, camera angles, mixed scenes.
- Documentation and source code handover.
- Post-launch support.
We guarantee reliability and stability of the solution, based on years of experience. The cost is calculated individually, but we guarantee a transparent budget with no hidden fees.
Stages and Timelines
| Stage | Timeline |
|---|---|
| Audit and requirements agreement | 3–5 days |
| Model development (built-in) | 1–2 weeks |
| IoT protocol integration | 1–3 weeks |
| Automation logic and testing | 1–2 weeks |
| Full cycle with custom ML model | 2–3 months |
Basic recognition with 5–10 scenes and MQTT commands: 2–4 weeks. Custom ML model with fine-tuning + full Matter/HomeKit integration: 2–3 months. Cost is calculated individually, depending on the number of supported IoT protocols and automation logic complexity.
Order turnkey development—we will assess your project in 1 day and offer the optimal solution. Contact us to discuss the details.







