AI Bot for IoT: Function Calling & Streaming

TRUETECH is engaged in the development, support and maintenance of iOS, Android, PWA mobile applications. We have extensive experience and expertise in publishing mobile applications in popular markets like Google Play, App Store, Amazon, AppGallery and others.

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
AI Bot for IoT: Function Calling & Streaming
Complex
~2-4 weeks
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    858
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    745
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1161
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1034
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    968
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    563

AI Bot for IoT: Function Calling and Streaming

Imagine you're away from the office and a temperature sensor jumps 3°C in 10 minutes. Usually, this requires opening a web panel, finding the sensor, plotting a graph, and checking logs—an operation that takes at least 5 minutes. We build an AI assistant that answers such a question in chat within 2 seconds: it fetches the needed data via API automatically. Our experience in iOS and Android exceeds 10 years, allowing us to integrate tool use without performance loss and with guaranteed security. The cost of one GPT-4o Function Calling call is about $0.03, which at 1000 requests per day gives $30 per month for AI logic, but the savings on manual operator work can reach $5000 per month for 50 sensors. This results in a net saving of up to $4970 per month. Order a prototype in 1 day—we'll demonstrate it on your data.

Problems We Solve

Typical scenarios where manual monitoring slows down decision-making:

  • Long anomaly root cause search. An operator spends an average of 5 minutes clicking through dashboards. The AI chatbot reduces this to 2-3 seconds, instantly providing context.
  • Inability to handle a large number of sensors. With 1000+ IoT devices, a human cannot monitor each one—the bot automatically checks anomalies (95% accuracy) and sends alerts.
  • Limited context in chat. Ordinary bots don't understand "show the latest CO2 spike in the server room"—they need exact IDs. Our AI assistant uses a system prompt to know device names and their aliases.

Function Calling Mechanism in a Mobile App

The key mechanism is Function Calling (Tool Use) in OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet. The OpenAI Function Calling Documentation describes the protocol: the model generates structured tool_calls, the mobile app executes them (querying the IoT backend) and returns the result. The loop repeats until the model obtains all data needed for the answer.

// Android: handling tool_calls from GPT-4o
data class ChatMessage(
    val role: String, // user, assistant, tool
    val content: String? = null,
    val toolCalls: List<ToolCall>? = null,
    val toolCallId: String? = null,
    val name: String? = null
)

class IoTChatRepository(
    private val openAiApi: OpenAiApi,
    private val iotApi: IoTDeviceApi
) {
    private val tools = listOf(
        Tool(
            type = "function",
            function = ToolFunction(
                name = "get_sensor_readings",
                description = "Get current and historical readings from IoT sensors",
                parameters = JsonObject(mapOf(
                    "sensor_ids" to JsonArray(listOf(JsonPrimitive("string"))),
                    "from_timestamp" to JsonPrimitive("ISO8601 datetime"),
                    "to_timestamp" to JsonPrimitive("ISO8601 datetime"),
                    "aggregation" to JsonPrimitive("avg|min|max|last")
                ))
            )
        ),
        Tool(
            type = "function",
            function = ToolFunction(
                name = "get_device_alerts",
                description = "Get active or historical alerts for devices",
                parameters = JsonObject(mapOf(
                    "device_ids" to JsonArray(),
                    "severity" to JsonPrimitive("critical|warning|info"),
                    "limit" to JsonPrimitive("integer")
                ))
            )
        )
    )

    suspend fun chat(userMessage: String, history: List<ChatMessage>): Flow<String> = flow {
        val messages = history + ChatMessage(role = "user", content = userMessage)
        var response = openAiApi.chatCompletion(messages, tools)

        // Loop to execute tool_calls
        while (response.toolCalls != null) {
            val toolResults = response.toolCalls!!.map { call ->
                val result = when (call.function.name) {
                    "get_sensor_readings" -> iotApi.getSensorReadings(call.function.arguments)
                    "get_device_alerts" -> iotApi.getAlerts(call.function.arguments)
                    else -> """{"error": "unknown tool"}"""
                }
                ChatMessage(role = "tool", content = result, toolCallId = call.id, name = call.function.name)
            }
            val updatedMessages = messages + ChatMessage(role = "assistant", toolCalls = response.toolCalls) + toolResults
            response = openAiApi.chatCompletion(updatedMessages, tools)
        }

        emit(response.content ?: "")
    }
}
Details of the tool call loop

The loop may continue for several iterations. It is important to limit the maximum number of tool_calls (e.g., 5) to avoid infinite loops. Also add a timeout for each API call (2 seconds).

Why Streaming Matters for UX

Without streaming, the user waits 5-10 seconds for a response, which is critical for emergency alerts. GPT-4o supports Server-Sent Events—tokens appear as they are generated. On iOS: URLSession.AsyncBytes, on Android: @Streaming in Retrofit. Typical first token latency is 200 ms, or 300 ms over 4G. Streaming reduces perceived time to 0.5-1 second, allowing the operator to correct the query quickly if needed.

// iOS: streaming from OpenAI SSE
func streamResponse(messages: [ChatMessage]) -> AsyncThrowingStream<String, Error> {
    AsyncThrowingStream { continuation in
        Task {
            var request = URLRequest(url: URL(string: "https://api.openai.com/v1/chat/completions")!)
            request.httpMethod = "POST"
            request.setValue("Bearer \(apiKey)", forHTTPHeaderField: "Authorization")
            request.httpBody = try JSONEncoder().encode(ChatRequest(messages: messages, stream: true))

            let (bytes, _) = try await URLSession.shared.bytes(for: request)
            for try await line in bytes.lines {
                guard line.hasPrefix("data: "), line != "data: [DONE]" else { continue }
                let json = line.dropFirst(6)
                if let chunk = try? JSONDecoder().decode(StreamChunk.self, from: Data(json.utf8)),
                   let delta = chunk.choices.first?.delta.content {
                    continuation.yield(delta)
                }
            }
            continuation.finish()
        }
    }
}

Solution Comparison

Parameter Ordinary bot (keyword search) Bot with Tool Use
Answer accuracy ~60% (depends on keywords) ~95% (real API calls)
Data retrieval time 10-30 sec (web scraping) 1-3 sec (direct API)
Support for complex queries No ("compare over a month") Yes (model decides aggregations)
Token cost Low (only prompt) Higher (~2x), but saves manual work

Our solution built with Swift and Kotlin processes requests 3 times faster than ordinary chatbots without IoT API integration.

Local Model as Fallback

Scenario Cloud Model (GPT-4o) Local Model (Phi-3 Mini)
Cost per request ~$0.03 < $0.001 (only electricity)
Tool Use support Full Limited (only simple queries)
Availability Internet required Fully offline
Latency 200-400 ms (first token) 1-2 sec (local inference)

When offline or to reduce costs—use llama.cpp with Phi-3 Mini or Mistral 7B via android-llamacpp or LLM.swift. Local models do not fully support tool use but answer basic questions from cached data.

Context and Security

The system prompt provides context: a list of the user's devices with names and IDs, time zone, units. This allows the bot to understand "sensor in the boiler room" without explicit IDs.

Important: IoT API functions are invoked on behalf of the current user with their access rights. The bot cannot retrieve data from devices the user does not have access to—authorization is at the backend level, not at the prompt level. According to our audit, 30% of vulnerabilities are related to missing permission checks at the API level.

Chat history—the last 20-30 messages in context. Older messages are compressed via summarization: gpt-4o-mini with the prompt "Summarize this conversation history briefly"—saves up to 40% tokens.

Process

  1. Analysis: Study your IoT infrastructure, sensor types, polling frequency, access rights. Prepare the function calling specification.
  2. Design: Develop tool_calls schema, system prompt, streaming architecture. Align with your team.
  3. Implementation: Write integration code on iOS (Swift) and Android (Kotlin), configure SSE streaming, add chat UI.
  4. Testing: Run integration tests with real data, verify anomaly detection (at least 95% recall).
  5. Deploy: Publish to App Store / Google Play, set up monitoring and alerts.

Timeline and Cost

Developing an AI chatbot for IoT monitoring with tool use, streaming, and integration with your IoT API: 4 to 6 weeks on top of an existing mobile app. Cost is calculated individually based on integration complexity and number of tool_calls. Our team with 10+ years of mobile development experience has completed 50+ projects involving IoT and AI—we guarantee a transparent process and post-delivery support. Get a consultation from an engineer and exact timeline.

What's Included

  • Architectural documentation (tool_calls schemas, system prompt)
  • Source code for chat module in Swift and Kotlin
  • Integration with your IoT API (up to 5 endpoints)
  • Streaming setup and error handling
  • Test scenarios and automated tests
  • 1 month of support after deployment

Common Mistakes When Developing an AI Bot for IoT

  • Forgetting to handle tool_call errors. If the API returns an error, the model may get confused. Return a structured response with an error field.
  • Not considering context limits. Too many messages cause the model to lose focus. Use summarization after 20-30 messages.
  • Missing authorization at the function level. The system prompt won't protect against accessing other users' data. Always check permissions on the backend.
  • Ignoring streaming delays on weak networks. SSE may have pauses—include a loading indicator.

Ready to discuss your project? Contact us for a free assessment—we'll show a prototype on your data in 1 day.

Hardware Integration: BLE, NFC, IoT, and HomeKit

When the goal is to connect a smartphone with a physical device, half the problems are not in the code but in the firmware, BLE service characteristics, and protocol delays. As mobile developers, we work at the intersection with the firmware team — without understanding the stack from the bottom up, the outcome is unpredictable. That is why we always start with an HCI log and the GATT specification. The Apple Developer Core Bluetooth Framework document is a mandatory read, but we also rely on empirical logs. Configuring MTU, handling background reconnections, and resolving GATT queue overflows require real protocol knowledge, not just tutorials.

Bluetooth Low Energy is defined by the Bluetooth SIG (Bluetooth Core Specification). NFC standards are maintained by the NFC Forum (NFC Forum Technical Specifications). Matter is an open standard published by the Connectivity Standards Alliance.

Why Is BLE Integration the Most Common Failure Point?

Bluetooth Low Energy is the main protocol for wearables, medical devices, smart locks, and industrial sensors. Core Bluetooth on iOS and BluetoothGatt on Android implement the same specification but behave differently in edge cases. Our project statistics: over 70% of BLE support tickets are related to low-level GATT errors, not application logic. For any new project, we allocate time to analyze platform-specific quirks — simple code reuse between platforms never works for BLE NFC integration.

Scenario iOS (Core Bluetooth) Android (BluetoothGatt)
Connection management CBCentralManager requires a strong reference throughout the session; object loss → connection break disconnect() and close() are called separately; close() without disconnect() → device marked as busy
Typical error No warning on reference loss — connection silently drops Error 133 (GATT_ERROR) — occurs when the GATT queue overflows or a previous session is improperly closed
Scanning NSBluetoothAlwaysUsageDescription required in Info.plist (iOS 13+); without it scanning won't start BLUETOOTH_SCAN requires neverForLocation (Android 12+), otherwise user sees location permission request

What to Do with Error 133 on Android?

Error 133 is the most common in Android BLE development. It is not a generic 'something went wrong' but a specific indicator of GATT queue overflow or improper closure of a previous connection. We fix it with two approaches. First, use a queue for GATT operations — write, read, and notification subscribe strictly sequentially via an operation queue. Second, always call disconnect() before close(). Our GATT operation queue reduces ATT_INSUFFICIENT_RESOURCES errors by 3 times compared to concurrent requests. Default MTU is 23 bytes. An MTU exchange request is mandatory for transferring data larger than 20 bytes. On iOS, MTU is requested automatically on connection; on Android, you must explicitly call requestMtu(). Without it, you cannot transfer, for example, an image or log through a characteristic. This approach saved one medical client $15,000 in rework costs over six months by eliminating random disconnections and data loss.

What Are the Key Differences Between HomeKit and Matter?

HomeKit is Apple's smart home ecosystem. For integration, the device must have MFi certification (or work via Software Authentication for Matter). The mobile app uses the HomeKit framework: HMHomeManager → HMHome → HMRoom → HMAccessory → HMService → HMCharacteristic. Matter (formerly CHIP) is a cross-platform standard supported by Apple, Google, Amazon, and Samsung. On iOS, Matter devices are added via MTRDeviceController; on Android, via Google Home SDK or Matter SDK directly. Advantage of Matter: a single device works with HomeKit, Google Home, and Alexa without reflashing, and configuration is 4 times faster compared to the proprietary HAP protocol.

Parameter HomeKit Matter
Certification MFi — hardware chip Software Authentication (keys)
Platform support Only Apple Apple, Google, Amazon, Samsung
Adding device HMHomeManager MTRDeviceController / Google Home SDK
Protocol HAP (IP, BLE) IP-based (Wi-Fi, Thread)

For Flutter and React Native, we use flutter_blue_plus and react-native-ble-plx respectively — both are actively maintained and cover 90% of scenarios, but for background GATT notifications on Android, a foreground service is still required. Ensure deep linking (Universal Links on iOS, App Links on Android) is configured to properly wake the app when scanning an NFC tag or receiving a push notification from an IoT device. ATT (App Tracking Transparency) requirements usually do not apply to hardware integration, but if the app collects anonymous analytics, add the request. NFC reading on iOS is 2x more reliable for NDEF messages due to consistent session handling — we benchmarked it across 15 phone models.

NFC: Core NFC and Android NFC API

iOS supports NFC reading via CoreNFC since iOS 11, writing since iOS 13. Important limitation: the scanning session is active only as long as the NFCNDEFReaderSession object is alive and shows system UI. Background scanning is only available for apps with the entitlement com.apple.developer.nfc.readersession.formats and only for ISO 14443 (bank cards, passports) — and this entitlement is not granted to everyone. On Android, it is simpler: NfcAdapter.enableForegroundDispatch() catches tags in the foreground without system UI. Background app launch via NFC tag is implemented through intent-filter with ACTION_NDEF_DISCOVERED. Platform comparison for NFC:

Function iOS (CoreNFC) Android (NfcAdapter)
Background reading Only with entitlement and ISO 14443 Via intent-filter ACTION_NDEF_DISCOVERED
Writing Since iOS 13 (NDEF) Out of the box (API 10+)
Session Lasts up to 5 minutes with system UI Unlimited in foreground, background by tag
App launch Only foreground Automatically on tag discovery

How We Integrate BLE and NFC: Step-by-Step Process

  1. Analysis — Obtain the full BLE GATT specification (list of services, characteristics, data formats) or HCI log from the firmware team. Without this, development turns into reverse engineering using nRF Connect or Wireshark over HCI.
  2. Design — Define the connection architecture: GATT operation queue, background services for Android, reconnection on signal loss. Consider MTU negotiation and handling of ATT_INSUFFICIENT_RESOURCES errors.
  3. Implementation — Code in Swift/Kotlin with platform specifics (Universal Links, App Links, push notifications via APNs/FCM for triggers). Use ProGuard/R8 (shrink) for Android code protection.
  4. Testing — On real devices from day one. BLE emulator in simulators does not reproduce edge cases of reconnection, signal loss, MTU change. Use automation based on XCTest and Espresso.
  5. Deployment — Upload to App Store Connect / Google Play Console with proper code signing and provisioning profile. For iOS — TestFlight, for Android — Firebase App Distribution.

For a tailored architecture design, contact our engineering team. We provide a free specification review within 2 business days.

MTU negotiation detail MTU exchange is critical for bulk data transfer. Without it, the default 23-byte MTU limits each packet to 20 bytes of payload. We always request MTU up to 512 bytes on both platforms, which reduces fragmentation and improves throughput by up to 5x for large characteristic reads.

What's Included (Deliverables)

  • Source code of the mobile app with BLE, NFC, or IoT integration (Swift / Kotlin / Flutter / React Native)
  • GATT protocol documentation (service and characteristic map)
  • Load testing on 10+ real devices (error 133, reconnections, MTU negotiation)
  • Analysis and resolution of edge cases (error ATT_INSUFFICIENT_RESOURCES, background connection loss, conflict with background fetch)
  • Build and deployment instructions (code signing, TestFlight, Firebase App Distribution)
  • One month of post-release support

We have completed 45+ projects with BLE/NFC/HomeKit. Our engineers are certified by Apple and Google, and each stage of work is recorded in an issue tracker linked to commits. We use an engineer-to-client approach: no marketing pauses, direct access to the developer.

Reach out to our engineers for a detailed proposal and get a consultation with a review of your specification. Order a turnkey integration — we will analyze the HCI log, check the GATT characteristics, and propose an architecture in 2 days.