AI Bot for IoT: Function Calling and Streaming
Imagine you're away from the office and a temperature sensor jumps 3°C in 10 minutes. Usually, this requires opening a web panel, finding the sensor, plotting a graph, and checking logs—an operation that takes at least 5 minutes. We build an AI assistant that answers such a question in chat within 2 seconds: it fetches the needed data via API automatically. Our experience in iOS and Android exceeds 10 years, allowing us to integrate tool use without performance loss and with guaranteed security. The cost of one GPT-4o Function Calling call is about $0.03, which at 1000 requests per day gives $30 per month for AI logic, but the savings on manual operator work can reach $5000 per month for 50 sensors. This results in a net saving of up to $4970 per month. Order a prototype in 1 day—we'll demonstrate it on your data.
Problems We Solve
Typical scenarios where manual monitoring slows down decision-making:
- Long anomaly root cause search. An operator spends an average of 5 minutes clicking through dashboards. The AI chatbot reduces this to 2-3 seconds, instantly providing context.
- Inability to handle a large number of sensors. With 1000+ IoT devices, a human cannot monitor each one—the bot automatically checks anomalies (95% accuracy) and sends alerts.
- Limited context in chat. Ordinary bots don't understand "show the latest CO2 spike in the server room"—they need exact IDs. Our AI assistant uses a system prompt to know device names and their aliases.
Function Calling Mechanism in a Mobile App
The key mechanism is Function Calling (Tool Use) in OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet. The OpenAI Function Calling Documentation describes the protocol: the model generates structured tool_calls, the mobile app executes them (querying the IoT backend) and returns the result. The loop repeats until the model obtains all data needed for the answer.
// Android: handling tool_calls from GPT-4o
data class ChatMessage(
val role: String, // user, assistant, tool
val content: String? = null,
val toolCalls: List<ToolCall>? = null,
val toolCallId: String? = null,
val name: String? = null
)
class IoTChatRepository(
private val openAiApi: OpenAiApi,
private val iotApi: IoTDeviceApi
) {
private val tools = listOf(
Tool(
type = "function",
function = ToolFunction(
name = "get_sensor_readings",
description = "Get current and historical readings from IoT sensors",
parameters = JsonObject(mapOf(
"sensor_ids" to JsonArray(listOf(JsonPrimitive("string"))),
"from_timestamp" to JsonPrimitive("ISO8601 datetime"),
"to_timestamp" to JsonPrimitive("ISO8601 datetime"),
"aggregation" to JsonPrimitive("avg|min|max|last")
))
)
),
Tool(
type = "function",
function = ToolFunction(
name = "get_device_alerts",
description = "Get active or historical alerts for devices",
parameters = JsonObject(mapOf(
"device_ids" to JsonArray(),
"severity" to JsonPrimitive("critical|warning|info"),
"limit" to JsonPrimitive("integer")
))
)
)
)
suspend fun chat(userMessage: String, history: List<ChatMessage>): Flow<String> = flow {
val messages = history + ChatMessage(role = "user", content = userMessage)
var response = openAiApi.chatCompletion(messages, tools)
// Loop to execute tool_calls
while (response.toolCalls != null) {
val toolResults = response.toolCalls!!.map { call ->
val result = when (call.function.name) {
"get_sensor_readings" -> iotApi.getSensorReadings(call.function.arguments)
"get_device_alerts" -> iotApi.getAlerts(call.function.arguments)
else -> """{"error": "unknown tool"}"""
}
ChatMessage(role = "tool", content = result, toolCallId = call.id, name = call.function.name)
}
val updatedMessages = messages + ChatMessage(role = "assistant", toolCalls = response.toolCalls) + toolResults
response = openAiApi.chatCompletion(updatedMessages, tools)
}
emit(response.content ?: "")
}
}
Details of the tool call loop
The loop may continue for several iterations. It is important to limit the maximum number of tool_calls (e.g., 5) to avoid infinite loops. Also add a timeout for each API call (2 seconds).
Why Streaming Matters for UX
Without streaming, the user waits 5-10 seconds for a response, which is critical for emergency alerts. GPT-4o supports Server-Sent Events—tokens appear as they are generated. On iOS: URLSession.AsyncBytes, on Android: @Streaming in Retrofit. Typical first token latency is 200 ms, or 300 ms over 4G. Streaming reduces perceived time to 0.5-1 second, allowing the operator to correct the query quickly if needed.
// iOS: streaming from OpenAI SSE
func streamResponse(messages: [ChatMessage]) -> AsyncThrowingStream<String, Error> {
AsyncThrowingStream { continuation in
Task {
var request = URLRequest(url: URL(string: "https://api.openai.com/v1/chat/completions")!)
request.httpMethod = "POST"
request.setValue("Bearer \(apiKey)", forHTTPHeaderField: "Authorization")
request.httpBody = try JSONEncoder().encode(ChatRequest(messages: messages, stream: true))
let (bytes, _) = try await URLSession.shared.bytes(for: request)
for try await line in bytes.lines {
guard line.hasPrefix("data: "), line != "data: [DONE]" else { continue }
let json = line.dropFirst(6)
if let chunk = try? JSONDecoder().decode(StreamChunk.self, from: Data(json.utf8)),
let delta = chunk.choices.first?.delta.content {
continuation.yield(delta)
}
}
continuation.finish()
}
}
}
Solution Comparison
| Parameter | Ordinary bot (keyword search) | Bot with Tool Use |
|---|---|---|
| Answer accuracy | ~60% (depends on keywords) | ~95% (real API calls) |
| Data retrieval time | 10-30 sec (web scraping) | 1-3 sec (direct API) |
| Support for complex queries | No ("compare over a month") | Yes (model decides aggregations) |
| Token cost | Low (only prompt) | Higher (~2x), but saves manual work |
Our solution built with Swift and Kotlin processes requests 3 times faster than ordinary chatbots without IoT API integration.
Local Model as Fallback
| Scenario | Cloud Model (GPT-4o) | Local Model (Phi-3 Mini) |
|---|---|---|
| Cost per request | ~$0.03 | < $0.001 (only electricity) |
| Tool Use support | Full | Limited (only simple queries) |
| Availability | Internet required | Fully offline |
| Latency | 200-400 ms (first token) | 1-2 sec (local inference) |
When offline or to reduce costs—use llama.cpp with Phi-3 Mini or Mistral 7B via android-llamacpp or LLM.swift. Local models do not fully support tool use but answer basic questions from cached data.
Context and Security
The system prompt provides context: a list of the user's devices with names and IDs, time zone, units. This allows the bot to understand "sensor in the boiler room" without explicit IDs.
Important: IoT API functions are invoked on behalf of the current user with their access rights. The bot cannot retrieve data from devices the user does not have access to—authorization is at the backend level, not at the prompt level. According to our audit, 30% of vulnerabilities are related to missing permission checks at the API level.
Chat history—the last 20-30 messages in context. Older messages are compressed via summarization: gpt-4o-mini with the prompt "Summarize this conversation history briefly"—saves up to 40% tokens.
Process
- Analysis: Study your IoT infrastructure, sensor types, polling frequency, access rights. Prepare the function calling specification.
- Design: Develop tool_calls schema, system prompt, streaming architecture. Align with your team.
- Implementation: Write integration code on iOS (Swift) and Android (Kotlin), configure SSE streaming, add chat UI.
- Testing: Run integration tests with real data, verify anomaly detection (at least 95% recall).
- Deploy: Publish to App Store / Google Play, set up monitoring and alerts.
Timeline and Cost
Developing an AI chatbot for IoT monitoring with tool use, streaming, and integration with your IoT API: 4 to 6 weeks on top of an existing mobile app. Cost is calculated individually based on integration complexity and number of tool_calls. Our team with 10+ years of mobile development experience has completed 50+ projects involving IoT and AI—we guarantee a transparent process and post-delivery support. Get a consultation from an engineer and exact timeline.
What's Included
- Architectural documentation (tool_calls schemas, system prompt)
- Source code for chat module in Swift and Kotlin
- Integration with your IoT API (up to 5 endpoints)
- Streaming setup and error handling
- Test scenarios and automated tests
- 1 month of support after deployment
Common Mistakes When Developing an AI Bot for IoT
- Forgetting to handle tool_call errors. If the API returns an error, the model may get confused. Return a structured response with an error field.
- Not considering context limits. Too many messages cause the model to lose focus. Use summarization after 20-30 messages.
- Missing authorization at the function level. The system prompt won't protect against accessing other users' data. Always check permissions on the backend.
- Ignoring streaming delays on weak networks. SSE may have pauses—include a loading indicator.
Ready to discuss your project? Contact us for a free assessment—we'll show a prototype on your data in 1 day.







