How Smart Reply Solves the Speed Problem in Chat Communication?
We integrate Smart Reply into messengers, CRMs, and e-commerce platforms, and we see how this feature dramatically accelerates conversations. The user gets three ready-made reply options below each message — one tap is enough. Without Smart Reply, they must manually type "Okay", "Got it", "Thanks". This slows down the dialogue and increases the number of unfinished messages. Research by Google ML Kit shows that the feature boosts engagement and sent message volume by 15–30%. For businesses, this means faster query processing and higher customer loyalty. In our practice, clients save 2–3 seconds per reply, which accumulates to hours saved per day for the entire support team. In one project for an online store, implementing Smart Reply reduced the average operator response time from 45 to 15 seconds, and the number of messages per dialogue increased by 20%. If you are evaluating Smart Reply implementation, contact us for a preliminary assessment.
Why ML Kit Smart Reply Doesn't Fit Russian Language?
Google's ready-made SmartReply model works only with English. For Android, integration takes an hour:
val smartReply = SmartReply.getClient()
val conversation = messages.takeLast(10).map { msg ->
if (msg.isFromUser) {
TextMessage.createForLocalUser(msg.text, msg.timestamp)
} else {
TextMessage.createForRemoteUser(msg.text, msg.timestamp, msg.senderId)
}
}
smartReply.suggestReplies(conversation)
.addOnSuccessListener { result ->
if (result.status == SmartReplySuggestionResult.STATUS_SUCCESS) {
val suggestions = result.suggestions.map { it.text }
showSuggestions(suggestions)
}
}
.addOnFailureListener { /* hide UI */ }
Plus side: speed <20 ms. Minus: limited template set and only English. For iOS, the alternative via Natural Language or Apple Intelligence API exists but offers poorer functionality. If your app targets Russian-speaking audience, ML Kit is not suitable — a custom model is needed. Also consider App Store Review Guidelines: section 5.1 permits on-device ML without special permission, but custom LLMs processing user data must comply with privacy rules.
How to Implement Contextual Smart Reply on a Custom LLM?
For Russian and specific domains (support, healthcare, B2B), we use an LLM with a prompt. Example in Swift:
func generateReplySuggestions(
lastMessages: [ChatMessage],
count: Int = 3
) async -> [String] {
let context = lastMessages.suffix(5)
.map { "\($0.role): \($0.text)" }
.joined(separator: "\n")
let prompt = """
You help the user quickly reply to a chat message.
Dialogue history:
\(context)
Suggest \(count) short reply options for the user.
Each reply is one sentence, maximum 10 words.
Format: JSON array of strings.
"""
let response = try await llmClient.complete(prompt: prompt, maxTokens: 100)
return parseJSONArray(response) ?? []
}
A similar implementation on Android with Kotlin and ML Kit or a custom model:
suspend fun generateReplySuggestions(
lastMessages: List<ChatMessage>,
count: Int = 3
): List<String> {
val context = lastMessages.takeLast(5)
.joinToString("\n") { "${it.role}: ${it.text}" }
val prompt = """
You help the user quickly reply to a chat message.
Dialogue history:
$context
Suggest $count short reply options for the user.
Each reply is one sentence, maximum 10 words.
Format: JSON array of strings.
"""
val response = llmClient.complete(prompt, maxTokens = 100)
return parseJSONArray(response) ?: emptyList()
}
Latency of 1–2 seconds is acceptable. We preload options while the user reads — by the time they're ready to reply, suggestions are ready. Custom LLM is 50x slower than ML Kit but provides Russian language and context flexibility. On-device models (TensorFlow Lite, Core ML) are faster but require more memory and are less configurable.
What to Choose: ML Kit or Custom LLM?
| Criterion | ML Kit Smart Reply | Custom LLM |
|---|---|---|
| Russian language support | No | Yes |
| Latency | <20 ms | 1–2 sec |
| Customization | Low | High (prompt, context) |
| Network dependency | No | Yes |
| Integration complexity | Low (1 day) | Medium (5–8 days) |
| Cost | Free | API or inference costs |
For English and simple scenarios, ML Kit is better (50x faster). For Russian and specific needs, custom model.
How Does Smart Reply Differ on iOS and Android?
| Platform | Stack | Native Smart Reply | Custom Smart Reply |
|---|---|---|---|
| iOS | Swift/SwiftUI | NaturalLanguage (limited) | LLM + CoreML |
| Android | Kotlin/Compose | ML Kit (English only) | LLM + TensorFlow Lite |
When to Show Smart Reply?
Smart Reply appears after an incoming message and disappears when the user starts typing. Three suggestions is optimal (Google Research). More overloads, less gives no choice. We use chips (horizontal scroll): MaterialChip on Android, custom Chip in SwiftUI.
// Android: hide when typing
editText.addTextChangedListener(object : TextWatcher {
override fun onTextChanged(s: CharSequence?, start: Int, before: Int, count: Int) {
smartReplyChips.isVisible = s.isNullOrEmpty()
}
override fun afterTextChanged(s: Editable?) {}
override fun beforeTextChanged(s: CharSequence?, start: Int, count: Int, after: Int) {}
})
Also important to set up analytics: track how many users use suggested replies, and A/B test the number of chips.
How Does the Implementation Process Work?
We work in stages:
- Analyze Smart Reply usage scenarios in your app.
- Choose approach: ML Kit or custom LLM.
- Design architecture (preloading, caching).
- Implement on Android (Kotlin/Compose) and iOS (Swift/SwiftUI).
- Integrate with chat and analytics system.
- Code documentation and instructions for your team.
- Test on real dialogues.
- Post-implementation support: bug fixes and refinements.
Deliverables include:
- Integration and setup documentation.
- Access to test environment for validation.
- Team training (1–2 hours) on using the solution.
- One month of technical support after launch.
What Are the Timelines and Results?
Smart Reply via ML Kit (Android, English) — 1–2 days. Custom on LLM with context classification — 5–8 days. Full integration on both platforms — up to 2 weeks. Based on our projects, implementation increases engagement by 20–40% and reduces average response time by 2–3 times. For a support team of 10 people, time savings amount to up to 500 person-hours per month, converting to financial savings of $3,000 to $10,000 monthly. Contact us to discuss your app's needs.
Our Experience
We have implemented Smart Reply for messengers, CRMs, and e-commerce. We work with ML Kit and custom models. Five years in the market, over 50 projects. We guarantee correct operation on Android and iOS. Request a consultation — we'll evaluate your scenario and propose the optimal solution. Get demo access to a working example today.







