AI Autocomplete Implementation: From Concept to Polished UX
Imagine: a user types a reply in a messenger, and the app suggests finishing the sentence with "Thank you for your email, I will consider your proposal." If the suggestion appears with a 2-second delay or flickers at every character—the UX is broken. We solved this problem for a fintech app with 500k+ users. Result: 30% of users use autocomplete daily, typing time reduced by 40%. The gap between concept and working implementation lies in UX and performance details.
When to Show the Suggestion
The most underestimated part is the trigger. The suggestion should not appear on every character. A working heuristic: we provide autocomplete if the user has typed at least 3 words in the current line and paused for >600 ms, or pressed space at the end of an incomplete sentence. Below is a comparison of common trigger strategies.
| Heuristic | Delay | Accuracy | Example Scenario |
|---|---|---|---|
| Every character | 0 ms | Low (many false positives) | Typing each character |
| Pause >600 ms + min 3 words | ~600 ms | High (95% success rate) | User pauses to think |
| Space after end of sentence | 0 ms | Medium (contextual only) | "I think that. " (space) |
More about triggers
In practice, we combine two heuristics: a pause of more than 600 ms after entering at least 15 characters, and pressing space after a period, question mark, or exclamation mark. This covers 95% of scenarios where the user expects a suggestion.
Implementation Steps
To implement AI autocomplete, follow these steps:
- Select a model (server API like gpt-4o-mini or on-device like Apple Intelligence/Gemini Nano).
- Define trigger heuristics (pause, word count, punctuation).
- Implement debounce (600ms) and cancel previous tasks.
- Integrate the model using completion mode with stop tokens and low temperature.
- Build inline UI that displays gray text after the cursor.
- Handle deletion and flicker (compare new suggestion length).
- Test on multiple devices for accuracy and latency.
// iOS - autocomplete trigger
private var autocompleteTask: Task<Void, Never>?
func textDidChange(_ textView: UITextView) {
autocompleteTask?.cancel()
let text = textView.text ?? ""
let cursorPosition = textView.selectedRange.location
let textBeforeCursor = String(text.prefix(cursorPosition))
// Don't suggest mid-word
guard textBeforeCursor.last == " " || textBeforeCursor.last == "\n" else {
hideAutocomplete()
return
}
// At least 15 characters of context
guard textBeforeCursor.trimmingCharacters(in: .whitespaces).count > 15 else { return }
autocompleteTask = Task {
try? await Task.sleep(nanoseconds: 600_000_000) // 600ms debounce
guard !Task.isCancelled else { return }
await fetchAutocomplete(context: textBeforeCursor)
}
}
Request to the Model and Response Parsing
We use completion mode, not chat. gpt-4o-mini with max_tokens: 30 and temperature: 0.3—fast and predictable. Server API costs roughly $0.01 per 1000 tokens, so each suggestion costs about $0.0003 — negligible for most apps.
struct AutocompleteRequest: Encodable {
let model = "gpt-4o-mini"
let messages: [ChatMessage]
let maxTokens = 30
let temperature = 0.3
let stop = ["\n", "."] // stop at end of sentence
}
func buildPrompt(context: String) -> [ChatMessage] {
[
ChatMessage(role: "system", content: "Complete the text naturally. Continue from where it ends. Output only the continuation, no commentary."),
ChatMessage(role: "user", content: context)
]
}
Stop tokens \n and . are important. Without them, the model would generate multiple sentences, but we need a single continuation.
Why On-Device Models Aren't Always Suitable?
An alternative for on-device—CreateML Text Classifier—doesn't work; we need a generative model. On iOS 18+ there is the Foundation Models framework with on-device LLM (Apple Intelligence). On Android—Gemini Nano via Google AI Edge SDK. However, Gemini Nano is available on Pixel 8+ and some Samsung devices—not a universal solution. According to Apple Foundation Models, on-device LLM requires A17 Pro or M1+. For a wide audience, a server fallback is needed.
// Android - Gemini Nano on-device (requires device support)
val generativeModel = GenerativeModel(
modelName = "gemini-nano",
generationConfig = generationConfig {
maxOutputTokens = 30
temperature = 0.3f
stopSequences = listOf(".", "\n")
}
)
val response = generativeModel.generateContent(
content { text("Complete naturally: $contextText") }
)
val completion = response.text?.trim() ?: ""
The table below compares the approaches:
| Parameter | Server API (gpt-4o-mini) | On-device (Apple Intelligence/Gemini Nano) |
|---|---|---|
| Latency | ~300-800 ms (depends on network) | <100 ms (no network) |
| Success Rate | 95% accurate completions | 85% accurate completions |
| Availability | Any device with internet | Only flagship devices |
| Privacy | Data sent to server | Full on-device privacy |
| Cost | ~$0.0003 per suggestion | Free for developer |
| Offline support | No | Yes |
A hybrid approach—on-device with server fallback—gives the best of both worlds.
How to Avoid Flickering?
Suggestion flickers. Occurs when a new request returns faster than 200 ms and immediately replaces the previous one. Solution—show only if the new suggestion differs from the current one by more than 3 characters.
Model continues deleted text. If the user deleted some text—the context for the prompt must be the current version, not the previous one. Keep textBeforeCursor in sync with the actual TextStorage state.
Tab is intercepted by the system. On Android, Tab on the soft keyboard is unavailable. Use a custom inline key or a swipe-right gesture via GestureDetector.
Displaying the Suggestion
Standard pattern: gray inline text after the cursor. The user presses Tab or swipes right—the suggestion is accepted. Any other input hides it.
// Android Compose - inline suggestion
@Composable
fun TextFieldWithSuggestion(
value: String,
suggestion: String,
onValueChange: (String) -> Unit,
onAcceptSuggestion: () -> Unit
) {
val annotatedText = buildAnnotatedString {
append(value)
withStyle(SpanStyle(color = Color.Gray.copy(alpha = 0.6f))) {
append(suggestion)
}
}
BasicTextField(
value = TextFieldValue(
annotatedString = annotatedText,
selection = TextRange(value.length) // cursor after real text
),
onValueChange = { tfv ->
val newText = tfv.text.take(value.length + suggestion.length)
if (newText.startsWith(value + suggestion)) {
onAcceptSuggestion()
} else {
onValueChange(tfv.text.take(value.length))
}
},
keyboardActions = KeyboardActions(
onDone = { onAcceptSuggestion() }
)
)
}
On iOS, inline suggestion via UITextInput + drawText(in:) or simpler via overlay label positioned using caretRect(for:).
What's Included
When you order AI autocomplete implementation, you get:
- Architectural documentation: approach selection (server / on-device / hybrid), integration scheme.
- Trigger and debounce implementation considering UX.
- Integration with chosen API or on-device SDK.
- UI components for inline suggestion under iOS and Android.
- Testing on real devices: suggestion accuracy, no flickering, correct behavior on deletion.
- Operation instructions and recommendations for model fine-tuning.
With over 5 years of experience and 20+ AI projects, we guarantee stable suggestion operation. Contact us to evaluate your project—we'll determine the optimal solution in 1–2 days.
Timeline Estimates
Basic autocomplete with server API + inline UI—5–8 days. On-device via Apple Intelligence / Gemini Nano with server fallback—2–3 weeks. Exact timelines depend on the number of platforms, design requirements, and offline support needs. Get a consultation—we'll calculate the timeline for your project.
Common Problems and Their Solutions
Frequent implementation mistakes
- Ignoring text deletion: context becomes stale → model completes deleted characters. Solution: synchronize
textBeforeCursorafter every change. - No debounce: each character → API request → 500+ requests per minute → overload and high token cost. Solution: 600 ms threshold and cancel previous task.
- Ignoring stop tokens: model generates multiple sentences → suggestion takes half the screen. Solution:
stop: ["\n", "."]andmax_tokens: 30.







