AI Recommendation System for Mobile Apps: Implementation Guide
The Problem: Slow or Irrelevant Recommendations
Recently, an e-commerce client approached us: their iOS app showed recommendations with a 2-second delay — users scrolled past before they loaded. The conversion rate in the recommendation block was 1.2%. We moved the final reranking to the device using CoreML — response time dropped to 50 ms, conversion rose to 4.1%. Savings on server resources: 40% (about $3,500/month) due to reduced requests. The cold start for new users was solved with an onboarding quiz (2 preference questions) and a popularity-based fallback. After 10 sessions, personalized recommendations worked.
We build AI recommendation systems not as black boxes but as pipelines: collecting behavioral events, feeding them into ML models, ranking, and embedding into the UI without performance loss. Our team has 7+ years of experience and 15+ projects for iOS and Android. A hybrid architecture outperforms pure server-side: CTR is 2–3 times higher with the same data.
How to Choose Architecture: On-Device or Server?
| Criterion | Server-Side | Client-Side (CoreML/TFLite) |
|---|---|---|
| Quality | High (sees all users) | Medium (only device) |
| Latency | Network delay | Instant, offline |
| Privacy | Data on server | Data on device |
| Model Updates | Once per day | Possible without release |
On-device reranking cuts latency by 3–5 times and saves up to 60% server resources (up to $4,000/month). Testing shows a hybrid approach improves CTR 2–3 times over pure server-side.
Why Event Collection Is the Foundation of Quality?
A recommendation system is only as good as its data. On mobile, you must log at minimum:
-
item_view— object view (with dwell time, not just impression) -
item_click— tap/click on object -
item_purchase/item_save— conversion action -
item_skip— scrolled past (important negative signal)
// Android: batched event logger class RecoEventLogger(private val api: RecoApi) { private val buffer = mutableListOf<RecoEvent>() private val flushInterval = 30_000L // 30 seconds fun log(event: RecoEvent) { buffer.add(event.copy(timestamp = System.currentTimeMillis())) if (buffer.size >= 20) flush() // or by timer } private fun flush() { if (buffer.isEmpty()) return val batch = buffer.toList() buffer.clear() viewModelScope.launch(Dispatchers.IO) { runCatching { api.sendEvents(batch) } // On error — write to Room for retry } } } Important: dwell time is often a missed signal. Track when a card enters the viewport (RecyclerView.OnScrollListener or LazyList.onVisibleItemsChanged) and when it leaves. A view under 2 seconds is probably a scroll-through. In one project, adding dwell time increased CTR by 18%.
How On-Device CoreML/TFLite Reranking Works?
If the server returns top-200 candidates, final ranking can happen on device. This eliminates an extra network request on every screen open.
On iOS with CoreML:
// Load model (bundled or via Core ML Model Deployment) let model = try MLModel(contentsOf: modelURL) let input = RerankerInput( userVector: userEmbedding, // Float32 array 64d itemVectors: itemEmbeddings, // [Float32 array 64d] sessionFeatures: sessionContext // last 10 actions ) let output = try model.prediction(from: input) let scores = output.featureValue(for: "scores")?.multiArrayValue TensorFlow Lite on Android uses Interpreter with ByteBuffer input. For models >10 MB, use GPU delegate (GpuDelegate) — acceleration of 3–8x on flagships.
Updating the model without an app release: on iOS — Core ML Model Deployment via CloudKit or custom CDN with MLModel.compileModel(at:). On Android — Firebase ML with RemoteModel or direct .tflite download into filesDir with hash verification.
Steps to Implement a Recommendation System
- Data & Event Audit — check which events are already logged, add missing ones (dwell time, skip).
- Architecture Selection — decide what lives on server vs. on device.
- Develop Event Tracker — with batching, retry mechanism, Room storage for offline.
- Server Model — collaborative filtering or a ready service (Amazon Personalize, Google Recommendations AI).
- Client Model Integration — CoreML/TFLite, reranking candidates.
- UI Components — adaptive blocks with lazy loading.
- A/B Testing — Firebase Remote Config, Amplitude Experiment.
- Documentation & 6-Month Guarantee.
What On-Device Reranking Delivers (Comparison)
| Parameter | Server Only | Hybrid (Server + On-Device) |
|---|---|---|
| Display latency | 200–500 ms | 20–50 ms |
| Number of requests | 1 per view | 1 per day |
| Server resource savings | — | up to 60% (up to $4,000/month) |
| Personalization quality | High | Very high (with session signals) |
How to Handle Cold Start?
First 5–10 sessions lack data for personalization. Standard approach — hybrid:
- Onboarding quiz (2–3 preference questions) gives initial profile.
- Popularity-based recommendations as fallback.
- Implicit feedback from first interactions quickly shifts profile.
Avoid showing “recommendations for you” until minimal history is collected — it’s fair to the user and keeps metric quality.
Which Quality Metrics to Track?
Click-through rate (CTR) and conversion are basic. But for mobile UX, also track “recommendation blindness”: if the block is ignored, it’s worse than low CTR. A/B testing via Firebase Remote Config or Amplitude Experiment is mandatory when changing algorithms. Minimum sample for statistical significance: 1000+ unique users per variant.
What’s Included (Deliverables)
- Technical documentation: event tracker architecture, data model.
- Source code of event tracker with batching and retry (Swift/Kotlin).
- Integration of server recommendation model (or custom).
- In-app UI recommendation component with lazy loading.
- A/B testing and metric monitoring setup.
- Team training on system usage.
- 6 months of technical support.
Timeline Guidelines
Integration of a ready server recommendation service with event tracker — 2–3 weeks. Hybrid system with on-device reranking, custom events, and A/B testing — 6–10 weeks. Cost is determined individually.
Contact us for a free audit of your app — we will assess your architecture and propose the optimal solution. Request implementation and get a consultation on model selection.







