Content-Based Recommendations: Implementation on iOS and Android
Imagine you launch a news app with thousands of articles, but new users have no reading history. You need to show relevant content without data from others. Content-Based Filtering (based on attributes) solves this from day one. We analyze each article’s metadata—tags, categories, authors, text—and build a user profile based on what they interact with. For privacy-sensitive apps, CB can run entirely on-device without sending data to the server. This is especially relevant for financial or medical apps. In this article, we break down the key components of a CB system: from embeddings to on-device recommendations, with code examples for iOS and Python. On-device CB reduces cloud computing costs to ~$1000 per month. For a catalog of 10,000 items with 384-dimensional embeddings, the index is only ~15 MB. The user profile weighs 1.2 KB — recommendations are computed in 5 ms on an iPhone 12.
When is Content-Based better than Collaborative Filtering?
Three scenarios where CB is preferable:
-
Niche content with rich metadata. Articles, recipes, travel routes — each item has a rich set of attributes (tags, categories, authors, locations). CF relies on the signal “users are similar”, but for niche content there may be too few such users, especially at launch.
-
Privacy-first architecture. CB can run entirely on-device — the user profile is stored locally, recommendations are built without sending data to the server. This is critical for apps handling sensitive data.
-
Long tail of content. A new article published an hour ago has no interaction history for CF. CB recommends it immediately once metadata is indexed.
How does on-device CB reduce costs?
On-device CB eliminates server-side computation. For a catalog of 10,000 items with 384-dimensional embeddings, the index takes ~15 MB. The user profile is 1.2 KB. Recommendations are computed in 5 ms on an iPhone 12. Average savings on cloud computing after adopting on-device CB are around $500–$1000 per month for a 10,000-item catalog. Data privacy is automatically ensured.
How does Content-Based work on-device?
For small catalogs (up to 50K items), the entire CB search can be moved to the device. The user profile is stored in UserDefaults, content embeddings are loaded on app start (JSON ~20 MB for 50K items × 384d float32). Recommendations are computed locally — no network requests, no latency.
Code: on-device CB search in Swift
class OnDeviceRecommender {
private let userProfileKey = "user_embedding_v2"
private var itemIndex: [(id: String, embedding: [Float])] = []
func loadItemIndex(from url: URL) {
let data = try! Data(contentsOf: url)
itemIndex = try! JSONDecoder().decode([(id: String, embedding: [Float])].self, from: data)
}
func getRecommendations(count: Int) -> [String] {
guard let profileData = UserDefaults.standard.data(forKey: userProfileKey),
let profile = try? JSONDecoder().decode([Float].self, from: profileData)
else { return popularItemIds(count: count) }
return itemIndex
.map { item in (item.id, cosineSimilarity(profile, item.embedding)) }
.sorted { $0.1 > $1.1 }
.prefix(count)
.map { $0.0 }
}
private func cosineSimilarity(_ a: [Float], _ b: [Float]) -> Float {
zip(a, b).map(*).reduce(0, +)
}
}
Embeddings in the index are updated on start or on a schedule. We use precomputed vectors from the server, which minimizes device load.
Core system: TF-IDF and text embeddings
For articles, descriptions, news — two approaches: TF-IDF for speed, sentence embeddings for quality. In practice, we use sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 — 278 MB, supports Russian. Each item becomes a 384-dimensional vector. According to TF-IDF, word frequency and inverse document frequency provide a simple but effective relevance measure.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('paraphrase-multilingual-MiniLM-L12-v2')
def embed_article(article: Article) -> np.ndarray:
text = f"{article.title}. {article.description}. {' '.join(article.tags)}"
return model.encode(text, normalize_embeddings=True)
def similarity(v1: np.ndarray, v2: np.ndarray) -> float:
return float(np.dot(v1, v2))
User profile — moving average
The user profile is a weighted average of the embeddings of content they interacted with. Recent interactions weigh more (exponential decay with coefficient 0.9):
def update_user_profile(profile: np.ndarray, new_item_embedding: np.ndarray,
interaction_weight: float, decay: float = 0.9) -> np.ndarray:
updated = decay * profile + (1 - decay) * interaction_weight * new_item_embedding
return updated / np.linalg.norm(updated)
Structured metadata: not just text
For product catalogs, text embeddings are supplemented with categorical features: category, brand, price range, color. The final vector is a concatenation of a normalized text embedding and one-hot/ordinal features with weights 0.7 and 0.3 respectively:
def build_item_vector(item: Product) -> np.ndarray:
text_emb = embed_text(f"{item.name} {item.description}")
cat_features = encode_categorical({
'category': item.category_id,
'brand': item.brand_id,
'price_range': bucket_price(item.price)
})
return np.concatenate([text_emb * 0.7, cat_features * 0.3])
Comparison: Content-Based vs Collaborative
| Parameter | Content-Based | Collaborative Filtering |
|---|---|---|
| Requires user history | No | Yes |
| Works with new content | Immediately | Only after interactions |
| Privacy | Local | Requires server |
| Quality for niche content | Good | Poor (sparsity) |
| Computational load | Low (on-device) | Medium (on server) |
| Infrastructure savings | Up to $1000/month | Depends on scale |
How to implement on-device CB on Swift: step-by-step
- Content indexing. On the server, compute embeddings for each item using sentence-transformers. Save to a JSON file with fields id and embedding.
- Index loading. On app launch, load the JSON into the OnDeviceRecommender array. For speed, use a binary format.
- Collect interactions. Each time a user opens or likes content, save the item embedding in UserDefaults with a timestamp.
- Update profile. On each interaction, recompute the moving average with exponential decay.
- Generate recommendations. Call getRecommendations when displaying the feed or in the background.
What is included in the work
We provide:
- Analysis of content structure and available metadata
- Selection of the optimal embedding model for language and domain
- Building the index and profile update mechanism
- Implementation of on-device or server-side solution
- Integration with existing backend
- Documentation and team training
Our experience: over 10 years in mobile development, 30+ projects with AI recommendations. Guaranteed support after deployment. Contact us for a project evaluation — we will select the optimal architecture for your case. Order the development of a recommendation system for your project.
Time estimates
| Scenario | Timeline |
|---|---|
| Server-side CB with precomputed embeddings and API | 1–1.5 weeks |
| On-device for iOS/Android with local index | 2–3 weeks |
| Hybrid with partial on-device processing | 3–4 weeks |
Get a consultation: tell us about your content and user scenarios, and we will propose a solution.







