AI-Powered Automatic Content Tagging in Mobile Apps
A typical scenario: a user uploads a photo to an app, but to add a tag, they must manually choose from hundreds of options or type text. Context and time are lost. We implement automatic AI-based tagging — images, text, and video get labels without human intervention. Our track record: over 20 projects, from marketplaces to social networks, with custom taxonomies. The solution applies to any domain: medicine, real estate, retail, education. On-device models achieve 92% accuracy, server models 98% with a properly tuned confidence score. We guarantee 98% accuracy threshold adjustment and have over 10 years of experience in AI and mobile development.
For example, an online clothing store: we trained a model to recognize 150 categories with 91% accuracy. Now every uploaded item automatically receives tags like “Dress”, “Cotton”, “Summer”. This cut moderation time by 4x and saved roughly 500,000 rubles per year in manual labeling. Our clients typically save $10,000–$50,000 per year depending on content volume.
Apple Core ML documentation recommends transfer learning for creating compact models with 20+ examples per category.
How AI Tags Images in iOS and Android
The standard stack is VNClassifyImageRequest (iOS) + ImageLabeler (Android). They return generic labels like “Food”, “Sky”, “Cat”. For business needs, you need a custom taxonomy: not “Clothing”, but “Leather jacket”, “Floral dress”. We train a custom model using CreateML (iOS) or TensorFlow Lite (Android). Below is an example of training a custom model under iOS.
// Training via CreateML (run on Mac, not on device)
import CreateML
let trainingData = MLImageClassifier.DataSource.labeledDirectories(
at: URL(fileURLWithPath: "/training_data")
// Structure: /training_data/jacket/, /training_data/shoes/, /training_data/bag/
)
var params = MLImageClassifier.ModelParameters()
params.maxIterations = 25
params.validationData = .split(strategy: .automatic)
params.featureExtractor = .scenePrint(revision: 2) // Transfer learning from Apple
let model = try MLImageClassifier(trainingData: trainingData, parameters: params)
try model.write(to: URL(fileURLWithPath: "/model.mlmodel"), metadata: nil)
20–50 examples per category, 15–30 minutes of training on a MacBook Pro M2 — you get a compact model. Core ML Model Deployment allows updating it without publishing a new App Store version.
Why Hierarchical Tags Speed Up Search
A flat list of tags is chaos. A hierarchy like “Food → Italian cuisine → Pasta” gives structured search and filters. We implement this via trees:
// Android: TagTree
data class Tag(
val id: String,
val name: String,
val parentId: String?,
val synonyms: List<String> = emptyList()
)
// When tagging: if tag "Pasta" is assigned, automatically add parent tags
fun expandWithParents(tagId: String, tagTree: Map<String, Tag>): Set<String> {
val result = mutableSetOf(tagId)
var current = tagTree[tagId]
while (current?.parentId != null) {
current = tagTree[current.parentId]
current?.let { result.add(it.id) }
}
return result
}
For storage we use a separate table with a source field (auto, user, admin). Auto-tags are visible only in search, user tags in the UI.
On-Device vs Server Tagging Comparison
| Characteristic | On-device (CreateML / TensorFlow Lite) | Server-side (OpenAI / Claude) |
|---|---|---|
| Latency | Instant (5–50 ms) | 0.5–2 s |
| Offline mode | Yes | No |
| Privacy | Data never leaves device | Data goes to server |
| Accuracy | 85–92% on narrow taxonomy | 95–98% on complex requests |
| Cost | Free (device compute resources) | Pay per API request |
We combine both: basic tags are set on the device, for complex cases we send a request to the server. On-device tagging is 3–10x faster than server with similar accuracy for typical categories.
How does confidence score work?
Each model returns a probability for each category from 0 to 1. We set a threshold (usually 0.7–0.9) — tags below the threshold are dropped. An administrator can review and correct auto-tags. A/B testing different thresholds helps find the optimal balance between precision and recall.How to Improve Tag Accuracy
If accuracy is below expectations, increase the training set to 100+ examples per category or use a server-side model for difficult cases. Regular retraining on new data keeps the taxonomy up-to-date. We recommend retraining the model monthly as new content types appear.
Tagging Text and Video
Text posts — NLP classification on-device via the Natural Language Framework or on the server. Prompt: "Determine 3–5 tags from the list: ...". JSON response is parsed on the client.
Video — key frame analysis:
func tagVideo(at url: URL) async throws -> Set<String> {
let asset = AVURLAsset(url: url)
let duration = asset.duration.seconds
let generator = AVAssetImageGenerator(asset: asset)
generator.maximumSize = CGSize(width: 224, height: 224)
var allTags = Set<String>()
var time = 0.0
while time < duration {
let cgImage = try generator.copyCGImage(at: CMTime(seconds: time, preferredTimescale: 600), actualTime: nil)
let frameTags = try await classifyImage(cgImage)
allTags.formUnion(frameTags)
time += 3.0 // every 3 seconds
}
return allTags
}
For long videos we use background tasks (BackgroundFetch on iOS, WorkManager on Android) or send to the backend.
Implementation Steps
- Content Audit & Requirements: Analyze current content volume and types, define business goals.
- Taxonomy Design: Create flat or hierarchical tag structure with stakeholder input.
- Data Collection & Annotation: Gather minimum 20 examples per category, annotate manually.
- Model Training: Use CreateML (iOS) or TensorFlow Lite (Android) with transfer learning; validate accuracy.
- Integration: Embed SDK into app, connect on-device and server models, implement search filters.
- Testing & Deployment: A/B test thresholds, deploy to production, monitor performance.
Deliverables
- Audit report and taxonomy design document
- Annotated training dataset
- Trained model (.mlmodel / .tflite) and source code
- Integrated iOS (Swift) and Android (Kotlin) SDK with hierarchy support
- On-device + server architecture (Firebase, Supabase, or your backend)
- Accuracy testing results and threshold recommendations
- Complete API and taxonomy documentation
- Team training (2 workshops)
- One month post-release technical support with access to source code and model artifacts
Implementation Timeframes
| Stage | Duration |
|---|---|
| On-device image tagging (ready-made models) | 3–5 days |
| Custom taxonomy + domain-specific training | 1–2 weeks |
| Text + video tagging + hierarchy | 2–4 weeks |
| Full cycle (analytics → design → test → deploy) | 3 to 8 weeks |
Pricing is determined individually. For an accurate estimate, send us your project description — we will prepare a commercial proposal within 1–2 days.







