Technical Challenge: From Simple Translation to Context and Offline
A mobile translation bot seems like a simple API call. But when you need to account for dialogue context, work in the subway, and translate through the camera, a simple request to DeepL or Google won't cut it. We—a mobile development team with 5 years of experience—have faced this many times. Choosing the wrong architecture leads to user churn: translations become nonsensical, the app crashes offline, or requires constant internet. In this article, we'll show you how to build a reliable translator with offline mode, voice input, and camera, using modern APIs and ML Kit. Over the years we have delivered 50+ projects for clients across the US, Europe, and CIS, including specialized translators for medical and legal domains. Our experience shows that the key decisions—which Translation API and offline architecture you pick—shape the product quality. We'll go step by step: from API selection to camera integration, so you can make an informed choice and avoid common pitfalls.
Mobile Translation Bot API Selection: DeepL vs Google vs Yandex vs LLM
DeepL, Google Cloud Translation, Yandex Translate, and LLMs each have their strengths. The choice depends on the language pair and context requirements.
DeepL API offers the best quality for European languages—30% better than Google on some tests (measured by BLEU score). Free tier: 500K characters per month. Supports formal/informal tone (formality parameter). Important: DeepL does not support translation from Russian into some languages—in such cases, use Google or Yandex.
Google Cloud Translation API covers 100+ languages, high quality for Russian. v3 supports a glossary—a list of terms that must not be translated or must be translated in a specific way. For medicine, legal texts, brands—this is a must.
Yandex Translate API delivers the best results for Russian ↔ European languages. Built-in language detection.
LLM (GPT-4o / Claude) provides contextual translation with tone and style awareness. It excels on specialized texts, idioms, and humor. More expensive by 2-3x for large volumes. Recommended for dialogue bots.
| API | Languages | Quality | Highlights |
|---|---|---|---|
| DeepL | 30+ European | Excellent (30% better than Google) | formality, free 500K chars |
| 100+ | High | glossary, autodetect | |
| Yandex | 100+ | High for RU | built-in detect |
| LLM | Any | Contextual (20-30% improvement) | tone, idioms |
Language Detection
The user inputs text—the bot needs to figure out the source language. Either a manual picker or auto-detection. Google Translation API returns detectedSourceLanguage. For short queries (1–2 words) auto-detection often fails—better to let the user choose manually.
Dialogue Context
If the user translates a series of related messages—an LLM with translation history gives more coherent results. Names, pronouns, and terms stay in context. To save tokens, pass the last N messages.
Terminology Glossary
Google Translation v3 Glossary API lets you create a list of terms that the model must not translate or must translate strictly:
from google.cloud import translate_v3
client = translate_v3.TranslationServiceClient()
glossary = client.create_glossary(
parent=f"projects/{project_id}/locations/us-central1",
glossary=translate_v3.Glossary(
name=glossary_name,
language_pair=translate_v3.Glossary.LanguageCodePair(
source_language_code="en",
target_language_code="ru"
),
input_config=translate_v3.GlossaryInputConfig(
gcs_source=translate_v3.GcsSource(input_uri=glossary_gcs_uri)
)
)
)
How to Implement Offline Translation?
For users in areas with unstable internet—translate on the device without sending text to the server. Data privacy is a key advantage. Offline models run locally—user text never leaves the device. They also speed up translation (10–50 ms vs. 200–500 ms for an API) and save up to 70% of costs under high load. For high-volume apps, offline mode can save thousands per month compared to cloud APIs. According to Google ML Kit documentation, language models weigh about 30 MB each.
iOS. Use MLKit Translation from Google. TranslateLanguage.allLanguages() lists available languages. Each model is ~30 MB. Before translating, check if the model is downloaded via translator.downloadModelIfNeeded(with:).
import MLKitTranslate
let options = TranslatorOptions(
sourceLanguage: .russian,
targetLanguage: .english
)
let translator = Translator.translator(options: options)
let conditions = ModelDownloadConditions(allowsCellularAccess: true)
translator.downloadModelIfNeeded(with: conditions) { error in
guard error == nil else { return }
translator.translate("Привет, мир") { result, error in
print(result ?? "")
}
}
Android. Same via TranslatorOptions from com.google.mlkit:translate. Offline models are quantized for efficiency—they run locally and never send data to the cloud.
Comparison: Online vs. Offline
| Parameter | Online (API) | Offline (ML Kit) |
|---|---|---|
| Languages | 100+ | 50+ |
| Quality | Maximum | Slightly lower (5–10%) |
| Latency | 200–500 ms | 10–50 ms (10x faster) |
| Privacy | Text leaves device | On-device only |
| Internet required | Yes | No |
| Cost per 100K chars | $20 | $0 (free) |
Voice Input and Translation Playback
A logical addition: user speaks → translator reacts → reads aloud. STT for input: native APIs (Android SpeechRecognizer, iOS SFSpeechRecognizer) or Whisper. TTS for output: AVSpeechSynthesizer on iOS (AVSpeechSynthesisVoice(language: "fr-FR")) or TextToSpeech on Android. Important: verify that the desired voice is available on the device before speaking.
Camera: Real-time Translation
The most impressive scenario—point the camera at a menu, sign, or document and see the translation overlaid on the image. Technically: ML Kit TextRecognizer → translate blocks → render on top of the camera preview with OCR bounding boxes. A hidden pitfall: OCR text coordinates are tied to the frame, which changes 30 times per second. Stabilizing results (comparing bounding boxes via IoU from the previous frame) reduces flickering.
Our Work Process
- Requirement analysis and API selection.
- Backend development: API keys, caching, glossary.
- Mobile UI: input, history, copy, share.
- Integrate offline models, voice, camera (optional).
- Testing and publishing (App Store / Google Play).
Deliverables
- Architecture decision and stack selection document.
- Backend and client-side source code.
- Glossary configuration and contextual translation setup.
- API integration documentation and user guides.
- Post-launch support (1 month).
- Training session for your team (1 day).
- Access to deployment scripts and performance benchmarks.
- User testing report with BLEU score evaluation.
Estimated Timelines and Pricing
A basic cloud-based translator: 2–3 days, from $1,000. With offline mode, voice input, and camera translation: 1.5–2 weeks, from $5,000. Full-featured app with all bells and whistles: up to $10,000. Accurate estimate after project analysis. We guarantee quality and adherence to deadlines—over 5 years of work, we have never missed a single deadline. Contact us to discuss your project. Order development of a mobile translation bot today.







