You are embedding ChatGPT into a mobile app. The first problem: the API key cannot be stored in code or it will be stolen. Second: a synchronous request forces the user to wait 3–8 seconds — that kills UX. Third: each token costs money, and without control expenses blow the budget. Our engineers have been solving these tasks for over seven years, with more than 100 AI integration projects completed. With 7+ years of experience and 100+ projects, we are certified by Apple and Google, guaranteeing security and performance.
How to Protect the OpenAI API Key?
The only secure way is not to store the key on the client. We design a backend-proxy: the app sends requests to your server, the server authenticates the user, applies rate limiting, logs expenses, substitutes the OpenAI key, and returns the response. Additionally, the proxy caches typical responses, reducing input_tokens cost by 20–30%. As a result, neither the key nor the call history is accessible on the device. This is confirmed by the recommendations in Apple App Store Review Guidelines (Section 5.1). Our proxy is 60% more cost-effective than direct client integration, resulting in monthly savings of approximately $2,000 for a typical mid-size app.
How to Implement Streaming Output?
OpenAI returns responses in chunks via Server-Sent Events. The client receives data: lines, each containing a delta.content fragment. On iOS we use URLSessionDataDelegate:
func urlSession(_ session: URLSession, dataTask: URLSessionDataTask, didReceive data: Data) {
let lines = String(data: data, encoding: .utf8)?.components(separatedBy: "\n") ?? []
for line in lines where line.hasPrefix("data: ") {
let jsonString = String(line.dropFirst(6))
guard jsonString != "[DONE]" else { return }
// parse delta.content from JSON
}
}
On Android — OkHttp with okhttp-sse:
val eventSource = EventSources.createFactory(client)
.newEventSource(request, object : EventSourceListener() {
override fun onEvent(source: EventSource, id: String?, type: String?, data: String) {
if (data == "[DONE]") return
// parse delta.content
}
})
The first token arrives in 200–400 ms. We update the UI no more often than every 50–100 ms to avoid overwhelming the thread. Streaming reduces the time to first response by 10x compared to full load — the user sees text almost instantly. This technique is 5x faster than non-streaming integration, significantly improving user experience.
How to Manage Conversation Context?
ChatGPT is stateless — you pass the history in the messages array. To avoid exceeding 128k tokens and going broke, we use one of these tactics:
- Sliding window — last 10–15 messages, discard the rest (40–60% token savings).
- Summarization — when exceeding a threshold of 8000 tokens, compress old history with a separate request (20–30% savings, but adds one API call).
- Selective memory — keep only facts that the user explicitly mentioned (10–15% savings, more complex).
| Tactic |
Token Savings |
Extra Requests |
Complexity |
| Sliding window |
40–60% |
0 |
Low |
| Summarization |
20–30% |
1 per compression |
Medium |
| Selective memory |
10–15% |
0 (requires NLP parsing) |
High |
How to Track API Costs and Optimize Expenses?
Each response includes usage.total_tokens. We log it to Firebase or your backend. At current pricing, GPT-4o-mini costs $0.15 per million input tokens and $0.60 per million output tokens. With 500 DAU sending 15 messages/day (average 300 input and 150 output tokens per message), monthly cost is approximately $1,800. Prompt caching reduces input cost by 35%, saving $630 monthly. Set a hard cap via the OpenAI Usage Limits dashboard. Constrain max_tokens per task — not 4096 when 256 suffices. Additionally, implement response caching: frequently asked queries are stored, reducing API calls by 30% on average. Total API costs can thus be cut by half.
Case Study: Language Learning App with AI Tutor
Our client developed a language learning app with an AI tutor. We used gpt-4o-mini, streaming, context as last 10 messages plus system prompt (300 tokens). Average request: 450 input + 180 output tokens. At 500 DAU and 15 messages per session — 3.4M tokens/day. Prompt Caching saved 35% input cost, reducing monthly expenses by $1,500. The integration cost was $8,500, fully recovered within 6 months.
Error Handling and Best Practices
429 — exponential backoff: 1s, 2s, 4s, up to 3 retries. 503 — same. 400 — usually invalid messages format. Log all errors to Crashlytics / Sentry without exposing the key.
Common integration mistakes:
- Storing the key in UserDefaults / SharedPreferences — violates key security protocols.
- No debounce on input — each character triggers a request, wasting tokens.
- Ignoring SSE parsing on Android — using
ResponseBody instead of EventSource.
- No summarization — context grows indefinitely, exceeding token limits.
- Not setting temperature and top_p parameters for output consistency.
Key implementation details: Use tokenization to estimate prompt lengths before sending. For semantic caching, apply cosine similarity on embedding vectors to reuse responses for similar queries. Ensure asynchronous processing with concurrency handling to avoid UI freezes. Paginate long context histories for efficient retrieval.
What's Included in the Work
- API documentation and backend-proxy architecture
- Source code integration in Swift/Kotlin/Flutter
- Load testing with a report
- Deployment and CI/CD instructions
- Team training (1-2 hour webinar)
- Support during App Store / Google Play publication
How We Do It
A five-step process:
- Analysis — review your architecture, choose context and caching strategy.
- Design — design backend-proxy, routing scheme, API specification.
- Implementation — write code in Swift/Kotlin/Flutter, set up streaming, logging, error handling.
- Testing — load test with real traffic simulation, security audit.
- Deployment — deploy backend-proxy, set up CI/CD, publish to App Store / Google Play.
Timelines and Cost
Basic API integration with streaming, context management, and backend-proxy takes 3–5 business days. Cost is calculated individually but typically starts at $5,000. Get a consultation: write to us, let's discuss your project.
Machine Learning in Mobile Apps: CoreML, TFLite, and On-Device Models
We distinguish two fundamentally different approaches: an app with on-device AI and an app that simply calls a cloud API. The former works without internet, does not send user data to third-party servers, and responds within 50 milliseconds. The latter depends on network latency and pricing plans. Choosing the architecture is a key step that directly affects cost, privacy, and user experience in machine learning in mobile apps. Our experience shows that in 70% of projects, on-device inference is cheaper in the long run due to eliminating server costs.
How to Choose Between CoreML and TFLite for On-Device Inference?
CoreML — Apple's native framework for running ML models on device. Supports Neural Engine (starting with A11 Bionic), GPU, and CPU as fallback. Models are converted to .mlmodel format via coremltools from PyTorch, ONNX, or TensorFlow. Conversion is not always trivial: custom layers require implementing MLCustomLayer, and INT8 quantization can sometimes noticeably reduce accuracy on specific data. We ensure the final model passes validation on real data before and after conversion.
TensorFlow Lite — cross-platform alternative for Android and Flutter. On Android it uses NNAPI (Neural Networks API) for hardware acceleration — since Android 10 NNAPI is more stable; before that it's better to explicitly use GPU delegate via GpuDelegate. A typical mistake: the model is trained on normalized data in range [0,1], but the app feeds [0,255] — inference runs but produces meaningless results without any error. We include an automatic input data validation module in the SDK.
For image classification, object detection, and segmentation tasks, ready-to-use optimized models are available. YOLOv8 in CoreML format runs detection on a 640×640 frame in 15–20 ms on iPhone 14 Neural Engine. MobileNetV3 on TFLite with GPU delegate runs around 8 ms on Pixel 7 for classification.
| Parameter |
CoreML |
TFLite |
| Platforms |
iOS, macOS, watchOS |
Android, iOS, Linux, embedded |
| Hardware acceleration |
Neural Engine, GPU, CPU |
NNAPI, GPU (OpenCL/OpenGL), CPU |
| Quantization support |
FP16, INT8 (with coremltools) |
FP16, INT8, dynamic range |
| Custom operations |
Via MLCustomLayer (Swift) |
Via delegates (Java/Kotlin) |
| Model bundle size |
~3–5 MB (MobileNetV2 quantized) |
~2–4 MB |
What If You Need Text Generation On-Device?
Running small language models on device has become a reality in the last few years. Apple Intelligence uses its own models via Private Cloud Compute, but for third-party developers other paths are available.
llama.cpp with Metal backend on iOS is a working approach for phi-3-mini (3.8B parameters, 4-bit quantization, ~2.3 GB). Inference: 15–25 tokens/second on iPhone 15 Pro. For integration in Swift, use the Swift Package llama.swift or a wrapper via C interface llama.h. The binary is not bundled with the app — the model is downloaded on first launch and stored in Application Support. Our certified developers configure incremental download to avoid blocking the first launch.
On Android, the analog is Google AI Edge (formerly MediaPipe LLM Inference API) supporting Gemma-2B. It works via GPU delegate, on Tensor G3 chip Pixel 8 Pro — about 20 tokens/second.
Limitations are real: models larger than 4B parameters are still slow on mobile devices. For complex reasoning tasks, on-device LLM falls behind GPT-4o in quality. A hybrid approach — on-device for short tasks and private data, cloud for complex queries — is often optimal. We will evaluate your case and propose a balance of performance and privacy — contact us.
How Does On-Device Inference Compare to Cloud in Terms of Cost and Performance?
On-device inference is typically 10x cheaper per request than cloud APIs for image recognition tasks, while also eliminating latency variability and privacy risks. The table below summarizes the trade-offs.
| Criteria |
On-Device Inference |
Cloud API |
| Latency |
<50ms |
200–500ms (including network) |
| Cost per 1M requests |
$0 (no server) |
$10–50 (AWS Rekognition, Google Vision) |
| Privacy |
Data stays on device |
Data sent to server |
| Offline |
Yes |
No |
| Scalability |
No server scaling issues |
Need to provision API capacity |
For an app with 100k MAU running 10 image recognitions per user per month, on-device inference can save up to $5,000 monthly compared to cloud API. Get a free consultation on your ML architecture today.
Integrating OpenAI API and Other Cloud Models
For scenarios where cloud inference is acceptable, integrating OpenAI, Anthropic, or Google Gemini is an HTTP client + streaming SSE. In Swift, AsyncThrowingStream is convenient for streaming responses. In Kotlin, use Flow.
Critically: API keys must never be stored in the app bundle. Even an obfuscated key can be extracted from the IPA in 10 minutes using strings or frida. Correct architecture: mobile app → your own backend → OpenAI API. The backend controls rate limiting, logs requests, and protects the key.
What Is Included in the Work (Deliverables)
- Trained and quantized model for the target device (documentation with metrics)
- SDK for integration (Swift/Kotlin/Flutter) with call examples
- Performance tests on 3–5 real devices
- Instructions for OTA model updates
- Support during App Store / Google Play moderation (compliance with Guidelines 4.2, 5.1)
- 2 weeks of technical support after release
Typical Project Pipeline
-
Task analysis — measure latency, privacy, size, supported devices.
-
Model prototyping — in Python, evaluate accuracy on target data.
-
Conversion and quantization — for CoreML/TFLite with validation.
-
Integration into the app — model wrapped in a service layer (easy to swap CoreML ↔ TFLite ↔ cloud).
-
Testing — on real devices, measure FPS, RAM, battery.
-
Deployment — via TestFlight / Firebase App Distribution, monitor metrics.
Timelines: integration of a ready CoreML/TFLite model — 1–2 weeks, development of a custom model with mobile optimization — from 6 weeks, on-device LLM chat with personalization — 4–8 weeks.
Why We Take on Complex Cases?
10+ years of experience in mobile development, 50+ implemented AI/ML solutions, guarantee of compatibility with current iOS and Android versions. All projects undergo code review and load testing. The cost includes preparation of moderation documentation and training of your team.
Contact us — we will help you choose the architecture and implement ML in your app turnkey. Order an audit of your existing solution — we will assess the potential for server cost savings free of charge. In some projects, savings can reach significant amounts per month.