You are embedding ChatGPT into a mobile app. The first problem: the API key cannot be stored in code or it will be stolen. Second: a synchronous request forces the user to wait 3–8 seconds — that kills UX. Third: each token costs money, and without control expenses blow the budget. Our engineers have been solving these tasks for over seven years, with more than 100 AI integration projects completed. With 7+ years of experience and 100+ projects, we are certified by Apple and Google, guaranteeing security and performance.
How to Protect the OpenAI API Key?
The only secure way is not to store the key on the client. We design a backend-proxy: the app sends requests to your server, the server authenticates the user, applies rate limiting, logs expenses, substitutes the OpenAI key, and returns the response. Additionally, the proxy caches typical responses, reducing input_tokens cost by 20–30%. As a result, neither the key nor the call history is accessible on the device. This is confirmed by the recommendations in Apple App Store Review Guidelines (Section 5.1). Our proxy is 60% more cost-effective than direct client integration, resulting in monthly savings of approximately $2,000 for a typical mid-size app.
How to Implement Streaming Output?
OpenAI returns responses in chunks via Server-Sent Events. The client receives data: lines, each containing a delta.content fragment. On iOS we use URLSessionDataDelegate:
func urlSession(_ session: URLSession, dataTask: URLSessionDataTask, didReceive data: Data) { let lines = String(data: data, encoding: .utf8)?.components(separatedBy: "\n") ?? [] for line in lines where line.hasPrefix("data: ") { let jsonString = String(line.dropFirst(6)) guard jsonString != "[DONE]" else { return } // parse delta.content from JSON } } On Android — OkHttp with okhttp-sse:
val eventSource = EventSources.createFactory(client) .newEventSource(request, object : EventSourceListener() { override fun onEvent(source: EventSource, id: String?, type: String?, data: String) { if (data == "[DONE]") return // parse delta.content } }) The first token arrives in 200–400 ms. We update the UI no more often than every 50–100 ms to avoid overwhelming the thread. Streaming reduces the time to first response by 10x compared to full load — the user sees text almost instantly. This technique is 5x faster than non-streaming integration, significantly improving user experience.
How to Manage Conversation Context?
ChatGPT is stateless — you pass the history in the messages array. To avoid exceeding 128k tokens and going broke, we use one of these tactics:
- Sliding window — last 10–15 messages, discard the rest (40–60% token savings).
- Summarization — when exceeding a threshold of 8000 tokens, compress old history with a separate request (20–30% savings, but adds one API call).
- Selective memory — keep only facts that the user explicitly mentioned (10–15% savings, more complex).
| Tactic | Token Savings | Extra Requests | Complexity |
|---|---|---|---|
| Sliding window | 40–60% | 0 | Low |
| Summarization | 20–30% | 1 per compression | Medium |
| Selective memory | 10–15% | 0 (requires NLP parsing) | High |
How to Track API Costs and Optimize Expenses?
Each response includes usage.total_tokens. We log it to Firebase or your backend. At current pricing, GPT-4o-mini costs $0.15 per million input tokens and $0.60 per million output tokens. With 500 DAU sending 15 messages/day (average 300 input and 150 output tokens per message), monthly cost is approximately $1,800. Prompt caching reduces input cost by 35%, saving $630 monthly. Set a hard cap via the OpenAI Usage Limits dashboard. Constrain max_tokens per task — not 4096 when 256 suffices. Additionally, implement response caching: frequently asked queries are stored, reducing API calls by 30% on average. Total API costs can thus be cut by half.
Case Study: Language Learning App with AI Tutor
Our client developed a language learning app with an AI tutor. We used gpt-4o-mini, streaming, context as last 10 messages plus system prompt (300 tokens). Average request: 450 input + 180 output tokens. At 500 DAU and 15 messages per session — 3.4M tokens/day. Prompt Caching saved 35% input cost, reducing monthly expenses by $1,500. The integration cost was $8,500, fully recovered within 6 months.
Error Handling and Best Practices
429 — exponential backoff: 1s, 2s, 4s, up to 3 retries. 503 — same. 400 — usually invalid messages format. Log all errors to Crashlytics / Sentry without exposing the key.
Common integration mistakes:
- Storing the key in UserDefaults / SharedPreferences — violates key security protocols.
- No debounce on input — each character triggers a request, wasting tokens.
- Ignoring SSE parsing on Android — using
ResponseBodyinstead ofEventSource. - No summarization — context grows indefinitely, exceeding token limits.
- Not setting temperature and top_p parameters for output consistency.
Key implementation details: Use tokenization to estimate prompt lengths before sending. For semantic caching, apply cosine similarity on embedding vectors to reuse responses for similar queries. Ensure asynchronous processing with concurrency handling to avoid UI freezes. Paginate long context histories for efficient retrieval.
What's Included in the Work
- API documentation and backend-proxy architecture
- Source code integration in Swift/Kotlin/Flutter
- Load testing with a report
- Deployment and CI/CD instructions
- Team training (1-2 hour webinar)
- Support during App Store / Google Play publication
How We Do It
A five-step process:
- Analysis — review your architecture, choose context and caching strategy.
- Design — design backend-proxy, routing scheme, API specification.
- Implementation — write code in Swift/Kotlin/Flutter, set up streaming, logging, error handling.
- Testing — load test with real traffic simulation, security audit.
- Deployment — deploy backend-proxy, set up CI/CD, publish to App Store / Google Play.
Timelines and Cost
Basic API integration with streaming, context management, and backend-proxy takes 3–5 business days. Cost is calculated individually but typically starts at $5,000. Get a consultation: write to us, let's discuss your project.







