Task: Extract Data from a Document and Pass to an LLM
A user opens a mobile app, attaches a PDF contract, and asks: "What is the termination period of the agreement?" At first glance, a typical scenario. But between file_picker and a meaningful model response, there are a dozen non-trivial decisions: from rendering each page of a scan to chunking a 100-page file that doesn't fit into the context. We implemented such functionality for several projects—financial, legal, medical—and know every pitfall. A mistake in one link and the model's response becomes meaningless, or you risk getting a rejection for non-compliance with App Store Review Guidelines regarding user content. To avoid this, let's break down the key stages: choosing the transfer method, handling scans, and organizing RAG for large volumes.
How to Pass a Document to an LLM?
Most LLMs accept text, not PDF. Conversion is needed. We consider two main approaches.
Direct Upload via Files API
OpenAI Assistants API and Gemini Files API accept PDF, DOCX, TXT directly. For a mobile app, this is the cleanest path: upload file, get file_id, insert into messages[]. However, there are limitations—OpenAI has a 512 MB per file limit and 100 files per assistant, and Files API is tied to Assistants/Batch, not Chat Completions.
Client-Side Text Extraction
For PDF on Android—PdfRenderer (built-in since API 21) for rendering pages to Bitmap + OCR via ML Kit TextRecognizer, or an Apache PDFBox port. On iOS—PDFKit + PDFPage.string for typewritten PDF; for scans—Vision framework with VNRecognizeTextRequest. Text goes into content[] as a string. PDFKit documentation
Problem with Scanned Documents
PDFKit.string returns an empty string for PDFs consisting of scanned pages—there is no text layer. ML Kit TextRecognizer handles it, but you need to render each page to Bitmap/CGImage and run OCR. For a 50-page document, this takes 2–5 seconds on device.
What to Do with Scans and Large Files?
Text Extraction: Pitfalls
On Android, PdfRenderer requires a ParcelFileDescriptor with the MODE_READ_ONLY flag. If the file arrives via a content:// URI from FileProvider, you need contentResolver.openFileDescriptor(). A direct File() from content:// throws FileNotFoundException—a common mistake for those unfamiliar with SAF (Storage Access Framework).
Multi-page documents must be processed page by page without loading everything into memory at once. PdfRenderer.Page must be closed after each page—page.close() is mandatory, otherwise an IllegalStateException occurs on the next iteration.
On iOS, PDFDocument(url:) can return nil for encrypted PDFs. Handle isEncrypted and request the password via UI rather than crashing silently.
Architectural Solution for Large Documents
The full text of a 100-page contract won't fit into the context window of most models—or it will fit, but at a high cost. The right path for large documents is RAG: split into chunks of 500–1000 tokens with an overlap of 50–100 tokens, index in a vector DB, retrieve the top 5 relevant chunks on query, and pass only those into context. Token savings of up to 40% compared to passing the full text directly. For documents up to 10 pages, client-side extraction works 3 times faster than uploading via Files API with response waiting.
For a mobile app, this usually means server-side processing: the client uploads the file to the backend, which handles chunking and embeddings. Only the query UI and response rendering remain on the client. Implementing vector search directly on the phone makes sense only for offline scenarios.
Comparison of Approaches
| Approach | Speed | Token Cost | Scan Support |
|---|---|---|---|
| Direct Files API | Fast (server) | High (full text) | Yes (if text layer exists) |
| Client extraction + text | Medium (depends on volume) | Medium (only text) | Yes (OCR on client) |
| RAG with server-side chunking | Slow (indexing), Fast (query) | Low (only relevant chunks) | Yes (if OCR exists) |
Formats and Limits
| Format | Android | iOS | API Limit (OpenAI) |
|---|---|---|---|
| PDF (text) | PdfRenderer + PDFBox | PDFKit | 512 MB |
| PDF (scanned) | ML Kit OCR | Vision VNRecognizeTextRequest | — (preprocessing needed) |
| DOCX | Apache POI (Java) | — | 512 MB (via Files API) |
| TXT / MD | Native | Native | No limits |
| XLSX | Apache POI | — | 512 MB |
DOCX on iOS without third-party libraries is painful. Either server-side conversion (LibreOffice headless) or limit format support to PDF + TXT for the mobile client.
What Is Included in Turnkey Work
- Audit of document formats in your product
- Selection of optimal strategy (Files API vs client extraction vs RAG)
- Implementation of file upload (
file_picker, SAF,UIDocumentPickerViewController) - Text conversion and cleaning (OCR for scans)
- Integration with LLM (OpenAI / Gemini / Anthropic)
- Progress indicators for long operations
- Testing on real documents of varying quality
- Documentation and team training
Why Choose Us
We are a team of mobile developers with 5+ years of experience in creating AI solutions for iOS and Android. We have implemented over 20 integrations of multimodal input for financial, legal, and medical projects. We guarantee code quality and adherence to deadlines. Contact us to assess your project—we will select the optimal solution for your budget.
Timelines: basic support for PDF + TXT with direct transfer — 1–2 weeks. Full pipeline with OCR, multiple formats, and RAG for large documents — 4–6 weeks. We will evaluate your project for free—reach out.







