Rasa NLP Integration for Mobile Chatbots
Standard Rasa config on Russian yields about 50 % accuracy — on short phrases, this means every second request is misrecognized. Clients with strict confidentiality requirements (medicine, finance) cannot use cloud NLP services like Dialogflow or Google NL. For them, Rasa is the only option, but it needs proper tuning. We have been doing this for many years: over 20 projects, average accuracy improvement from 50% to 85% through custom pipeline and high-quality dataset.
Why Rasa is Better than Dialogflow for Sensitive Data
Rasa is self-hosted, Dialogflow is cloud-based. For companies where data must stay on-premise, the difference is fundamental. Rasa gives full control over the pipeline: you can replace any component, add your own tokenizer, use custom models. Dialogflow limits you to built-in mechanisms. At scale of 10 000+ dialogues per day, Rasa is 3–4 times cheaper due to no per-request fees.
| Characteristic | Rasa | Dialogflow |
|---|---|---|
| Deployment | Self-hosted (your server) | Cloud (Google) |
| Data control | Full | Limited |
| Cost at 10 000 requests/day | ~$200/month (VPS) | ~$400/month (Standard) |
| Accuracy on Russian (custom pipeline) | 85 %+ | 80 % |
| Custom actions support | Yes (Python) | Yes (Webhook, slower) |
How We Configure Rasa for Russian
The main issue is that the default WhitespaceTokenizer works poorly with Russian morphology. We replace it with SpacyNLP using the ru_core_news_md model and add char n-gram features via CountVectorsFeaturizer. Here is the production config we use:
language: ru pipeline: - name: SpacyNLP model: ru_core_news_md - name: SpacyTokenizer - name: SpacyFeaturizer - name: RegexFeaturizer - name: LexicalSyntacticFeaturizer - name: CountVectorsFeaturizer - name: CountVectorsFeaturizer analyzer: char_wb min_ngram: 1 max_ngram: 4 - name: DIETClassifier epochs: 150 - name: EntitySynonymMapper - name: ResponseSelector epochs: 100 - name: FallbackClassifier threshold: 0.7 ambiguity_threshold: 0.1 This pipeline yields 85 % accuracy on a set of 20–50 intents. DIETClassifier with 150 epochs trains in 5–10 minutes on a VPS with 4 GB RAM.
How Rasa Core Manages Dialogue
Rasa separates NLU (intent recognition) and Core (dialogue management). Core operates on rules (rules.yml) and stories (stories.yml). The most common mistake is trying to describe all scenarios via rules and ending up with a fragile system that breaks on non-standard utterance order.
Rule: rigid commands (cancel, help, restart) go in rules. Multi-step scenarios with variability go in stories. Rasa Core learns from stories and generalizes unseen scenarios — that is its main advantage over decision trees.
Custom Actions. Dynamic responses (order status, available slots) are implemented via action_server — a separate Python service that Rasa calls over HTTP. The mobile app does not interact directly with it:
class ActionCheckOrderStatus(Action): def name(self) -> str: return "action_check_order_status" async def run(self, dispatcher, tracker, domain) -> list: order_id = tracker.get_slot("order_id") status = await order_service.get_status(order_id) dispatcher.utter_message(text=f"Your order #{order_id}: {status}") return [SlotSet("order_status", status)] How to Integrate Rasa with a Mobile App
Rasa Server provides a REST channel at /webhooks/rest/webhook. The mobile app sends a POST request:
{ "sender": "user_device_id_or_session_uuid", "message": "message text" } The response is an array of messages, each can be text, image, buttons, or custom payload.
On Android, this is a standard Retrofit call. Important: sender must be a stable session identifier — Rasa stores slots between requests within the same sender. Generating a new ID each time will lose dialogue context.
For production, Rasa should not be exposed directly to the internet — we place nginx in front with rate limiting and token authentication.
Deployment
Rasa Server + Action Server are easily run via Docker Compose. The model is trained with rasa train and mounted into the container. On a small VPS, training 50 intents takes 5–10 minutes — perfectly acceptable for CI/CD.
Rasa Enterprise (commercial version) adds analytics and A/B testing of dialogues, but for most tasks the open-source version is sufficient.
What's Included in Our Work
- Domain audit and collection of example utterances for the NLU dataset.
- Pipeline configuration, base model training, accuracy evaluation via
rasa test. - Development of dialogue scenarios (rules + stories), custom actions.
- REST channel integration with the mobile client, infrastructure setup.
- Documentation and training for your team.
- 30-day warranty support.
Timeframes
Integration with an existing Rasa server — 3–4 days. Full cycle including model training from scratch, scenario writing, and deployment — 2–4 weeks depending on the number of intents.
Contact us for a project assessment. Get a consultation on configuring Rasa for your task. Rasa Documentation | Wikipedia: Rasa







