A/B Testing of Chatbot Scenarios in a Mobile App
A product manager wants to test which bot welcome message converts better to purchase: "Hi, how can I help?" or "I'll show products for your query right away." An A/B test at the UI level is straightforward. But a bot is not just text: it's a dialog graph, a set of intents, escalation logic to a human agent. We've faced this dozens of times — setting up A/B testing for bot scenarios requires separate infrastructure. Our experience shows that a proper end-to-end A/B testing implementation takes 3–5 days for two variants, and up to 2 weeks for complex server-side logic.
What We Test in a Bot
Bot scenarios differ from UI elements: a variant is not a button color but an entire dialog graph. A user might complete 7 steps in variant A and 3 steps in variant B to reach the same outcome. The metric is not a click but the completion of a target action (purchase, request, resolved issue). This complicates measurement and requires event tracking at each dialog step.
Typical hypotheses for A/B testing on a bot:
- Different greetings and tone of voice
- Quick replies vs. text input on the first step
- Timing of escalation offer to an operator (immediately vs. after 2 failed intents)
- Different CTA phrasings within the dialog
How We Implement A/B Testing for Bots
We offer a proven approach that includes platform selection, configuration setup, and event tracking integration. Each step is documented and handled by our certified engineers.
Platform Selection. The table below compares popular solutions:
| Platform | Type | Statistics | Self-hosted | SDK |
|---|---|---|---|---|
| Firebase Remote Config | Client-side | Automatic | No | iOS, Android, Web |
| Growthbook | Client-side/Server-side | Advanced | Yes | iOS, Android, Web |
| Statsig | Client-side | Powerful, with caching | No | iOS, Android, Web |
| Custom Server-side | Server-side | Full control | Yes | Any via API |
Integration with Firebase. For a quick start, we use Firebase Remote Config. Bot configuration parameters (scenario ID, prompt version, escalation threshold) are read at app launch:
let remoteConfig = RemoteConfig.remoteConfig() remoteConfig.fetch(withExpirationDuration: 3600) { [weak self] status, error in guard status == .success else { return } remoteConfig.activate { _, _ in let botVariant = remoteConfig["bot_scenario_variant"].stringValue ?? "control" self?.chatViewModel.loadScenario(variant: botVariant) } } Firebase automatically splits the audience into groups by traffic percentage. Additional conditions (country, app version) can be set. Analytics is handled via Firebase Analytics with conversion events.
Server-side A/B vs. Client-side. If the bot is implemented through a server-side dialog engine (Rasa, Dialogflow CX, custom), we recommend managing the variant on the server. The client sends userId + sessionId, the server selects the experiment group and returns responses for the appropriate variant. This prevents cheating and simplifies analytics. We used this approach in a project with over 500,000 users.
Why Statistical Significance Is Critical
The main mistake in A/B tests is stopping the test at the first promising numbers. A minimum sample size must be calculated in advance. For a desired effect of 5%, baseline conversion of 15%, and test power of 80%, at least 2,800 users per group are required. Firebase A/B Testing calculates this automatically, but we additionally verify the calculations.
Our engineers, with over 5 years of experience in mobile development, guarantee the test will only be stopped after reaching statistical significance. Otherwise, we re-analyze at no charge.
Dialog Event Tracking
Without detailed tracking of each step, it's impossible to understand where the user dropped out of the funnel. Minimum set of events:
-
bot_session_start— {variant, userId, sessionId} -
bot_message_sent— {variant, stepId, messageType} -
bot_message_received— {variant, stepId, intentId, confidence} -
bot_intent_failed— {variant, stepId, userInput} — when NLU didn't recognize intent -
bot_escalated— {variant, stepId, reason} -
bot_goal_completed— {variant, goalType} — conversion event
All events with variant and sessionId allow restoring the full user path in any variant. We set up this tracking as part of the service — you get ready analytics in your chosen platform.
What's Included in the Work
- Audit of the current bot scenario and hypothesis formulation
- Selection of the A/B testing platform (Firebase, Growthbook, Statsig, or server-side)
- Configuration setup (Remote Config, feature flags)
- Development of event tracking for each dialog step
- Integration and test launch
- Monitoring and statistical significance calculation
- Automatic winner selection (configurable)
- Results documentation and scaling recommendations
Timeline Estimates
Implementation of A/B testing for two scenario variants using Firebase Remote Config and event tracking: 3 to 5 days. If integration with a server-side dialog engine and more complex audience segmentation is required: 1 to 2 weeks.
Want to know how long your project will take? Contact us for a free estimate and optimal solution.







