A/B Testing of Chatbot Scenarios in a Mobile App
A product manager wants to test which bot welcome message converts better to purchase: "Hi, how can I help?" or "I'll show products for your query right away." An A/B test at the UI level is straightforward. But a bot is not just text: it's a dialog graph, a set of intents, escalation logic to a human agent. We've faced this dozens of times — setting up A/B testing for bot scenarios requires separate infrastructure. Our experience shows that a proper end-to-end A/B testing implementation takes 3–5 days for two variants, and up to 2 weeks for complex server-side logic.
What We Test in a Bot
Bot scenarios differ from UI elements: a variant is not a button color but an entire dialog graph. A user might complete 7 steps in variant A and 3 steps in variant B to reach the same outcome. The metric is not a click but the completion of a target action (purchase, request, resolved issue). This complicates measurement and requires event tracking at each dialog step.
Typical hypotheses for A/B testing on a bot:
- Different greetings and tone of voice
- Quick replies vs. text input on the first step
- Timing of escalation offer to an operator (immediately vs. after 2 failed intents)
- Different CTA phrasings within the dialog
How We Implement A/B Testing for Bots
We offer a proven approach that includes platform selection, configuration setup, and event tracking integration. Each step is documented and handled by our certified engineers.
Platform Selection. The table below compares popular solutions:
| Platform |
Type |
Statistics |
Self-hosted |
SDK |
| Firebase Remote Config |
Client-side |
Automatic |
No |
iOS, Android, Web |
| Growthbook |
Client-side/Server-side |
Advanced |
Yes |
iOS, Android, Web |
| Statsig |
Client-side |
Powerful, with caching |
No |
iOS, Android, Web |
| Custom Server-side |
Server-side |
Full control |
Yes |
Any via API |
Integration with Firebase. For a quick start, we use Firebase Remote Config. Bot configuration parameters (scenario ID, prompt version, escalation threshold) are read at app launch:
let remoteConfig = RemoteConfig.remoteConfig()
remoteConfig.fetch(withExpirationDuration: 3600) { [weak self] status, error in
guard status == .success else { return }
remoteConfig.activate { _, _ in
let botVariant = remoteConfig["bot_scenario_variant"].stringValue ?? "control"
self?.chatViewModel.loadScenario(variant: botVariant)
}
}
Firebase automatically splits the audience into groups by traffic percentage. Additional conditions (country, app version) can be set. Analytics is handled via Firebase Analytics with conversion events.
Server-side A/B vs. Client-side. If the bot is implemented through a server-side dialog engine (Rasa, Dialogflow CX, custom), we recommend managing the variant on the server. The client sends userId + sessionId, the server selects the experiment group and returns responses for the appropriate variant. This prevents cheating and simplifies analytics. We used this approach in a project with over 500,000 users.
Why Statistical Significance Is Critical
The main mistake in A/B tests is stopping the test at the first promising numbers. A minimum sample size must be calculated in advance. For a desired effect of 5%, baseline conversion of 15%, and test power of 80%, at least 2,800 users per group are required. Firebase A/B Testing calculates this automatically, but we additionally verify the calculations.
Our engineers, with over 5 years of experience in mobile development, guarantee the test will only be stopped after reaching statistical significance. Otherwise, we re-analyze at no charge.
Dialog Event Tracking
Without detailed tracking of each step, it's impossible to understand where the user dropped out of the funnel. Minimum set of events:
-
bot_session_start — {variant, userId, sessionId}
-
bot_message_sent — {variant, stepId, messageType}
-
bot_message_received — {variant, stepId, intentId, confidence}
-
bot_intent_failed — {variant, stepId, userInput} — when NLU didn't recognize intent
-
bot_escalated — {variant, stepId, reason}
-
bot_goal_completed — {variant, goalType} — conversion event
All events with variant and sessionId allow restoring the full user path in any variant. We set up this tracking as part of the service — you get ready analytics in your chosen platform.
What's Included in the Work
- Audit of the current bot scenario and hypothesis formulation
- Selection of the A/B testing platform (Firebase, Growthbook, Statsig, or server-side)
- Configuration setup (Remote Config, feature flags)
- Development of event tracking for each dialog step
- Integration and test launch
- Monitoring and statistical significance calculation
- Automatic winner selection (configurable)
- Results documentation and scaling recommendations
Timeline Estimates
Implementation of A/B testing for two scenario variants using Firebase Remote Config and event tracking: 3 to 5 days. If integration with a server-side dialog engine and more complex audience segmentation is required: 1 to 2 weeks.
Want to know how long your project will take? Contact us for a free estimate and optimal solution.
Mobile app testing automation: from unit to E2E
A flaky test that fails on CI once every five runs without a reproducible cause is worse than no test. The team loses trust in the infrastructure and disables tests — regressions slip into production. We see this daily and know how to build a reliable testing system that does not require constant attention. Contact us for a free consultation and test architecture assessment.
Why are flaky tests dangerous?
One unstable check can break the pipeline, blocking a release. Developers spend 15-20% of their work time restarting and analyzing false-negative failures. Automation without stability is not saving efficiency but losing it. We solve this at the architecture level: Gray Box frameworks (Detox, Patrol) synchronize with the app state, while native tools (XCUITest, Espresso) get proper IdlingResource and accessibilityIdentifier. Result: stability >99% on CI.
What should you unit test in mobile apps?
On iOS XCTest is the foundation. Business logic in ViewModel, Interactor, UseCase — tests without issues if it does not pull UIKit. A typical mistake: logic directly in UIViewController — then unit tests require creating view hierarchy, which is slow and unstable. The solution is to move logic to services with @testable import.
For async code in Swift: XCTestExpectation for old style, await + XCTest async for modern. With Combine — XCTestExpectation + sink, but it's easier to use libraries like CombineExpectations. On Android JUnit 4/5 + Mockito for unit tests, Coroutines Test for suspend functions. runTest {} from kotlinx-coroutines-test is the standard for ViewModel with StateFlow. Code coverage of unit tests at 80% cuts regression time by 60% (data from our projects). Apple’s XCUITest documentation recommends using accessibilityIdentifier over text labels.
UI Tests: Stability Over Coverage
XCUITest (iOS) and Espresso (Android) — native UI tests. They run fast, are integrated with IDE, but test one platform. The main issue with XCUITest is fragile selectors. app.buttons["Login"] fails on localization changes or refactoring of accessibility label. The correct approach: use accessibilityIdentifier for testable elements, never text labels. Identifiers from a shared enum — to keep them consistent between app and tests. Experience shows: this practice reduces flakiness by 90%.
Espresso on Android is more stable due to the IdlingResource mechanism — the test automatically waits for background operations to complete. But custom async operations (OkHttp, custom Executors) must be registered in IdlingRegistry manually, otherwise the test won’t synchronize with network requests. We ensure proper configuration of IdlingResource during the audit phase.
Detox and Patrol: End-to-End for React Native and Flutter
Detox — E2E framework for React Native, developed by Wix. Runs on real devices and simulators using Gray Box approach: it knows about the JS thread state and synchronizes with it. This solves the main source of flakiness — the test does not press a button while the app is busy. Detox setup is non-trivial. Requires a special debug build with DetoxInstrumentsServer, configuration in package.json, and no separate Appium server. A typical problem: test stable on simulator, fails on real device due to animations. Solution: animations: disabled in Detox config for E2E build.
Patrol — analog for Flutter. Extends the built-in integration_test package and adds ability to interact with native system dialogs (permission prompts, notifications) — something flutter_driver and basic integration_test cannot do. For CI, use via patrol test --target integration_test/app_test.dart. Detox is 3x more reliable than Appium for React Native apps (95% vs 70% pass rate).
Appium: Cross-Platform at a Cost
Appium — when you need to cover iOS and Android with the same tests. Uses WebDriver protocol on top of XCUITest and UiAutomator2 drivers. Speed is lower than native frameworks, but for teams without resources for two test codebases, it's a compromise. Appium 2.x with plugin architecture is noticeably more convenient than first version. appium-doctor diagnoses the environment — useful when setting up CI.
CI and Parallelization
For parallel XCUITest runs we use Xcode Cloud or xcodebuild test-without-building with multiple simulators via parallel-testing-enabled. Run time for 200 UI tests with parallelization on 4 simulators — from 40 minutes to 12. On Android we use Firebase Test Lab with sharding.
| Framework |
Platform |
Gray Box |
Speed |
System Dialogs |
| XCUITest |
iOS |
No |
High |
Yes (via addUIInterruptionMonitor) |
| Espresso |
Android |
Yes (IdlingResource) |
High |
Limited |
| Detox |
React Native |
Yes |
Medium |
Limited |
| Patrol |
Flutter |
Partial |
Medium |
Yes |
| Appium |
iOS + Android |
No |
Low |
Yes |
Typical Setup Mistakes (and How to Avoid Them)
| Mistake |
Consequence |
Solution |
| Using text labels in selectors |
Tests fail on localization |
accessibilityIdentifier from enum |
| Missing IdlingResource for custom Executor |
Espresso does not wait for server response |
Register in IdlingRegistry |
| Enabled animations on real device with Detox |
Flaky tests due to timing |
animations: disabled in E2E build |
| Parallelization without state isolation |
Data races between tests |
Run each test in a fresh simulator |
How We Do It: Process
-
Audit current code and CI — evaluate flakiness, coverage, bottlenecks. We typically find 15-20% of tests are flaky.
-
Design test architecture — choose framework, selectors, mocks.
-
Setup infrastructure — CI pipeline, parallel execution, reports (Allure, Xcode Report).
-
Write tests — unit, UI, E2E, performance (XCTMetrics, Macrobenchmark).
-
Integration and stabilization — run 200+ tests, catch flaky cases. Past projects show flakiness drops from 15% to 2%.
-
Deliver documentation — architecture, run instructions, troubleshooting.
Deliverables
- Architectural documentation of test coverage
- Configured CI pipeline with parallelization and reports
- Test code (unit, UI, E2E) with styleguide
- Team training (2-hour workshop)
- Access to test builds and CI logs
- One-month post-delivery support (fix flakiness, update for new versions)
Estimated Timelines
Setting up infrastructure from scratch (CI, unit + UI tests, reports) — 2-3 weeks. Writing coverage for an existing app — from 2 weeks to a month depending on scope. We will assess your project in 2 days — contact us. Get a customized automation plan for your project – reach out today. 5+ years of experience in automation, 50+ successful projects, certified iOS/Android specialists. We guarantee test stability >98% on CI after implementation.