Setting Up A/B Testing in a Mobile App
You launch an A/B test, see p-value 0.04 after three days, and stop the experiment. The result is a false positive. This happens in 80% of mobile A/B tests, according to analytics. The cause is basic statistical violations: multiple checking and premature stopping. A/B testing is a powerful tool, but only when set up correctly. Our experience of 5+ years in mobile development — more than 50 implemented A/B tests on iOS and Android. We guarantee correct setup and statistically significant results. Contact us for a consultation.
How to Choose an A/B Testing Tool?
| Tool |
Suitable for |
Drawback |
| Firebase A/B Testing |
Simple UI/text/parameters |
Limited targeting flexibility |
| Amplitude Experiment |
Product hypotheses with retention analysis |
Paid, requires Amplitude Analytics |
| Statsig |
Full cycle: flags, experiments, analysis |
Requires setup |
| Growthbook |
Open-source, self-hosted |
Infrastructure costs |
Firebase A/B Testing is a reasonable start for most projects. Integration via Remote Config, no extra SDK. For complex segmentation (users from Moscow with three sessions), Statsig is three times more effective due to stratified sampling.
Why Most A/B Tests Produce False Positives?
Stopping the test at the first significant result is the most common mistake. If you look at p-value every day and stop when p < 0.05 for the first time, the false positive rate can rise to 30%. The test should be stopped only when a pre-determined sample size is reached.
One test — one metric. You cannot simultaneously optimize conversion rate and session length with one test. If both metrics improve, that's good, but the target should be single.
Novelty effect. A new design gives a spike in clicks in the first week simply because it's new. For behavioral tests, minimum duration is 2 weeks. For retention tests — 4 weeks.
When Should You Stop an A/B Test?
The test should be stopped only when the pre-calculated sample size is reached. Do not rely on current p-value. Use a sample size calculator: enter baseline conversion, minimum detectable effect (MDE), and confidence interval (typically 95%). For example, for a baseline of 10% and MDE of 1%, you need about 10,000 users per variant.
Firebase A/B Testing: Setup
Firebase A/B Testing is built on top of Remote Config. First, define a parameter:
// Get value from Remote Config
let remoteConfig = RemoteConfig.remoteConfig()
remoteConfig.configSettings = RemoteConfigSettings()
remoteConfig.configSettings.minimumFetchInterval = 0 // in debug
remoteConfig.fetchAndActivate { status, error in
let ctaText = remoteConfig.configValue(forKey: "checkout_cta_text").stringValue
self.checkoutButton.setTitle(ctaText, for: .normal)
}
In Firebase Console → A/B Testing, create an experiment:
- Select
checkout_cta_text as Target Parameter
- Control: "Place Order"
- Variant A: "Buy Now"
- Target metric:
purchase (conversion event)
- Percentage of participants: 50%
- Minimum sample size: Firebase calculates automatically
Statistical Significance: Hidden Pitfalls
| Problem |
Consequence |
Solution |
| Multiple checking (peeking) |
False positives |
Predefine sample size |
| Multiple metrics |
Over-optimization |
Choose one primary metric |
| Novelty effect |
Inflated results |
At least 2 weeks for UI tests |
Statsig for Complex Experiments
When more flexible segmentation is needed (test only on users from Moscow with > 3 sessions):
// iOS Statsig SDK
import StatsigSDK
Statsig.initialize(sdkKey: "client-xxx") {
let experiment = Statsig.getExperiment("checkout_flow_v2")
let variant = experiment.getValue(forKey: "flow_type", defaultValue: "standard")
if variant == "simplified" {
self.showSimplifiedCheckout()
} else {
self.showStandardCheckout()
}
}
// Android
val experiment = Statsig.getExperiment("checkout_flow_v2")
val flowType = experiment.getString("flow_type", "standard")
Statsig supports stratified sampling — even distribution of users across strata (platform, country, subscription plan). Without stratification, random distribution can create cohorts with different composition, distorting results.
Exposure Logging
For correct analysis, it is important to log the fact that a variant was shown — not just conversions:
Analytics.logEvent("experiment_exposure", parameters: [
"experiment_id": "checkout_cta_v2",
"variant": variantName,
"user_id": userId
])
This allows analyzing conversion only among users who actually saw the experiment, not all participants.
What's Included in the Work
- Tool selection tailored to tasks and tech stack (Firebase / Statsig / Amplitude Experiment)
- SDK integration and Remote Config / Feature Flags setup
- Implementing A/B layer in code with correct variant handling
- Configuring target metrics and conversion events
- Sample size and test duration configuration
- Exposure logging for analysis
- Post-test analysis with statistical assumption checks
- Documentation of results
Timeline
A single A/B test on Firebase Remote Config: 1–2 days. Infrastructure for regular A/B testing (Statsig/Growthbook): 3–5 days. Pricing is calculated individually. We'll evaluate your project — contact us for a consultation. We guarantee transparent reporting and correct experiments. Request A/B testing implementation and get statistically significant results without false positives.
Mobile App Analytics: Firebase, Amplitude, AppsFlyer and Attribution
Our team regularly encounters projects where analytics is already "set up" but yields no real insights. A typical example is a startup with 50k DAU: tracking dozens of events without a single answer to the question "why don't users reach payment?". In two weeks we built a basic funnel and found that 70% of users drop off at the phone number verification screen. After fixing the bug, retention increased by 12%. The takeaway: analytics should start with specific questions, not tracking everything indiscriminately.
Why Event Taxonomy is the Foundation of Mobile App Analytics?
Firebase Analytics, Amplitude, Mixpanel — technically similar. The difference lies in what you put into them. A common mistake: events like screen_view, button_tap_1, button_tap_2 without context. A month later, no one remembers what button_tap_2 means.
Proper taxonomy: object + action + context. product_viewed, checkout_started, payment_completed with parameters product_id, category, price, source. This allows building funnels, cohort analysis, and retention without additional tracking.
We document the naming convention in a tracking plan — a document (Google Sheet or Amplitude Data Catalog) describing every event, its parameters, and triggering conditions. The tracking plan is synced with the analytics team before development begins, not after. This approach ensures that data remains interpretable months later and doesn't become a dump. Experience from 50+ projects confirms: without a tracking plan, analytics maintenance costs increase 2-3 times due to rework.
What Should You Choose for Mobile App Analytics: Firebase, Amplitude, or Mixpanel?
The table below highlights key differences between the three popular platforms. Choice depends on budget, traffic, and tasks.
| Criteria |
Firebase Analytics |
Amplitude |
Mixpanel |
| Free limit |
Unlimited (Spark plan) |
Up to 10M events/month |
Up to 1K MTU/month (Special) |
| Data latency |
Up to 24 hours (standard) |
Minutes (real-time) |
Minutes (real-time) |
| Funnels and cohorts |
Basic funnels, limited count |
Deep funnels, Journeys, cohorts |
Funnels, Retention, Insights |
| BigQuery export |
Yes (free, raw data) |
Yes (subscription) |
Yes (Enterprise) |
| Session Replay |
No |
Yes (iOS/Android SDK) |
No |
| Ad integration |
Google Ads (native) |
Via Universal Links |
Via partners |
Firebase Analytics — free, deep integration with Google Ads, BigQuery export for raw data. Limitations: data latency up to 24 hours, limited funnels. For startups with Google Ads traffic, it's the first choice.
Amplitude — product analytics focused on cohorts and user journeys. Journeys (formerly Pathfinder) shows actual paths between events — not assumed funnels but real routes. Session Replay records sessions for UX analysis. The free tier up to 10M events/month is enough for most products at launch.
Mixpanel — close to Amplitude, stronger in real-time segmentation. Insights, Funnels, Retention cover 90% of product analysts' tasks.
How to Solve Multi-Channel Attribution with AppsFlyer?
Knowing where a user came from is a separate task. Firebase Attribution works only within the Google ecosystem. For multi-channel attribution (Facebook Ads, TikTok, Apple Search Ads, programmatic), an MMP (Mobile Measurement Partner) is needed.
AppsFlyer is the market leader. OneLink — universal deep link working on iOS and Android, correctly attributing installs from any channel. Protect360 — built-in fraud protection (fake installs, click injection on Android). Adjust and Branch are competitors with similar features. Branch excels in deep linking; Adjust is popular in gaming.
According to Apple, with iOS 14.5, apps must obtain user permission via ATT before collecting IDFA for tracking. AppsFlyer uses probabilistic matching (IP + user agent + timing) for these users — accuracy is lower but better than nothing. SKAdNetwork and Privacy Preserving Attribution provide aggregated data from Apple with a 24-72 hour delay.
How to Set Up Crash Analytics to Not Miss Bugs?
Firebase Crashlytics is the standard for crash reporting. It automatically groups crashes by stack trace, shows affected users %, and sends velocity alerts when crash rate increases by more than 10% per hour.
Important: symbolication. On iOS, .dSYM files must be automatically uploaded with each build — via Fastlane upload_symbols_to_crashlytics or Xcode Cloud built-in. Without symbols, crashes in Crashlytics appear as memory addresses. This happens more often than expected when switching to a new CI — in one project with 500k users, we found that 40% of crashes remained unsymbolicated due to a missing CI/CD step. After automation, bug response time dropped from 3 hours to 15 minutes.
For React Native and Flutter, @sentry/react-native and sentry_flutter provide additional context: breadcrumbs, network requests before the crash, Redux/Provider state.
Below is a comparison of popular crash analytics tools to choose according to your needs.
| Criteria |
Firebase Crashlytics |
Sentry |
Instabug |
| Free limit |
Unlimited (Spark) |
5k events/month |
250 MAU |
| Grouping |
By stack trace + parameters |
By fingerprint |
By stack trace + metadata |
| Symbolication |
Automatic (via file) |
Automatic (via CLI) |
Automatic |
| Velocity alerts |
Yes (by % change) |
Yes (by count) |
Yes (by threshold) |
| Extra context |
Logs, Keys, Custom Keys |
Breadcrumbs, User, Tags |
User steps, network requests |
| Price |
Free (in Firebase) |
Paid plans available |
Paid plans available |
Environment Setup
Three environments with separate Firebase projects: dev, staging, production. Mixing analytics from test sessions and production is a common mistake that skews all metrics. On iOS via GoogleService-Info.plist per scheme, on Android via google-services.json in each flavor folder.
Timelines: basic analytics with Firebase + Crashlytics — 3-5 days. Full tracking plan + Amplitude/Mixpanel with funnels and cohorts — 2-3 weeks. Attribution via AppsFlyer with deep linking and fraud protection — 1-2 weeks. Cost is calculated individually based on integration complexity.
What Is Included in Our Work
As part of analytics implementation, we provide:
- Development and approval of a tracking plan with product and marketing teams.
- SDK integration (Firebase, Amplitude, Mixpanel, AppsFlyer) considering your stack (Swift/Kotlin/Flutter/React Native).
- Setup of funnels, cohorts, dashboards, and alerts.
- Automation of symbolication and .dSYM upload via Fastlane.
- Documentation of events and parameters.
- Team training on the analytics platform.
- Two weeks of post-release support and tracking adjustments.
Our experience: 7 years of analytics implementation and over 80 successful projects in mobile development. We guarantee data correctness and transparency at every stage.
Contact us for a consultation on setting up analytics for your app. Request an audit of your current analytics — and we will show you which metrics you are losing.