Recently, a client revamped their checkout — conversion dropped by 20%. We ran an A/B test that showed the old version performed 15% better. At full rollout, this would have cost the business $24,000 monthly. Split testing is the only way to make design decisions based on data, not intuition. Over our work, we've conducted more than 200 experiments for e-commerce stores, landing pages, and SaaS products. The average conversion lift is 15–30%. Some tests brought significant additional profit.
Without A/B tests, every change is a lottery. One of our clients spent $30,000 on a new homepage design that dropped conversion by 8%. A test would have shown this in 2 weeks, saving the entire budget.
Suppose you change a landing page without a test. If conversion drops by 10% with 1,000 daily visitors, that's 100 lost leads per day. In a month — 3,000 leads, each costing an average of $6 — losses are $18,000. The cost of the test is much lower: typically $600 to $1,800.
What problems does A/B testing solve?
Often, teams are confident that a new design or CTA will improve conversion, but statistics show the opposite. For example, we tested a button color change — expecting a 20% increase, we got a 5% drop. Technical errors also occur: incorrect user segmentation, data leakage between variants, improper tracking. Once we found that due to faulty implementation, 30% of users were in both variants — the test had to be restarted. And the classic trap: premature test stopping when the difference seems obvious but the sample size hasn't been reached. According to Nielsen Norman Group, 73% of tests are stopped early, leading to false conclusions.
How we do it
Each experiment follows the scheme: analytics → design → implementation → tracking → analysis. We use a modern stack: React 18, Next.js 14, TypeScript, Node.js, Docker. For data storage — PostgreSQL and Redis. Growthbook allows iterating hypotheses 3x faster compared to VWO: you write logic on the client or server, not through a visual editor. Bayesian statistics further improve decision-making under uncertainty.
Implementation via Vercel Edge Middleware
// middleware.ts
import { NextResponse } from 'next/server';
import type { NextRequest } from 'next/server';
const EXPERIMENT_COOKIE = 'exp_checkout_v2';
const VARIANTS = ['control', 'variant-a', 'variant-b'];
function assignVariant(): string {
const rand = Math.random();
if (rand < 0.34) return 'control';
if (rand < 0.67) return 'variant-a';
return 'variant-b';
}
export function middleware(request: NextRequest) {
const response = NextResponse.next();
const existing = request.cookies.get(EXPERIMENT_COOKIE)?.value;
if (existing && VARIANTS.includes(existing)) {
return response;
}
const variant = assignVariant();
response.cookies.set(EXPERIMENT_COOKIE, variant, {
maxAge: 60 * 60 * 24 * 30,
httpOnly: true,
sameSite: 'lax',
});
response.headers.set('x-ab-checkout', variant);
return response;
}
export const config = {
matcher: ['/checkout/:path*'],
};
// app/checkout/page.tsx
import { cookies, headers } from 'next/headers';
export default function CheckoutPage() {
const variant = headers().get('x-ab-checkout') ??
cookies().get('exp_checkout_v2')?.value ??
'control';
return (
<>
{variant === 'control' && <CheckoutV1 />}
{variant === 'variant-a' && <CheckoutV2OneStep />}
{variant === 'variant-b' && <CheckoutV2TwoStep />}
<ABTracker experiment="checkout_v2" variant={variant} />
</>
);
}
Tracking results
// components/ABTracker.tsx (Client Component)
'use client';
import { useEffect } from 'react';
export function ABTracker({ experiment, variant }: {
experiment: string;
variant: string;
}) {
useEffect(() => {
gtag('event', 'experiment_impression', {
experiment_id: experiment,
variant_id: variant,
});
posthog.capture('$experiment_started', {
'$experiment_id': experiment,
'$variant_key': variant,
});
}, [experiment, variant]);
return null;
}
function trackConversion(variant: string) {
gtag('event', 'purchase', {
experiment_id: 'checkout_v2',
variant_id: variant,
value: orderTotal,
});
}
Statsig: fast integration
// Statsig SDK (server and client parts)
import Statsig from 'statsig-node';
await Statsig.initialize(process.env.STATSIG_SERVER_KEY!);
const experiment = Statsig.getExperiment(
{ userID: userId, email: userEmail },
'checkout_redesign'
);
const checkoutLayout = experiment.get('layout', 'single-page');
const ctaColor = experiment.get('cta_color', 'blue');
// Client side (React SDK)
import { useExperiment } from 'statsig-react';
function PricingCTA() {
const { config } = useExperiment('pricing_cta');
const buttonText = config.get('button_text', 'Get Started');
const buttonVariant = config.get('button_variant', 'primary');
return (
<Button variant={buttonVariant} onClick={() => {
statsig.logEvent('cta_clicked', buttonText);
}}>
{buttonText}
</Button>
);
}
Why statistical significance is critical
Without it, you risk mistaking random fluctuation for a win. Before launch, we calculate the required sample size using the frequentist approach:
# Python: sample size calculation
from statsmodels.stats.power import zt_ind_solve_power
baseline_rate = 0.03
expected_effect = 0.15
lift = baseline_rate * expected_effect
n = zt_ind_solve_power(
effect_size=lift / (baseline_rate * (1 - baseline_rate)) ** 0.5,
alpha=0.05,
power=0.8,
)
print(f"Sample size per variant: {int(n)}") # ~12,000
Rule: do not stop the test before reaching the planned sample size, even if results look good. For a test with a 5% conversion and expected improvement of 10%, you need 6,500 users per variant — that's 2–3 weeks of traffic for an average site.
Tip: don't peek at results daily — it skews statistics. Automatically calculate p-value and stop the test only when the planned sample size is achieved. Use sequential testing if intermediate decisions are needed.
How to choose an A/B testing tool
| Tool | Type | Best for |
|---|---|---|
| Growthbook | Open source / SaaS | Technical teams, self-hosted |
| Statsig | SaaS | Quick start, analytics integration |
| Optimizely | Enterprise SaaS | Large companies, complex experiments |
| VWO | SaaS | Marketing teams without dev |
| Vercel Edge Experiments | PaaS | Next.js on Vercel |
| Custom implementation | - | Full control, minimal overhead |
Which metrics to track in an A/B test
| Metric | Type | Example |
|---|---|---|
| Primary | Target action | Conversion to purchase, sign-up |
| Secondary | Engagement | Time on site, page views |
| Business | Revenue, LTV | Average order value, retention |
| Guardrail | Risk | Bounce rate, errors |
All metrics must be defined before the experiment starts. Track them in GA4: use events experiment_impression and experiment_conversion.
What's included in the work
- Setting up an A/B testing tool for your stack (Growthbook, Statsig, VWO, Optimizely, or custom).
- Implementing variant distribution on backend/edge with consistency guarantees.
- Integrating event tracking into GA4, PostHog, Amplitude.
- Calculating required sample size and test duration.
- Documenting results and recommendations for further experiments.
- Training your team on how to run tests and interpret results.
Our process
- Analytics: study current metrics, identify bottlenecks, formulate a hypothesis.
- Design: choose the tool, define variants and success metrics.
- Implementation: integrate distribution and tracking, set up dashboards.
- Launch: start the test, monitor data correctness.
- Analysis: after reaching sample size — statistical checking, report generation.
Timeline: 2 to 4 business days for a simple test, 5–10 days for a complex one with custom logic. Cost is calculated individually.
Order A/B testing setup and get data-driven conversion growth. Contact us for a consultation on tool selection and experiment execution.







