Imagine you launched an A/B test via Optimizely but after a week noticed LCP increased by 300 ms due to a third-party script. Or you cannot get raw data — only aggregated graphs. And the license cost for 10,000 visitors is a significant monthly expense. A custom A/B testing platform solves all these problems: you control the code, data, and budget. Savings on licensing fees can reach up to 70%, with a payback period of 3–6 months. Moreover, a custom platform runs 2–3 times faster because there are no external scripts.
We have accumulated over 5 years of experience developing such systems for online stores with 1M+ monthly visitors — more than 50 projects. Our platform is built on a modular principle and easily adapts to any architecture.
Advantages of a Custom A/B Testing Platform
Ready-made tools accelerate the start but become expensive and inflexible at scale. A custom platform pays off after just a few tests — in one project, conversion increased by 15% after the first experiment. With annual use, the savings on licensing can cover the development cost within a few months.
How We Solve Deterministic Assignment
The key requirement of an A/B test: a user must always fall into the same experiment variant. To achieve this, we use hash-based assignment. A hash of user_id and experiment name modulo 100% determines the variant number. The result is saved in the database, ensuring consistency across multiple visits.
Here’s the table schema for storing experiments and assignments:
CREATE TABLE experiments (
id SERIAL PRIMARY KEY,
slug VARCHAR(100) UNIQUE NOT NULL,
name VARCHAR(255) NOT NULL,
description TEXT,
status VARCHAR(20) DEFAULT 'draft', -- draft, running, paused, completed
traffic SMALLINT DEFAULT 100, -- % of traffic participating in the experiment
start_at TIMESTAMPTZ,
end_at TIMESTAMPTZ,
created_at TIMESTAMPTZ DEFAULT NOW(),
updated_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE TABLE experiment_variants (
id SERIAL PRIMARY KEY,
experiment_id INTEGER REFERENCES experiments(id),
slug VARCHAR(100) NOT NULL,
name VARCHAR(255),
weight SMALLINT DEFAULT 50,
config JSONB DEFAULT '{}',
UNIQUE(experiment_id, slug)
);
CREATE TABLE user_assignments (
user_id BIGINT NOT NULL,
experiment_id INTEGER REFERENCES experiments(id),
variant_id INTEGER REFERENCES experiment_variants(id),
assigned_at TIMESTAMPTZ DEFAULT NOW(),
PRIMARY KEY (user_id, experiment_id)
);
The PHP service implements deterministic distribution using crc32 and database persistence:
class ExperimentAssignmentService
{
public function getVariant(int $userId, string $experimentSlug): ?string
{
$experiment = $this->getActiveExperiment($experimentSlug);
if (!$experiment) return null;
$existing = $this->assignmentRepo->find($userId, $experiment['id']);
if ($existing) return $existing['variant_slug'];
$trafficBucket = $this->hashToBucket($userId, $experimentSlug . '_traffic');
if ($trafficBucket >= $experiment['traffic']) return null;
$variantBucket = $this->hashToBucket($userId, $experimentSlug);
$variant = $this->selectVariant($experiment['variants'], $variantBucket);
$this->assignmentRepo->assign($userId, $experiment['id'], $variant['id']);
$this->eventTracker->track($userId, 'experiment.assigned', [
'experiment' => $experimentSlug,
'variant' => $variant['slug'],
]);
return $variant['slug'];
}
private function hashToBucket(int $userId, string $salt): int
{
$hash = crc32($userId . '_' . $salt);
return abs($hash) % 100;
}
private function selectVariant(array $variants, int $bucket): array
{
$cumulative = 0;
foreach ($variants as $variant) {
$cumulative += $variant['weight'];
if ($bucket < $cumulative) return $variant;
}
return end($variants);
}
}
Event Tracking in ClickHouse
We send all significant user actions with experiment context. To avoid slowing down the user experience, events are written asynchronously via a queue.
class ExperimentEventTracker
{
public function track(int $userId, string $event, array $properties = []): void
{
$activeVariants = $this->assignmentRepo->getUserVariants($userId);
$payload = [
'event' => $event,
'user_id' => $userId,
'session_id' => session_id(),
'occurred_at' => now()->toIso8601String(),
'experiments' => $activeVariants,
'properties' => $properties,
];
$this->queue->push(new TrackExperimentEvent($payload));
}
}
Data is stored in ClickHouse — a columnar DBMS optimized for analytical queries. This allows fast conversion calculations and report generation even with millions of events.
Results Computation: Z-test
After collecting data, we use a two-sided Z-test for proportions (Wikipedia). It indicates whether the difference between control and test group conversions is statistically significant. The minimum detectable effect (MDE) is configured in advance — for example, 5% at 80% power.
import numpy as np
from scipy import stats
def calculate_significance(control, treatment):
p1 = control['conversions'] / control['users']
p2 = treatment['conversions'] / treatment['users']
n1, n2 = control['users'], treatment['users']
p_pool = (control['conversions'] + treatment['conversions']) / (n1 + n2)
se = np.sqrt(p_pool * (1 - p_pool) * (1/n1 + 1/n2))
if se == 0: return {'error': 'Insufficient data'}
z = (p2 - p1) / se
p_value = 2 * (1 - stats.norm.cdf(abs(z)))
diff = p2 - p1
se_diff = np.sqrt(p1*(1-p1)/n1 + p2*(1-p2)/n2)
ci = [diff - 1.96*se_diff, diff + 1.96*se_diff]
return {'significant': p_value < 0.05, 'p_value': round(p_value, 6), 'lift': round((p2-p1)/p1*100,2) if p1>0 else None}
To ensure reliable results, we also check for Sample Ratio Mismatch (SRM) — whether the actual user distribution deviates from expected. If the chi-square test p-value is below 0.01, the data is flagged as unreliable.
| Parameter | Calculation |
|---|---|
| Conversion | conversions / users |
| Lift | (p2-p1)/p1 * 100% |
| Confidence Interval | p ± 1.96 * SE |
| Stage | Duration | Result |
|---|---|---|
| Analytics | 1–2 days | Goals, metrics, architecture |
| Design | 2–3 days | DB schema, API, contracts |
| Implementation | 7–10 days | Code, tests |
| Testing | 2 days | Unit, integration, load |
| Deployment | 1–2 days | Release, monitoring |
How Feature Flags Work in A/B Tests?
A/B testing and feature flags are related concepts. We integrate them as follows: an experiment variant contains a JSON configuration that influences feature behavior. For example, variant treatment_a enables {"checkout_steps": 1, "show_trust_badges": true}. The frontend or backend code simply reads this config and changes behavior.
$variant = $experimentService->getVariant($userId, 'checkout-redesign');
$config = $experimentService->getVariantConfig('checkout-redesign', $variant);
$checkoutSteps = $config['checkout_steps'] ?? 3;
What's Included
- Development of assignment service with hash-based distribution and unit tests
- Event tracking with asynchronous ClickHouse writes
- Statistical significance computation (Z-test, confidence intervals)
- Admin panel for launching and monitoring experiments
- Integration documentation and team training
- First month of pilot launch support
Our Work Process
- Analytics — we analyze your goals, metrics, current architecture (1–2 days)
- Design — prepare DB schema, API, contracts (2–3 days)
- Implementation — write code, write tests (7–10 days)
- Testing — unit tests, integration testing, load (2 days)
- Deployment — deploy to your environment, configure monitoring (1–2 days)
For any inquiries, contact us — we will help estimate the work volume and calculate savings.
Timeline and Budget
Estimated development time: 14 to 21 days. The cost is calculated individually depending on integration complexity and additional requirements. Order a custom platform today — we will provide a proposal within 2 business days.
We guarantee code quality: we use code reviews, test coverage of at least 80%, and provide a 30-day bug fix warranty after launch. Get a consultation and find out how a custom A/B testing platform can improve your conversions without performance trade-offs.







