Your data science team trained a new model, shows AUC lift on test, but business won't deploy without proof on real data. A/B testing ML models is the only way to reliably measure business impact. A metric on a test dataset shows accuracy but doesn't answer the key question: will it bring more money or better UX? A properly configured A/B test provides a statistically sound answer with risk control. Our certified MLOps engineers with proven experience have delivered over 100 experiments across 30+ projects — from banking to e-commerce. Inference cost savings from correct model selection can reach 30% of the budget, translating to an average of $15,000 per month for typical enterprise deployments. Seldon Core sets up 3x faster than custom routing on Nginx. For instance, in a recent project for a retail client, we measured a 12% increase in average order value with a p-value of 0.003. With A/B testing ML models, you can confidently assess business impact. Contact us for an audit of your infrastructure.
Why ML A/B is harder than classic?
In classic A/B, users are randomly assigned to groups once. In ML A/B, additional complexities arise:
- Novelty effect: users react to novelty, not model quality.
- Long-term effects: recommendation systems influence behavior not visible in short tests.
- Carryover effect: previous prediction result affects current behavior.
- Network effects: in collaborative systems, one user's behavior influences others.
These effects require careful experimental design. For example, to combat novelty, use ramp-up with gradual traffic increase from 5% to 50%.
A/B Architecture for ML
Traffic Split Levels
| Method |
Description |
When to use |
| User-level split |
Same user always gets one model version |
Personalization, recommendations |
| Request-level split |
Each request randomly directed to a version |
Stateless services (search, pricing) |
| Cohort-based split |
Split by user segments |
Balance demographic characteristics |
Traffic routing:
import hashlib
def get_model_version(user_id: str, experiment_id: str) -> str:
# Deterministic hashing for stable assignment
hash_key = f"{experiment_id}:{user_id}"
hash_value = int(hashlib.md5(hash_key.encode()).hexdigest(), 16)
bucket = hash_value % 100 # 0-99
if bucket < 50: # 50% traffic
return "model_v2"
else:
return "model_v1_control"
Tools
Nginx / Envoy — infrastructure-level routing by headers or weights.
Seldon Core / KServe — Kubernetes-native inference with built-in A/B. Seldon Core is a Kubernetes-native inference platform for ML models, ideal for Kubernetes inference workloads:
apiVersion: machinelearning.seldon.io/v1
kind: SeldonDeployment
spec:
predictors:
- name: control
traffic: 70
graph:
name: model-v1
- name: treatment
traffic: 30
graph:
name: model-v2
Feature flags (LaunchDarkly, Unleash) — flexible experiment management without redeployment.
Tools Comparison
| Tool |
Traffic management |
Statistics |
Notes |
| Seldon Core |
Built-in A/B, canary |
Yes (prometheus) |
Kubernetes-native, supports any ML framework |
| KServe |
InferenceGraph with % traffic |
Not built-in |
Simple configuration, Knative integration |
| LaunchDarkly |
Feature flags |
No |
Flexible management, ML-infrastructure independent |
How to choose traffic split method?
Choice depends on service nature and goals. User-level split is good for personalization but may distort long-term effects. Request-level split is simpler but only for stateless loads. Cohort-based split gives control over biases but requires prior segmentation. We help select the optimal scheme for your architecture.
Statistical Methodology
Metrics for ML A/B:
-
Primary metric: business metric (conversion, ARPU, retention)
-
Guardrail metrics: latency, error rate — must not degrade
-
Secondary metrics: proxy indicators (CTR, engagement)
Sample size and test power:
To detect an effect of 2% at a baseline conversion of 5%, significance level α=0.05 and power 80%, approximately 15,000 users per group are needed. Conduct power analysis to ensure adequate sample size. Use a power calculator (scipy.stats.norm or online tools) before launching. Our experiments typically involve sample sizes from 10,000 to 1 million users per group. For more details, see Wikipedia: Statistical hypothesis testing.
Stopping the test:
- Do not stop early due to preliminary results (peeking problem)
- Minimum duration: 1-2 weeks to account for daily and weekly patterns
- Use sequential testing (e-values) if you need to make decisions earlier
Analysis of Results
from scipy import stats
control_conversions = [0, 1, 0, 1, ...] # 0/1 per user
treatment_conversions = [0, 1, 1, 0, ...]
# t-test for continuous metrics
t_stat, p_value = stats.ttest_ind(control_conversions, treatment_conversions)
# Chi-squared for binary metrics
from scipy.stats import chi2_contingency
contingency = [[control_success, control_fail],
[treatment_success, treatment_fail]]
chi2, p_value, dof, expected = chi2_contingency(contingency)
print(f"Relative lift: {(treatment_rate - control_rate) / control_rate:.2%}")
print(f"P-value: {p_value:.4f}")
print(f"Statistically significant: {p_value < 0.05}")
What's included in the work
- Audit of current infrastructure and metrics
- Design of traffic split scheme (user-level, request-level, cohort)
- Routing setup via Nginx/Envoy or Seldon Core
- Integration of feature flags for experiment management
- Calculation of required sample size and test duration
- Power analysis to validate sample sizes
- Monitoring of guardrail metrics (latency, error rate)
- Documentation of results and knowledge transfer to the team
Process
- Analytics — study your infrastructure, goals, and available data.
- Design — select split type, metrics, and tools.
- Implementation — set up routing, integrate feature flags, connect monitoring.
- Test — launch a pilot experiment on 10% traffic, verify correctness.
- Deploy — full-scale test with automatic stop rules.
Timelines and Cost
Setup timelines for A/B testing — from 2 to 4 weeks depending on infrastructure complexity. Cost is calculated individually after the audit. Our proven methodology guarantees reliable results. Get a consultation — contact us to assess your project. A properly configured A/B test enables you to make model deployment decisions based on data with a measurable level of confidence, not intuition.
Typical primary metrics include: conversion rate (CVR), average revenue per user (ARPU), retention rate at D1/D7/D30, click-through rate (CTR). Guardrail metrics often include p99 latency, error rate, and cost per inference.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.