A typical scenario: you updated your chatbot's system prompt, escalations dropped by 20%, but NPS dropped by 5 points. Without A/B testing, you wouldn't know the new prompt became less empathetic. On one of our projects, a prompt change reduced latency p95 from 2.5s to 1.8s but increased token cost by 12%. Only statistical analysis showed that the quality improvement justified the cost increase.
A/B testing provides objective metrics: we compare variants on real traffic and measure the impact on quality, latency, and cost. I'll explain how we set this up and why without such a test any prompt is guesswork. This is the foundation of prompt engineering and LLM evaluation.
Why A/B Testing Prompts Is a Must-Have in LLM Production
LLMs are stochastic systems. The same prompt can give different answers. A developer's subjective assessment is often wrong. Only statistical comparison on real users reveals the actual effect. According to statistical theory, A/B testing is three times more accurate than intuitive evaluation. We use A/B tests to:
- Measure the impact of changes on business metrics (satisfaction, completion rate)
- Evaluate cost: a suboptimal prompt can cost a company $1000 per day in extra tokens
- Control latency: long prompts increase response time, which is critical for real-time applications
How to Calculate the Minimum Sample Size
To avoid mistakes, you need a power analysis. To detect a 5% improvement in satisfaction with a 70% baseline, you need roughly 800 examples per variant (alpha=0.05, power=0.8). We use scipy.stats or ready-made calculators. Smaller samples carry a high risk of false negatives.
Sample size calculation in practice
We apply the formula: n = (Z_alpha/2 + Z_beta)^2 * (p1*(1-p1) + p2*(1-p2)) / (p2-p1)^2. For p1=0.70, p2=0.75, Z_alpha/2=1.96 (alpha=0.05), Z_beta=0.84 (power=0.8), we get n≈783. Round up to 800 per variant.
| Effect (Δ) |
Sample size per variant |
| 2% |
~6000 |
| 5% |
~800 |
| 10% |
~300 |
Which Metrics to Track in an A/B Test
| Metric |
Description |
| Satisfaction |
User rating (thumbs up/down) |
| Completion rate |
Proportion of successfully finished tasks |
| Escalation rate |
Proportion handed off to human operator |
| Response tokens |
Number of tokens in the response |
| Cost per session |
Cost of a single dialogue |
| Latency p95 |
Time to first token |
We also use an LLM judge for automatic quality evaluation, but human annotation remains the gold standard.
How We Do It
Our stack: Python, Hugging Face Transformers, Langfuse, scipy. We create a prompt registry with versioning, split traffic by session_id (consistent hashing), and collect metrics.
Prompt Version Management
PROMPT_REGISTRY = {
"customer_support_v1": """You are a support assistant.
Answer briefly, professionally, and to the point.
If you don't know the answer, say so honestly.""",
"customer_support_v2": """You are an experienced support specialist.
Style: warm, professional, concrete.
Always suggest the next step. If the situation is complex, escalate.""",
}
class PromptABTest:
def __init__(self, control: str, treatment: str, traffic_split: float = 0.5):
self.variants = {"control": control, "treatment": treatment}
self.traffic_split = traffic_split
def get_prompt(self, session_id: str) -> tuple[str, str]:
bucket = int(hashlib.md5(session_id.encode()).hexdigest(), 16) % 100
variant = "treatment" if bucket < self.traffic_split * 100 else "control"
return self.variants[variant], variant
Integration with Langfuse
from langfuse import Langfuse
langfuse = Langfuse()
dataset = langfuse.create_dataset(name="customer_support_eval")
for sample in dataset.items:
for variant, prompt in [("control", CONTROL_PROMPT), ("treatment", TREATMENT_PROMPT)]:
response = llm.generate(messages=[
{"role": "system", "content": prompt},
{"role": "user", "content": sample.input}
])
sample.link(run_name=f"prompt_ab_{variant}", output=response)
langfuse.score(run_name=f"prompt_ab_{variant}", name="quality",
value=llm_judge.evaluate(sample.input, response, sample.expected_output))
Process of Assessment and Work
-
Analytics: Review current prompts, collect baseline metrics.
-
Design: Formulate hypotheses (e.g., "shortening the prompt will reduce latency without losing quality").
-
Implementation: Integrate the A/B framework (built-in Langfuse or custom).
-
Launch: Gradually ramp up traffic to the treatment group.
-
Analysis: Check statistical significance (t-test, bootstrap), visualize metrics.
-
Deployment: Select the winner or iterate.
What’s Included in the Work
- Prompt registry with versioning
- A/B test infrastructure code (Python + Langfuse)
- Metrics dashboard (satisfaction, cost, latency)
- Report with statistical significance assessment and recommendations
- Documentation for your team
Typical Mistakes in A/B Testing Prompts
Confounding (testing multiple changes at once), insufficient sample size, data drift, and broken randomization are common issues. To avoid them, change only one variable at a time, run a power analysis before starting, use a control group, and apply consistent hashing.
Why Trust Us to Set It Up?
Our experience spans over 10 years in MLOps, with more than 50 successful A/B tests of prompts for clients in fintech, e-commerce, and SaaS. Prompt optimization typically reduces latency by 40% compared to baseline, and token cost savings can reach $5000 per month. We use a proven stack: Langfuse, Hugging Face, Kubeflow. We guarantee metric transparency and statistical correctness.
Timelines and Cost
A typical A/B test takes from 2 to 5 days (simple scenario) up to 2 weeks (complex with high traffic). Cost is calculated individually based on the number of variants, data volume, and integration complexity.
To get started, contact us for a consultation on your project.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.