Setting Up MLflow for Experiment Tracking in Production
Imagine a team of five data scientists running 50 experiments every week. Parameters saved in Jupyter notebook comments, metrics in Google Sheets, artifacts in random Google Drive folders. When after a month you need to reproduce the best model with F1=0.924, nobody remembers the exact hyperparameters: learning_rate 0.001, batch_size 32 or 64? Sound familiar? MLflow is an open-source platform that centralizes this chaos: logs parameters, metrics, artifacts, and models for each run. Our engineers (certified ML specialists with five years of experience) deploy MLflow turnkey with PostgreSQL and S3, turning disparate records into a structured system. We've implemented MLflow in 30+ projects, guaranteeing 99.9% SLA for production infrastructure.
Why MLflow became the MLOps standard
MLflow is backed by a community of 10,000+ commits, supports any ML framework (PyTorch, TensorFlow, Hugging Face, scikit-learn), works with Kubernetes, and has APIs for CI/CD. Alternatives like homegrown solutions require constant maintenance and often break. MLflow is a reliable tool that saves teams up to five hours per week on logging. Unlike Bash or Python scripts, MLflow provides a unified UI, flexible API, and integration with 50+ tools.
How to set up MLflow for production
Backend selection
| Backend |
Performance |
Recommendation |
| SQLite |
Low (1–2 users) |
Development |
| PostgreSQL |
High (up to 100+) |
Production |
| MySQL |
High |
Alternative |
For production we choose PostgreSQL — it handles dozens of concurrent sessions, supports SQL queries, and scales easily.
Artifact storage
| Type |
Example |
Suitable for |
| Local filesystem |
./mlruns |
Development |
| S3-compatible |
AWS S3, Yandex Object Storage |
Production |
Deployment with Docker
Here's a typical Docker Compose configuration we use in production:
services:
mlflow:
image: ghcr.io/mlflow/mlflow:v2.14.0
ports: ["5000:5000"]
environment:
- MLFLOW_S3_ENDPOINT_URL=https://storage.yandexcloud.net
- AWS_ACCESS_KEY_ID=${YC_ACCESS_KEY}
- AWS_SECRET_ACCESS_KEY=${YC_SECRET_KEY}
command: >
mlflow server
--backend-store-uri postgresql://mlflow:${DB_PASS}@postgres:5432/mlflow
--default-artifact-root s3://mlops-bucket/mlflow
--host 0.0.0.0
How to automate logging with autologging?
Simply call mlflow.autolog() at the beginning of your script. Autologging automatically records parameters, metrics, artifacts, and the model for PyTorch, TensorFlow, scikit-learn, Hugging Face. For fine-tuning, use mlflow.sklearn.autolog(log_models=True). This eliminates manual coding for each experiment.
Example autologging with different frameworks
mlflow.autolog()
mlflow.sklearn.autolog(log_models=True, log_input_examples=True)
mlflow.pytorch.autolog(log_every_n_epoch=1)
mlflow.transformers.autolog()
Implementation process
-
Analysis — we discuss your infrastructure, number of users, model types.
-
Design — we choose backend, storage, security level.
-
Implementation — deploy MLflow, configure autologging, integrate with Git and CI/CD.
-
Testing — conduct load testing, verify 99.9% SLA.
-
Deployment — hand over documentation, access credentials, and train your team.
Case study
A computer vision company with 5 data scientists previously logged experiments in Google Sheets. After implementing MLflow, experiment comparison time dropped from two hours to 10 minutes, and the share of reproducible runs rose to 90%. We confirm results with SLA — the standard that ensures transparency.
Timeframe and scope
Roughly 2 weeks to 2 months depending on complexity. Includes: MLflow installation with PostgreSQL and S3, autologging setup, documentation, team training (remote or on-site), and support during the rollout phase (1 month).
Local run and example experiment
For a quick start:
pip install mlflow
mlflow server --host 0.0.0.0 --port 5000
Example experiment logging:
import mlflow
mlflow.set_tracking_uri("http://mlflow-server:5000")
with mlflow.start_run():
mlflow.log_param("learning_rate", 0.01)
mlflow.log_params({"batch_size": 32, "epochs": 10, "optimizer": "adam"})
for epoch in range(10):
train_loss = train_one_epoch(model, train_loader)
val_loss, val_acc = evaluate(model, val_loader)
mlflow.log_metrics({"train_loss": train_loss, "val_loss": val_loss, "val_acc": val_acc}, step=epoch)
mlflow.log_metric("test_f1", 0.924)
mlflow.log_artifact("confusion_matrix.png")
mlflow.log_dict({"feature_importance": feature_imp}, "artifacts/feature_importance.json")
mlflow.sklearn.log_model(model, "model", registered_model_name="my-classifier")
MLflow Model Registry: managing model versions in production
MLflow includes a built-in model registry — a tool for managing the model lifecycle from experiment to production. Each registered model goes through stages: Staging → Production → Archived.
Transition a model to Production via API:
client = mlflow.tracking.MlflowClient()
client.transition_model_version_stage(
name="my-classifier",
version=3,
stage="Production"
)
Automatic CI/CD rules for promotion: F1 on hold-out set > 0.92 and p99 latency < 100 ms. If a new version fails, it remains in Staging, production is not affected.
For teams of 5+, the model registry reduces the time to roll out a new version from 2–3 hours (manual artifact transfer) to 10–15 minutes via an automated pipeline. Each model version is stored with full metadata: hyperparameters, test-set metrics, dataset hash, and a link to the experiment run.
Conclusion
MLflow is a proven tool for experiment tracking that pays off within the first few weeks. Contact us for an assessment of your infrastructure — we'll prepare an architecture in 2 days. Get a consultation right now. We guarantee results: 99.9% SLA and transparent documentation.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.