Automated Model Retraining Setup
A model trained once inevitably degrades: data changes, user behavior evolves, new patterns emerge. For example, in a movie recommendation system, a model trained last year suggests old movies, ignoring new trends. This leads to a 15-20% conversion drop over several months. Our team with 5+ years of experience in MLOps automates model retraining turnkey. We have implemented over 20 projects for recommendation services, fraud monitoring, and scoring. Automatic retraining is a system that monitors model quality and triggers a training cycle upon detecting degradation or on a schedule. You get up-to-date predictions without manual intervention and reduce the risk of business losses.
How to Set Up Retraining Triggers?
There are two approaches: schedule-based and trigger-based. Schedule-based — retraining on a schedule (daily, weekly) regardless of model quality. Simple to implement, predictable, suitable for fast-changing domains (news recommendations, dynamic pricing). Trigger-based — retraining when drift or metric degradation is detected. There are three types of drift: data drift (input data distribution changed), performance drift (metrics on labeled data fell below a threshold), concept drift (the relationship between features and target changed). In practice, a combination is used: soft drift triggers + a hard schedule as a fallback. We help select optimal thresholds based on historical data, e.g., KS-statistic < 0.1 or PSI < 0.2.
What Is Data Drift and How to Detect It?
Data drift is a change in the distribution of the model's input data. Detected by statistical tests: KS-test for numerical features, Chi-square for categorical. For multivariate data, the Population Stability Index is used. We also deploy a drift detector based on scipy.stats.ks_2samp, which automatically signals into MLflow. Drift monitoring saves up to 30% on GPU costs by retraining only when necessary.
Retraining System Architecture
[Monitoring] -> [Drift Detected / Schedule] -> [Data Collection]
-> [Data Validation] -> [Training Job] -> [Evaluation]
-> [A/B Test / Canary] -> [Promotion] -> [Monitoring]
Orchestrators: Airflow, Prefect, Kubeflow Pipelines, Vertex AI Pipelines. The choice depends on your stack: Airflow is convenient for complex DAGs with Python operators, Kubeflow for Kubernetes-native pipelines.
Example Airflow DAG:
from airflow import DAG
from airflow.operators.python import PythonOperator
dag = DAG(
'model_retraining',
schedule_interval='@weekly',
catchup=False
)
check_drift = PythonOperator(
task_id='check_data_drift',
python_callable=run_drift_detection,
dag=dag
)
collect_data = PythonOperator(
task_id='collect_training_data',
python_callable=prepare_dataset,
dag=dag
)
train = PythonOperator(
task_id='train_model',
python_callable=run_training,
dag=dag
)
check_drift >> collect_data >> train
Managing Training Data
Key question: what data to include in retraining? Options: full retrain (all historical data) — stable but expensive in time and computation; rolling window (only the last N days) — the model forgets history but adapts better; incremental learning (fine-tuning on new data without retraining from scratch) — saves resources but not suitable for all algorithms (e.g., linear models — yes, gradient boosting — limited). In practice, a rolling window of 1-3 months is chosen, but for seasonal data, weighted samples are added — older data with lower weight.
| Approach |
Speed |
Adaptation to Trends |
Resources |
| Full retrain |
Low |
Medium |
High |
| Rolling window |
High |
High |
Medium |
| Incremental |
Very high |
High |
Low |
Why Is Pre-release Validation Critical?
An automatically retrained model must not go into production without validation. We use a custom gateway that checks quality and latency:
def validate_new_model(new_model, current_model, test_dataset):
new_metrics = evaluate(new_model, test_dataset)
current_metrics = evaluate(current_model, test_dataset)
# New model must be no worse than current
if new_metrics['auc'] < current_metrics['auc'] * 0.99:
raise ValueError(f"New model AUC {new_metrics['auc']:.4f} "
f"worse than current {current_metrics['auc']:.4f}")
# Check latency
if new_metrics['p95_latency_ms'] > 100:
raise ValueError("Inference too slow")
return True
Without such a gateway, you risk degrading service quality unnoticed. We ensure that every release passes a comparison with the current model by AUC and latency, followed by an A/B test on 10% traffic. Only after confirming metrics does the model receive 100% traffic.
Experiment Management in Auto-retraining
Each retraining cycle is logged in MLflow with: data version (DVC hash), hyperparameters, metrics, training time. This allows retrospective analysis of degradation and identification of when the model started to decline. A typical result: the team transitions from manual retraining "when remembered" (every 2-3 months) to an automatic cycle with weekly updates and always-current quality metrics. Reduction in operational costs by 25% due to automation.
What Is Included in the Work
- Audit of current infrastructure and data (1-2 days)
- Design of trigger scheme and pipeline
- Implementation of DAGs and integration with MLflow
- Setup of drift monitoring (KS-test, PSI, metric drop)
- Validation gateway with A/B testing
- Documentation, team training, 2 weeks of support
Contact us for a free audit. Get a turnkey solution within 5–10 days depending on complexity. Request a consultation to discuss your project.
Definition of concept drift taken from Wikipedia.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.