In a production cluster with 8× NVIDIA A100, we faced GPU utilization of only 40% due to lack of MIG and improper scheduling. After implementing the NVIDIA GPU Operator and gang scheduling, utilization rose to 92%, and idle time dropped by 55%. Our team has 5+ years of experience configuring GPU clusters for AI/ML and has delivered over 20 projects. Without proper GPU scheduling, you risk losing up to 60% of resources — direct losses. Our experience shows that correct Kubernetes configuration for AI/ML workloads guarantees GPU utilization above 90% and reduces GPU infrastructure costs by 30–50%. Have your cluster audited — we'll identify bottlenecks and suggest optimization.
How the NVIDIA GPU Operator Simplifies GPU Management
The operator automatically deploys NVIDIA drivers, container toolkit, and device plugin — all via Custom Resources. Manual configuration on each node is eliminated. The result is a cluster with GPU support ready in one Helm release, 3 times faster than manual setup. The NVIDIA GPU Operator documentation confirms this.
Installation command
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install gpu-operator nvidia/gpu-operator \
--namespace gpu-operator --create-namespace \
--set driver.enabled=true --set driver.version="545.23.06" \
--set toolkit.enabled=true --set devicePlugin.enabled=true \
--set dcgmExporter.enabled=true --set gfd.enabled=true
kubectl get pods -n gpu-operator
kubectl get nodes -o custom-columns='NAME:.metadata.name,GPU:.status.capacity.nvidia\.com/gpu'
Why Proper GPU Scheduling Matters
Typical issues: a job cannot start because whole GPUs are unavailable, even though free MIG slices exist; or one task blocks the entire node. The solution is a combination of MIG, gang scheduling, and priorities.
| Mechanism |
When to use |
Typical utilization |
| Whole GPUs |
Models requiring >24 GB memory |
70–85% |
| MIG (1g.10gb) |
Small batch jobs, inference |
80–90% |
| Gang scheduling |
Distributed training (PyTorch DDP) |
85–95% |
The standard Kubernetes scheduler does not support gang scheduling, causing deadlock in distributed training. Volcano and Kueue solve this: Volcano is mature with fair sharing; Kueue is a native API. For large clusters, Volcano delivers 85–95% utilization versus 60–75% for the standard scheduler — a 30% improvement.
How to Avoid Deadlock in Distributed Training
Gang scheduling with Volcano ensures all pods of a distributed task start simultaneously:
apiVersion: batch.volcano.sh/v1alpha1
kind: Job
metadata:
name: distributed-training
spec:
minAvailable: 4
schedulerName: volcano
tasks:
- replicas: 1
name: master
template:
spec:
containers:
- name: master
image: your-registry/pytorch-trainer:v1
resources:
limits:
nvidia.com/gpu: 8
- replicas: 3
name: worker
template:
spec:
containers:
- name: worker
image: your-registry/pytorch-trainer:v1
resources:
limits:
nvidia.com/gpu: 8
Why Use MIG?
MIG (Multi-Instance GPU) on A100/H100 splits a GPU into up to 7 instances. Switching to MIG increases task placement density and reduces GPU costs. Request in a pod spec:
resources:
limits:
nvidia.com/mig-1g.10gb: 1
How to Set Priorities for ML Tasks
For production inference, set a PriorityClass with high priority (1000); for batch training, low priority (100) with PreemptLowerPriority. Critical services always get GPUs first; training runs on leftover resources.
What's Included in Cluster Setup
- Deployment of the NVIDIA GPU Operator with drivers and DCGM Exporter.
- PriorityClass for inference prioritization.
- Node labels for grouping GPU types (A100, V100, A10G).
- Gang scheduling (Volcano or Kueue).
- Cluster Autoscaler for dynamic GPU node scaling.
- Grafana dashboard with utilization, temperature, and NVLink metrics.
- Documentation and team training (optional).
How We Do It: Phases
-
Infrastructure audit — measure p99 latency, utilization, identify bottlenecks.
- Design — choose a scheduler, define resource profiles and MIG configurations.
- Implementation — Helm releases, monitoring setup, tests with synthetic loads.
- Testing — verify FLOPS, low-priority job preemption, gang scheduling correctness.
- Deployment and support — hand over documentation, activate monitoring, provide SLA.
Common mistakes when configuring a GPU cluster:
- Missing node labels — pods don't know which GPU type is on the node. Solution:
kubectl label node gpu-node-1 nvidia.com/gpu.product=A100-SXM4-80GB.
- Unconfigured Cluster Autoscaler — the cluster doesn't grow under peak load. Specify min/max GPU nodes.
- No monitoring — performance degradation goes unnoticed. DCGM Exporter + Grafana provide full visibility.
Comparison of GPU Scheduling Approaches
| Approach |
Advantages |
Disadvantages |
| Standard K8s scheduler |
Simplicity |
No gang scheduling, low utilization |
| Volcano |
Gang scheduling, fair sharing |
Additional component |
| Kueue |
Native API, easy integration |
Limited policies |
Proper GPU scheduling in Kubernetes is key to efficient AI/ML workloads. We guarantee GPU utilization above 85% and stable operation even under peak load. Contact us for a cluster audit — we'll assess current utilization and suggest optimization. Get a consultation and a detailed implementation plan.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.