GPU idling due to network bottlenecks or misconfigured drivers is a common scenario—an 8-node A100 cluster delivering only 30% utilization. The root cause is rarely the cards themselves but the environment: slow interconnect, improper NCCL parameters, or unoptimized storage. We have tuned 20+ GPU clusters over 5+ years and know how to squeeze every ampere-hour. Our approach is simple: eliminate bottlenecks across the entire chain—from drivers to job scheduler.
Why hardware is only half the story
Even top-tier GPUs won't boost performance if the rest of the system is unbalanced. The NVIDIA A100 (80GB SXM) with NVLink (600 GB/s intra-node) and H100 (80GB SXM5) with HBM3 (3.35 TB/s) are powerful, but they require matching infrastructure. Without InfiniBand and a parallel filesystem (GPFS, Lustre), you get 30–50% GPU utilization instead of 85%+. We design clusters around your workloads: for LLMs with tensor parallelism, NVLink speed is critical; for data parallelism, InfiniBand bandwidth between nodes matters most.
How to verify cluster performance after tuning
After every setup we run AllReduce tests using nccl-tests. For 8× A100, expected bandwidth is >280 GB/s at 1GB message size. If lower, we hunt for the bottleneck: NUMA affinity, driver version, switch configuration. We also launch a benchmark training run for your model (e.g., GPT-2 or BERT) and compare throughput against expectations. According to NVIDIA, proper NUMA tuning can yield up to 20% improvement.
Our GPU cluster tuning process
- Audit current infrastructure and requirements—dataset size, model types, training frequency.
- Design—select GPUs, number of nodes, interconnect type, filesystem.
- Install drivers and CUDA—production versions, enable persistence mode, optimize power limits.
- Tune NCCL—fine-tune parameters, test AllReduce bandwidth (target >280 GB/s on 8× A100).
- Integrate with scheduler—Slurm for batch training or Kubernetes + GPU Operator for containerization.
- Monitoring and optimization—DCGM, Prometheus, dashboards with key metrics.
- Documentation and team training—how to launch jobs, diagnose issues.
Example: driver and CUDA installation
# Ubuntu 22.04
apt install linux-headers-$(uname -r) nvidia-driver-535
wget https://developer.download.nvidia.com/compute/cuda/12.3.0/local_installers/cuda_12.3.0_545.23.06_linux.run
sh cuda_12.3.0_545.23.06_linux.run --silent --toolkit
# cuDNN
tar -xvf cudnn-linux-x86_64-8.9.7.29_cuda12-archive.tar.xz
cp cuda/include/cudnn*.h /usr/local/cuda/include
cp cuda/lib64/libcudnn* /usr/local/cuda/lib64
ldconfig
nvidia-smi; nvcc --version
NCCL tuning and interconnect testing
apt install libnccl2 libnccl-dev
git clone https://github.com/NVIDIA/nccl-tests
cd nccl-tests && make
./build/all_reduce_perf -b 1G -e 4G -f 2 -g 8
# Expected: 1GB ~280 GB/s, 4GB ~300 GB/s (algbw)
Interconnect choice: InfiniBand vs Ethernet
| Parameter |
InfiniBand HDR |
Ethernet 100GbE |
| Bandwidth |
200 Gbps |
100 Gbps |
| Latency |
~1 µs |
~3–5 µs |
| Scaling efficiency for LLM |
85–90% |
60–70% |
| RDMA support |
Native |
Requires RoCEv2 |
For multi-node training with tensor parallelism, InfiniBand is mandatory. Ethernet is acceptable only for small clusters (2–4 nodes) or inference.
Configuration comparison: single-node vs multi-node
| Parameter |
Single-node (8× GPU) |
Multi-node (32+ GPU) |
| Interconnect |
NVLink (600 GB/s) |
InfiniBand HDR (200 Gbps) |
| Storage |
Local NVMe |
Parallel FS (Lustre) |
| Scheduler |
Slurm / Kubernetes |
Slurm + gang scheduling |
| Typical task |
Fine-tuning LLaMA 7B |
Pre-training GPT-3 175B |
NCCL tuning details
NCCL uses Tree, Ring, and NVLS algorithms. For H100, we recommend enabling NVLS (NVLink Shared) to speed up all-reduce. The parameter NCCL_ALGO=NVLS can yield 10–15% improvement. Also important is NCCL_IB_HCA to specify InfiniBand interfaces. More details can be found in the official NCCL repository.
Orchestration: Slurm or Kubernetes?
Slurm is the HPC standard, best for long batch jobs with fixed GPU count. Kubernetes + GPU Operator suits containerized, dynamic resource allocation. We help you choose and configure gang scheduling so all GPU pods launch simultaneously.
GPU Operator installation (Helm)
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm install gpu-operator nvidia/gpu-operator --namespace gpu-operator --create-namespace --set driver.enabled=true --set toolkit.enabled=true
Example Slurm job
#!/bin/bash
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=8
#SBATCH --gres=gpu:8
#SBATCH --partition=a100
#SBATCH --time=48:00:00
srun python train.py --nproc_per_node=8 --nnodes=4
Monitoring: DCGM Exporter and metrics
helm install dcgm-exporter nvidia/dcgm-exporter
Key metrics: GPU utilisation (>85%), memory copy utilisation, NVLink bandwidth, power usage.
Common tuning mistakes
- Skipping NUMA affinity configuration—costs 10–20% performance.
- Using a single filesystem partition for both datasets and checkpoints—creates an IO bottleneck.
- Not running AllReduce tests between nodes—often only discovered in production.
- Wrong scheduler parameters (timeout, backfill)—GPUs sit idle.
Results and guarantees
After tuning, your cluster will deliver:
- GPU utilization ≥85% under standard loads.
- Scaling efficiency of 85–90% for multi-node training.
- Documented deployment and monitoring procedures.
We guarantee stable operation and provide support under a service agreement. We will estimate your project within 1–2 days. Contact us for a consultation and get a preliminary assessment. Order tuning and forget about GPU downtime.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.