Inference for AI models in production often costs more than expected. LLM API fees, GPU compute, and managed inference services scale non-linearly with traffic. A project with hundreds of thousands of daily requests can easily run tens of thousands of dollars per month. Clients in fintech and e-commerce have reported that 60–80% of their budget goes to inference. They often pay for unnecessarily high-quality responses for simple queries like "What's the dollar exchange rate today?" We systematically optimize these costs without degrading quality. Our experience shows that combining model routing, semantic caching, and quantization reduces the bill by 40–70%. This is confirmed in over 50 deployments. LLM inference optimization is critical for cost efficiency. Our methods significantly reduce LLM API costs. Model quantization is another key technique. Prompt compression drives token savings.
How model routing cuts costs
Model routing dynamically selects a model for each request. Simple tasks go to cheap models, complex ones to powerful models. The method yields 50–70% cost savings. Quality degradation is less than 5%. This is far better than using a single model for all requests.
class IntelligentModelRouter:
def route_request(self, request: dict) -> str:
query = request['query']
complexity = self.complexity_estimator.estimate(query)
if complexity < 0.3:
return "gpt-4o-mini" # $0.15/1M input tokens
elif complexity < 0.7:
return "gpt-4o" # $5.00/1M input tokens
else:
return "gpt-4-turbo" # $10.00/1M input tokens
def estimate_complexity(self, query: str) -> float:
# Heuristics: length, presence of code, math, multi-step reasoning
features = [
len(query.split()) / 200,
1.0 if any(kw in query for kw in ['calculate', 'code', 'step-by-step']) else 0,
1.0 if '```' in query else 0,
]
return min(np.mean(features) * 1.5, 1.0)
What is semantic caching?
Semantic caching stores responses and their embeddings. On a new query, cosine similarity is computed. If above 0.95, the cached response is returned. This covers paraphrases and achieves a 20–40% hit rate without quality loss.
from redis import Redis
import numpy as np
class SemanticCache:
def __init__(self, similarity_threshold: float = 0.95):
self.redis = Redis()
self.threshold = similarity_threshold
self.embedder = SentenceTransformer('all-MiniLM-L6-v2')
def get(self, query: str) -> str | None:
query_emb = self.embedder.encode(query)
cached_keys = self.redis.keys("cache:*")
for key in cached_keys:
cached_data = json.loads(self.redis.get(key))
cached_emb = np.array(cached_data['embedding'])
similarity = np.dot(query_emb, cached_emb) / (
np.linalg.norm(query_emb) * np.linalg.norm(cached_emb)
)
if similarity > self.threshold:
return cached_data['response']
return None
def set(self, query: str, response: str, ttl: int = 3600):
key = f"cache:{hashlib.md5(query.encode()).hexdigest()}"
self.redis.setex(key, ttl, json.dumps({
'embedding': self.embedder.encode(query).tolist(),
'response': response
}))
Prompt compression and quantization
Context compression via LLMLingua reduces tokens by 30–50% with 1–5% quality degradation. 4-bit quantization via bitsandbytes cuts VRAM by 4x: Llama-2-70B requires ~35GB instead of 140GB. Both methods suit self-hosted scenarios.
# LLMLingua for compressing long contexts
from llmlingua import PromptCompressor
compressor = PromptCompressor(
model_name="microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank",
use_llmlingua2=True
)
compressed = compressor.compress_prompt(
context,
rate=0.5, # Compress to 50% tokens
force_tokens=['\n', '?']
)
# 4-bit quantization via bitsandbytes (GPTQ)
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4"
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-70b-chat-hf",
quantization_config=quantization_config,
device_map="auto"
)
Step-by-step plan for implementing model routing
- Audit current requests: collect logs for a week, cluster by complexity (length, code presence, number of steps). Define thresholds for switching between models.
- Choose models: pick 2–3 models with different price/quality trade-offs (e.g., GPT-4o-mini, GPT-4o, GPT-4-turbo).
- Implement the router: use the
IntelligentModelRouter class above. Tune heuristics for your domain.
- A/B testing: run the router on 10% of traffic, compare response quality and costs. Optimize thresholds.
- Gradual rollout: increase traffic share to 100% once metrics are stable.
Expected cost savings
| Method |
Cost reduction |
Quality degradation |
| Model routing |
50–70% |
<5% |
| Semantic caching |
20–40% |
0% |
| Prompt compression |
30–50% |
1–5% |
| 4-bit quantization |
40–60% (self-hosted) |
1–3% |
| Batch inference |
30–50% |
0% |
Combining model routing and semantic caching gives the greatest effect without risk for most production scenarios. For example, an e-commerce client with 2M daily requests cut monthly costs from $28,000 to $8,400. Another fintech project reduced costs from $50,000 to $15,000 using the same methods. Typical monthly savings range from $10,000 to $50,000 for enterprise clients. Deployment evaluations show average savings of 60%.
Comparison of quantization tools
| Tool |
Model support |
Format |
Inference speed (tokens/s) |
| bitsandbytes |
HuggingFace, LLaMA, Mistral |
INT8/INT4 |
~85% of FP16 |
| GPTQ |
AutoGPTQ, HuggingFace |
INT4 |
~95% of FP16 |
| AWQ |
vLLM, HuggingFace |
INT4 |
~90% of FP16 |
More details on quantization can be found in the bitsandbytes documentation.
What our optimization service includes
- Cost audit: analyze spending structure by model, cache hit rate, context length, GPU utilization.
- Recommendations: select model routing strategy, configure cache, set up batch inference.
- Implementation: deploy code, tune pipeline, optimize infrastructure.
- Documentation: describe the schema, monitoring instructions, rollback plan.
- Training: we provide team training on the new pipeline.
- Access: we provide access to monitoring dashboards.
- Support: one month of post-deployment consultation.
Additional capabilities
We can also integrate prompt compression and quantization for self-hosted models. Deep GPU utilization analysis via NVIDIA Nsight and MLflow identifies bottlenecks.
We will assess your project in 2–3 days. Order an audit — we guarantee transparent pricing and a clear savings plan. Our team has 5+ years of MLOps experience and has delivered over 50 inference optimization projects for fintech, e-commerce, and SaaS clients. With over 5 years in the market and 50+ projects delivered, we are a trusted partner in AI optimization. AI inference cost optimization is our specialty. Get a free consultation to discuss your case.
Our optimization covers all key techniques: model routing, semantic caching, prompt compression, model quantization, batch inference, and GPU utilization. Typical monthly savings for enterprise clients range from $10,000 to $50,000.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.