Implementing an LLM Prompt Management Platform
We worked on a project where 80 prompts were scattered throughout the code: every change required a full application deployment, and rollback meant searching through git and a new release. After implementing a Prompt Registry, management time dropped by 80%, and token costs fell by 25%. But that's not the limit: with proper A/B testing and versioning, savings can reach 40%. Many companies still edit prompts manually, which leads to errors and overspending on LLM tokens. Implementing a full-fledged platform gives you control over every prompt, and integration with any LLM provider takes 2 to 6 weeks.
How does a prompt management platform solve problems?
Without a centralized registry, you don't see which prompt is used where, there's no versioning, and testing is done manually. The platform solves this through three components: a registry with hash versions, an API for deployment, and a metrics dashboard.
Let's compare approaches:
| Parameter |
Without platform |
With platform |
| Storage |
Hardcoded in code |
In registry with versions |
| Changes |
Requires CI/CD deployment |
Via API in 1 second |
| Rollback |
Git search + deployment |
One click |
| Metrics |
None |
A/B tracking, p99 latency, tokens |
| Security |
Full access |
Roles, approvals |
A/B testing on the platform identifies the best prompt 3 times faster. Each new prompt is first tested on 10% of traffic — response quality and tokens are compared. A sample of 1000 requests provides statistical significance.
Why is prompt versioning critical for LLM applications?
Even a small change can cause hallucinations or increase token usage. Without versioning, you don't know what changed or when. In one project, a production prompt was accidentally overwritten — quality dropped by 30%, and the fix took a day. With versioning, each version stores the hash, author, timestamp, and status (reviewed/deployed). OpenAI recommends using versioning to track prompt changes in production environments.
Prompt Registry Architecture
from dataclasses import dataclass
from typing import Optional
import hashlib
@dataclass
class PromptVersion:
id: str
name: str
version: int
content: str
variables: list[str] # Variables in the prompt {{variable}}
model: str
temperature: float
max_tokens: int
created_by: str
created_at: datetime
metadata: dict
hash: str = None
def __post_init__(self):
self.hash = hashlib.sha256(self.content.encode()).hexdigest()[:8]
class PromptRegistry:
def __init__(self, db_connection, cache):
self.db = db_connection
self.cache = cache
def register(self, name: str, content: str, model: str = "gpt-4o",
temperature: float = 0.0, **kwargs) -> PromptVersion:
"""Register a new prompt version"""
last_version = self.db.get_latest_version(name)
version_num = (last_version.version + 1) if last_version else 1
variables = self._extract_variables(content) # {{var}} → ['var']
prompt = PromptVersion(
id=str(uuid.uuid4()),
name=name,
version=version_num,
content=content,
variables=variables,
model=model,
temperature=temperature,
max_tokens=kwargs.get('max_tokens', 1000),
created_by=kwargs.get('created_by', 'system'),
created_at=datetime.utcnow(),
metadata=kwargs.get('metadata', {})
)
self.db.save(prompt)
return prompt
def get(self, name: str, version: str = "latest",
environment: str = "production") -> PromptVersion:
"""Retrieve a prompt by name and version"""
cache_key = f"prompt:{name}:{version}:{environment}"
cached = self.cache.get(cache_key)
if cached:
return cached
if version == "latest":
prompt = self.db.get_latest_deployed(name, environment)
else:
prompt = self.db.get_by_version(name, int(version))
self.cache.set(cache_key, prompt, ttl=300)
return prompt
def render(self, name: str, variables: dict, **kwargs) -> str:
"""Retrieve and render a prompt"""
prompt = self.get(name, **kwargs)
rendered = prompt.content
for var, value in variables.items():
rendered = rendered.replace(f"{{{{{var}}}}}", str(value))
# Check: all variables filled?
missing = [v for v in prompt.variables if f"{{{{{v}}}}}" in rendered]
if missing:
raise ValueError(f"Missing variables: {missing}")
return rendered
Deploying Prompts Across Environments
class PromptDeploymentManager:
def deploy(self, prompt_name: str, version: int,
environment: str, require_review: bool = True):
prompt = self.registry.get_by_version(prompt_name, version)
if require_review and not prompt.is_reviewed:
raise ValueError("Prompt requires review before deployment to production")
# Record deployment
self.db.create_deployment(
prompt_id=prompt.id,
environment=environment,
deployed_by=current_user(),
deployed_at=datetime.utcnow()
)
# Invalidate cache
self.cache.delete(f"prompt:{prompt_name}:latest:{environment}")
# Webhook notification
self.notify_team(
f"Prompt '{prompt_name}' v{version} deployed to {environment}"
)
Prompt Quality Metrics
For each prompt, we measure: p99 latency (target < 500 ms), token usage per request (15-25% savings after optimization), output quality score (LLM-judge rating 0-1), precision@k for RAG. Integration with LangSmith or W&B allows comparing versions and making data-driven decisions.
Example metrics dashboard:
| Metric |
Current v3 |
Previous v2 |
Change |
| p99 latency |
420 ms |
680 ms |
-38% |
| Tokens/request |
2450 |
3100 |
-21% |
| Quality score |
0.92 |
0.85 |
+8% |
| Hallucination rate |
2.1% |
4.5% |
-53% |
Token cost savings after optimization average $5,000–$15,000 per month for projects with 1 million tokens/day. For more intensive systems, savings reach $20,000 monthly. Implementation cost is recouped in 2–3 months due to reduced API expenses.
How does A/B testing of prompts improve response quality?
A/B testing allows comparing two prompt versions on real requests. We set up traffic splitting (e.g., 10% on the new version) and collect metrics: response quality (LLM judge score), tokens, latency. After reaching statistical significance (usually 1000 requests), the winner is automatically deployed. A/B testing cuts the time to choose the best prompt by a factor of 3.
What's included in the work
- Audit of current prompts: inventory, assessment of impact on business metrics.
- Registry schema design: data model, metadata, access rights.
- Integration development: API for all environments (dev/staging/prod), webhook notifications.
- Monitoring implementation: metric tracking, alerts on degradation.
- Documentation and team training: process descriptions, role model.
- Support during operation: platform warranty, optimization consulting.
Implementation process
- Analytics: measure current state — number of prompts, change frequency, latency and token usage.
- Design: describe the registry architecture, choose vector DB (ChromaDB, Qdrant) and cache (Redis).
- Implementation: configure prompt registry, integrations with LLM providers, CI/CD pipeline.
- Testing: A/B testing on staging, rollback check, load testing (1000+ RPS).
- Deployment: phased rollout to production, metric monitoring first 48 hours.
Timeline: 2 to 6 weeks depending on complexity of integrations and number of environments. We'll evaluate the project in 1-2 days after the audit.
We guarantee transparency of all changes and a reduction in prompt management time by 80%.
Get a consultation — we'll explain how to adapt the platform to your stack. Order an audit of your prompts — we'll estimate the savings potential in 1-2 days.
Case: Optimizing a support prompt
For a fintech client, we optimized a chat bot prompt: removed unnecessary instructions, added few-shot examples. Result: p99 latency dropped from 1.2 s to 400 ms, tokens per request fell from 3000 to 1800, and answer accuracy increased from 78% to 94%.
MLOps: Infrastructure for Training, Deploying, and Monitoring ML Models
The model is trained, metrics — F1 0.94 on validation. Three months later in production, quality drops by 12%. No one knows when — there is no monitoring. It's impossible to retrain quickly — the training script is in a Jupyter notebook of a data scientist who has already left. Data for retraining is collected manually from three disparate systems. About half of the projects come to us with this pain. We build a turnkey MLOps platform: from experiment tracking to automatic deployment and data drift monitoring. We will assess your infrastructure in 1–2 weeks, and in 4–6 weeks you will get a basic MLOps core running in production. Our team has 10+ years of experience in ML infrastructure, over 50 implementations.
How does MLOps infrastructure benefit your ML projects?
Experiment Tracking and Reproducibility
Without tracking, an ML project turns into chaos: it's unclear which checkpoint is better, which hyperparameters were used, which dataset. Reproducing a result a month later is a quest.
Why is experiment tracking the foundation of reproducibility?
MLflow is an open source standard for tracking. It logs parameters, metrics, artifacts (models, graphs), and code. MLflow Model Registry is a centralized model storage with versioning and lifecycle stages (Staging → Production → Archived). Deployment via MLflow Serving or integration with external systems.
Typical initialization in code:
import mlflow
mlflow.set_experiment("fraud-detection-v2")
with mlflow.start_run():
mlflow.log_params({"learning_rate": 3e-4, "batch_size": 64, "epochs": 10})
mlflow.log_metric("val_f1", val_f1, step=epoch)
mlflow.pytorch.log_model(model, "model")
This is the minimum. In production, we add logging of system metrics (GPU utilization, memory), dataset (hash, version), code (git commit hash). Weights & Biases — richer UI, collaboration features, sweep for hyperparameter optimization. MLflow — for on-premise deployment without external dependencies.
DVC (Data Version Control) — versioning of data and models on top of git. Data is stored in S3/GCS/Azure Blob, only metadata (hashes) in git. dvc repro reproduces the entire pipeline from raw data to metrics.
To ensure reproducibility of training, fix random seeds (torch.manual_seed, numpy.random.seed, random.seed) and record them in experiment metadata. Without this, debugging irregular results is painful. Log the dataset version (DVC hash) and git commit — then any experiment can be reproduced down to the byte.
Pipeline Orchestration: Kubeflow, Airflow, Prefect
A pipeline orchestrator becomes necessary when: A 100-line training script in cron is fine for simple tasks. But as soon as you have a multi-step pipeline (data loading → preprocessing → feature engineering → training → validation → deployment if quality above threshold), you need an orchestrator with retry logic, visualization, and alerts.
Kubeflow — Kubernetes-native orchestrator for ML (see Kubeflow). Each step is a Docker container. Supports parallel steps, conditional branches, artifacts between steps. Integrates with Katib (AutoML), KServe (serving), Feast (feature store).
Apache Airflow — more general DAG orchestrator. Wide ecosystem of operators (S3, Spark, DBT, Kubernetes). Easier to deploy if Airflow already exists in the company.
Prefect / Metaflow — less boilerplate. Prefect 2.x with @flow and @task decorators — quick start for small teams.
Typical training pipeline architecture on Kubeflow:
- Data ingestion component — fetches data from S3/DB, validates schema via Great Expectations
- Preprocessing component — transformations, normalization, train/val/test split
- Training component — training on GPU, logging to MLflow
- Evaluation component — metric calculation, comparison with baseline in Model Registry
- Conditional deployment — deploy only if new model is better than current by >2% F1
Each component is a separate Docker image. Pipeline is versioned in git. Scheduled run (retraining once a week on new data) or manual.
Model Registry and Lifecycle Management
Model Registry is not just a checkpoint store. It is a centralized system that knows:
- Which model is currently in production (and with what metrics)
- History of all versions with training parameters
- Metadata: dataset, git commit, validation results
- Lifecycle stage: None → Staging → Production → Archived
MLflow Model Registry — standard. For enterprise — Vertex AI Model Registry (GCP), SageMaker Model Registry (AWS), Azure ML Model Registry.
Model promotion through stages: automatically move model to Staging after successful eval, then manual or automatic (during A/B test) promotion to Production. Rollback — switch to previous Production version in seconds.
Serving: From FastAPI to Triton Inference Server
Simple case. FastAPI + PyTorch/ONNX on one server — 80% of production ML deployments are exactly that. Sufficient for most tasks with load up to 100 req/s.
from fastapi import FastAPI
import onnxruntime as ort
app = FastAPI()
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
@app.post("/predict")
async def predict(request: PredictRequest):
inputs = preprocess(request.text)
outputs = session.run(None, {"input_ids": inputs})
return {"label": postprocess(outputs)}
Triton Inference Server — production standard for high loads (500+ req/s). Dynamic batching, concurrent model execution, model ensemble. Supports TensorRT, ONNX, PyTorch TorchScript, TensorFlow SavedModel.
KServe — Kubernetes-native ML serving with autoscaling, canary deployments, A/B testing out of the box. Scale-to-zero for inactive models — savings on infrastructure up to 40% annually for a project with 10 models.
Monitoring: Data Drift, Model Drift, Infrastructure Metrics
Monitoring — what is usually done last and regretted first. Three levels.
Infrastructure monitoring. Latency (P50/P95/P99), throughput (req/s), error rate (4xx, 5xx), GPU/CPU utilization. Prometheus + Grafana — standard. Alert when P99 latency > threshold or error rate > 1%.
Data drift monitoring. Distribution of input data changes over time. Detect via PSI (Population Stability Index) for numerical features: PSI > 0.2 — strong drift. Chi-squared test for categorical, Kolmogorov-Smirnov test for continuous. Evidently AI — open source library with ready-made drift tests.
Model drift monitoring. If ground truth is delayed (e.g., we know conversion after a week) — monitor real metrics. If not — surrogate metrics: distribution of prediction scores, proportion of confident predictions.
Alerting. Three levels: INFO (minor drift, log it), WARNING (significant, notify team), CRITICAL (quality dropped below threshold — automatic switch to fallback model).
Why is data drift monitoring important?
Without it, you learn about model degradation only from user complaints or ringing SLA. A drift alert allows you to retrain the model in advance, before errors start causing losses. In one of our projects, PSI monitoring detected drift 2 days after a data source change — this saved the campaign.
| Common Mistake |
Consequences |
Solution |
| Lack of data versioning |
Irreproducible experiments |
Implement DVC or similar |
| Manual model deployment |
Human errors, slow rollback |
Automate CI/CD pipeline |
| Monitoring only by business metrics |
Late drift detection |
Add data drift monitoring (PSI, KS) |
Feature Store
Feature Store solves the training-serving skew problem. If preprocessing during training and inference is implemented in two different places — divergence is inevitable.
A Feature Store is needed when:
- Several models use the same features
- Features are computed from streaming data (real-time)
- Large team with different people on feature engineering and model training
Feast — open source Feature Store. Offline store (S3 + Parquet) for training, online store (Redis, DynamoDB) for low-latency inference. Feature definitions as code, materialization job syncs offline → online.
Tecton (commercial), Vertex AI Feature Store (GCP), SageMaker Feature Store (AWS) — managed options with less ops overhead.
CI/CD for ML
ML CI/CD is regular CI/CD plus specific ML steps.
ML-specific checks in CI:
- Reproducibility check: run training with a fixed seed, result must match
- Data validation: Great Expectations or Pandera on schema/distribution checks
- Model performance check: automatic eval on holdout, block merge if degradation > threshold
- Latency regression test: inference must meet SLA
GitOps for deployment. Merge to main → CI triggers training → eval → if passes → automatic deployment to Staging → smoke tests → manual promotion to Production or automatic upon successful canary.
Tools: GitHub Actions / GitLab CI for CI, ArgoCD for GitOps deployment on Kubernetes.
What's Included in MLOps Platform Development
We provide a full cycle of work, documentation, and team training.
| Stage |
Duration |
Result |
| Audit of current infrastructure and data pipeline |
1–2 weeks |
Roadmap with risks and priorities |
| Core deployment: MLflow, orchestrator, serving |
4–6 weeks |
Working training and deployment pipeline |
| Feature Store and CI/CD for ML |
2–3 months |
Feature Store, automatic retrain and deployment |
| Drift monitoring and alerting |
3–4 weeks |
Dashboards, alerts, incident playbook |
| Team training and documentation |
1–2 weeks |
Runbook, policies, training for data scientists |
Total time from audit to full MLOps platform: 3–5 months. Also possible phased launch: basic level (tracking + serving) in 4–6 weeks.
Cost is calculated individually based on data volume, number of models, and infrastructure requirements. Order an MLOps infrastructure audit — get a roadmap in 1–2 weeks. Contact us for a project assessment — we will send a preliminary estimate within 2 business days.
Note: warranty on architectural solutions — 12 months. We provide integration certificates with major cloud providers (AWS, GCP, Azure). During our work, we have not lost a single client after the first implementation — the experience of 50+ successful MLOps projects speaks for itself. Get a consultation on building an MLOps platform today.