Traders often complain: the model worked great in a trend, but as soon as the market turned sideways, it started to drain the deposit. The reason is not the model but the lack of MLOps that adapts to market regime. Standard MLOps, where 500 ms latency is considered normal, is unacceptable for HFT — every microsecond of delay means lost profit. Our infrastructure is designed for the stringent requirements of the financial industry: p99 latency under 5 ms, zero-downtime model switching, and a full audit trail.
According to a J.P. Morgan study, 70% of HFT firms use FPGAs for inference, achieving up to 10x acceleration over CPU. However, FPGAs are not always necessary: for intraday strategies, ONNX Runtime with optimizations is sufficient. Let's break down how to build MLOps that withstands real trading loads.
Why MLOps for trading is a separate discipline
Standard MLOps is designed for services with latencies of hundreds of milliseconds and the possibility of manual model switching. In trading, the cost of error is lost profit or regulatory sanctions. Here are four key differences:
Latency requirements: HFT models must deliver predictions in <1 ms. For intraday strategies, <100 ms. Standard REST API inference services often fall short. Compare approaches:
| Strategy Type | Target Latency | Inference Tool | Speedup over PyTorch |
|---|---|---|---|
| HFT | <1 ms | FPGA / ONNX Runtime + TensorRT | up to 5x |
| Intraday | <100 ms | Triton Inference Server / ONNX Runtime | up to 3x |
| Medium-term | <1 s | TorchServe / BentoML | 1.5x |
Zero-downtime switching: replacing a model during trading hours is risky. We need hot-swap mechanisms without interrupting trading — for example, via shadow deployment with an agreement rate check > 90%. This is 2-3 times more reliable than standard blue-green deployment with interruption.
Reproducibility: during an audit, you must be able to exactly reproduce a model's prediction at a specific point in time (which model version was active, what data was used). We guarantee this through a full audit trail — every prediction is logged with model_version, data_version, and code hash.
Market regime awareness: retraining must account for the current market regime (trend, mean-reversion, high-volatility). A model good for trending markets is dangerous in a sideways market. Our pipeline automatically detects regime changes using volatility and correlations.
How to ensure zero-downtime model deployment
Hot-swap without stopping trading is the key requirement. We use a multi-stage approach:
- Shadow deployment: A new model runs in parallel with the main one. For 2 hours, we compare predictions — if the agreement rate > 90%, the model can be promoted.
- Canary deployment: Switch 10% of trading volume to the new model, monitor P&L attribution. If deviations occur, rollback within 1 second.
- Full rollout: During non-trading hours (02:00-09:00), switch all volume. All actions are logged for audit.
Infrastructure architecture
┌─────────────────────────────────────────────────────────┐
│ Data Infrastructure │
│ [Market Data Vendor] → [Kafka] → [ClickHouse/TimescaleDB]│
│ [Alternative Data] → [Feature Store] ← [Feature Pipeline]│
└─────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────┐
│ Training Infrastructure │
│ [Airflow/Prefect] → [GPU Training Cluster] │
│ [MLflow] ← [Experiment Tracking] → [Model Registry] │
│ [DVC] → [Data Versioning] │
└─────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────┐
│ Inference Infrastructure │
│ [Model Loader] → [Low-Latency Inference Server] │
│ [Shadow Model] → [A/B Framework] → [Active Model] │
│ [Risk Management Layer] → [Execution Engine] │
└─────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────┐
│ Monitoring Stack │
│ [Prediction Logger] → [ClickHouse] → [Grafana] │
│ [Drift Detector] → [Alert Manager] → [PagerDuty] │
│ [P&L Attribution] → [Model Performance Dashboard] │
└─────────────────────────────────────────────────────────┘
How to ensure low-latency inference
import onnxruntime as ort
import numpy as np
import threading
class LowLatencyModelServer:
"""Inference with target latency <5ms"""
def __init__(self, model_path: str):
# ONNX Runtime with optimizations
opts = ort.SessionOptions()
opts.intra_op_num_threads = 4
opts.inter_op_num_threads = 1
opts.execution_mode = ort.ExecutionMode.ORT_SEQUENTIAL
opts.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
self.session = ort.InferenceSession(
model_path,
sess_options=opts,
providers=['CUDAExecutionProvider', 'CPUExecutionProvider']
)
self._lock = threading.RLock()
# Warm up
dummy_input = np.zeros((1, 50), dtype=np.float32)
for _ in range(10):
self.predict(dummy_input)
def predict(self, features: np.ndarray) -> float:
with self._lock:
result = self.session.run(
None,
{'input': features.astype(np.float32)}
)
return float(result[0][0])
This ONNX Runtime code gives up to 3x acceleration over PyTorch JIT on the same GPU. For HFT, we additionally use TensorRT and pinned memory.
Technical detail: model quantization
Using INT8 quantization via ONNX Runtime further reduces latency by 20-30% without loss of accuracy for most trading models. We apply QAT (Quantization Aware Training) during fine-tuning to minimize error.Retraining pipeline with market regime detection
class TradingModelRetrainingPipeline:
def __init__(self, model_registry, risk_manager):
self.registry = model_registry
self.risk = risk_manager
def should_retrain(self, performance_metrics: dict) -> tuple[bool, str]:
# Feature drift
if performance_metrics['feature_psi'] > 0.2:
return True, "Feature drift detected"
# Metric degradation
if performance_metrics['sharpe_ratio_7d'] < 0.5:
return True, "Sharpe degradation"
# Market regime change
if self._detect_regime_change():
return True, "Market regime change"
return False, None
def safe_model_swap(self, new_model_path: str):
"""Hot-swap model without interrupting trading"""
# 1. Start shadow deployment for 2 hours
self._start_shadow_deployment(new_model_path)
# 2. Check agreement rate
if self._shadow_agreement_rate() < 0.90:
raise ValueError("Shadow model agreement rate too low for safe swap")
# 3. Switch during non-trading hours (02:00-09:00)
if not self._is_safe_swap_window():
self._schedule_swap_for_night()
return
# 4. Swap
with self.risk.trading_pause(timeout_seconds=5):
self.active_model = load_model(new_model_path)
Which metrics to monitor to catch degradation early
Monitoring should cover not only model performance but also infrastructure metrics. We distinguish three levels:
| Level | Metrics | Alert Threshold |
|---|---|---|
| Infrastructure | p99 latency, throughput, GPU utilization | latency >5 ms, GPU util <50% |
| Model | Sharpe ratio (7d), drawdown, LIFT | Sharpe <0.5, drawdown >15% |
| Data | PSI, feature importance, correlation with regime | PSI >0.2 |
Audit trail and reproducibility
def log_prediction_for_audit(features, prediction, model_version, timestamp):
audit_store.insert({
'timestamp': timestamp,
'model_version': model_version,
'model_git_hash': get_model_code_hash(model_version),
'data_version': get_feature_data_version(timestamp),
'input_features': features.tolist(),
'prediction': float(prediction),
'prediction_id': str(uuid.uuid4())
})
Every prediction is saved with model version, data version, and code hash — sufficient for accurate reproduction in case of a regulatory request or investigation.
What's included
Our delivery includes:
- Documentation: detailed architecture, user manual, and audit trail guide.
- Access: to the CI/CD pipeline, model registry, and monitoring dashboards.
- Training: 2–3 sessions for your team, covering operations and troubleshooting.
- Support: 6 months of stability guarantee and priority bug fixes.
Company metrics
Our team brings over 5 years of MLOps experience, having completed 20+ trading projects. With 8 years on the market, we have built robust infrastructure for firms ranging from startups to large financial institutions.
Implementation timeline
Basic MLOps infrastructure (tracking, registry, deployment): 4-6 weeks. Full system with drift monitoring, auto-retraining, and audit trail: 3-4 months. The investment pays off at the first serious model degradation incident. Contact us — let's discuss details and timelines for your project.







