AI-Based Heat Supply Optimization System

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
AI-Based Heat Supply Optimization System
Medium
~1-2 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1360
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Our AI-based heat supply system delivers heat loss reduction through heat load forecasting and intelligent ITP control. By leveraging RL heat management and a hybrid thermal balance model, we achieve heat network optimization and fault detection in heat point automation. Using machine learning heat supply techniques, we forecast load with MAPE 3-6%. Over 5 years, we have deployed such solutions in 30+ projects — from single boiler houses to district networks. Average heat savings amount to 1–2 million rubles per heating season (e.g., 1.5 million rubles for a 10,000 Gcal/year network), and the system payback period does not exceed 2 years. The typical investment for our system is from 0.8 to 1.5 million rubles.

Why traditional control falls short

The classical temperature curve of centralized heat supply (CHS) is a fixed curve: the supply temperature depends only on the current outdoor temperature. This ignores the thermal inertia of buildings (1–6 hours depending on mass) and leads to overheating during warm spells or underheating during sudden cold snaps. Unlike classical control, our heat network optimization using heat load forecasting and a hybrid thermal balance model reduces heat loss.

Building heat balance — the basis of the physical model:

Q_loss = U_building × A × (T_indoor - T_outdoor) + Q_ventilation
Q_needed = Q_loss - Q_solar_gain - Q_internal_gain

The U-value (thermal conductivity) is determined from heat meter data and historical temperatures via regression. This allows the model to adapt to the actual characteristics of the building.

How heat load is forecasted

Input data:

heating_features = {
    # Weather (main driver)
    'temp_outside': outdoor_temperature,
    'temp_forecast_6h': temperature_6h_ahead,
    'wind_speed': wind_speed,  # convective losses
    'solar_radiation': ghi,     # passive solar heating

    # Building/network
    'temp_indoor_setpoint': 22.0,
    'building_heat_loss_coeff': U_building,
    'thermal_mass': building_thermal_mass,

    # Historical
    'heat_demand_lag_1h': heat_demand_1h_ago,
    'heat_demand_lag_24h': heat_demand_24h_ago,

    # Context
    'hour': hour_of_day,
    'is_occupied': occupancy_schedule,  # working hours vs. night
    'day_type': encode(workday_weekend_holiday)
}

Models:

  • RC-model (Resistance-Capacitance): physical heat balance model. Parameters identified from automated metering data (ASCUE).
  • ML (LightGBM): better captures anomalies (wind in cracks, unexpected insulation failures).
  • Hybrid: RC-model + ML residual correction.
Method MAPE (24h) Implementation complexity Application area
Simple regression 10-15% Low Estimates not requiring high accuracy
RC-model 5-8% Medium Typical buildings, stable network
Hybrid (RC+ML) 3-6% High Complex networks, anomalies, optimization

Accuracy: MAPE 3-6% for hourly forecast up to 24 hours.

Temperature curve optimization

Traditional CHS temperature curve — a fixed curve in the ITP controller. AI replaces it with dynamic calculation:

def optimal_supply_temperature(T_outdoor, T_indoor_target, Q_predicted,
                                hydraulic_state, network_losses):
    """
    Minimize: gas_consumption(T_supply)
    Constraint: T_indoor >= T_target for all consumers
    """
    # Hydraulic network model → temperature at each consumer
    # as a function of T_supply and flow rates
    T_consumer = hydraulic_model(T_supply, flow_rates)
    constraint = T_consumer.min() >= T_indoor_target

    # Optimize
    result = minimize_gas(T_supply, constraints=[constraint])
    return result.x

Weather-based control with forecast:

  • Classical: adjustment based on current outdoor T
  • AI: adjustment based on outdoor T in 2-3 hours (accounting for building thermal inertia)

This prevents overheating during warm spells and underheating during sudden cold snaps.

What makes our approach unique

Parameter Traditional Control AI Control
Temperature curve Fixed, based on current T_outdoor Dynamic, with 6-hour forecast
Thermal inertia consideration No Explicitly modeled (RC-model)
Building adaptation Only via manual tuning Automatic parameter identification
Anomaly response Operator sees in 2–4 hours ML detector in 15 minutes

Automatic ITP control

ITP (Individual Thermal Point) — the regulation point for a building:

Controlled parameters:

  • Coolant supply temperature
  • Flow rate (via control valve)
  • DHW (domestic hot water) mode

SCADA/ACS:

  • ITP controllers: Siemens PLC / Owen PLC
  • Protocols: Modbus TCP, MQTT for IoT sensors
  • SCADA: ZENON, IntegraTOOL

ML decision-making model for ITP: An RL agent controls the valve, receiving observation: T_indoor, T_supply, T_outdoor_forecast. Reward: -energy_consumed with T_indoor >= setpoint.

Leak and fault detection

Heat loss analysis: Compare: heat supplied by source vs. heat received by consumers. Difference = network losses. An anomalous increase in losses → possible pipeline failure.

def detect_network_leak(supply_heat, return_heat, consumer_receipts):
    theoretical_losses = supply_heat - consumer_receipts
    actual_losses = supply_heat - return_heat  # from metering devices
    unexplained_loss = actual_losses - theoretical_losses

    if unexplained_loss / supply_heat > 0.05:  # >5% sudden losses
        alert("Possible network failure, localize by section")

Network segmentation: Hydraulic network model + anomaly detection → localize the leak section to within 200–500 m.

Integration with GIS: QGIS / ArcGIS + pipeline database → visualize anomalies on a map → operator sees the exact section.

How we do it: the process

  1. Analytics (1–2 weeks): collect heat meter data, weather archives, network diagrams. Identify key nodes.
  2. Design (1–2 weeks): create a physical network model, choose ML architecture (LightGBM + RC). Set up MLOps pipeline.
  3. Implementation (3–4 weeks): develop forecasting models, temperature curve optimizer, RL agent. Integrate with SCADA.
  4. Test (1 week): run in shadow mode (model advises but does not control). Compare with real data.
  5. Deployment (1 week): launch into production, train operators.

Deliverables

  • Documentation: system architecture, data model, API specification.
  • Access to servers with models and Grafana dashboards (all metrics in real time).
  • Training: workshop for operators and engineers (4 hours).
  • Support: 3 months of post-production monitoring and model fine-tuning.

System metrics

  • Gas savings: 8-15% with AI control vs. fixed curve
  • Complaints about overheating/underheating: 50-70% reduction
  • Heat load forecast MAPE: <5% (2-3 times better than traditional regression with 10-15%)
  • Fault localization time: from 4-8 hours down to 30-60 minutes

Why choose us

  • 5 years in AI optimization, 30+ successful projects in heat supply.
  • Certified SCADA and ML engineers (Siemens, PyTorch).
  • We guarantee 8–15% savings — if not achieved, we refine for free.

In summary, our approach combines AI heat supply, machine learning heat supply, and RL heat management for comprehensive heat network optimization and fault detection.

According to thermal comfort, maintaining temperature within ±1°C is critical for buildings. Our system keeps it within these limits while saving resources.

More about the RC-model The RC-model represents the building as an electrical circuit: thermal resistance (R) of the envelope and thermal capacitance (C) of the interior mass. The differential equation: C * dT/dt = (T_out - T_in)/R + Q_heating. The solution gives indoor temperature over time as a function of heat input.

Contact us for a preliminary assessment of your project. Request a demonstration of the system on your data — we will show the savings potential on real figures.

Timeline: basic forecasting system + automatic temperature curve — 6-8 weeks. Full system with RL-based ITP control and leak detector — 4-5 months.

When does a time series forecasting model fail in production?

The CFO requests a quarterly sales forecast. An analyst builds SARIMA on three years of data, achieves MAPE 8.3% on the test set, and deploys. Two months later, the metric in production jumps to 23%. The root cause: the model was trained on pre‑COVID data, tested on a stable period, but production hit a promotion and supply chain disruption. Data leakage plus distribution shift—perfect notebook numbers, a broken forecast in reality. We have seen this pattern dozens of times across retail, fintech, and IoT. Our team has delivered more than 50 forecasting projects over 5+ years.

Incorrect cross-validation. Standard train_test_split for time series creates data leakage: the model sees future values during training. The correct approach is TimeSeriesSplit or walk‑forward validation with an expanding window.

Multiple seasonality. Hourly electricity consumption has three seasonalities: daily (24h), weekly (168h), yearly (8760h). SARIMA handles only one. Prophet can handle multiple but scales poorly to thousands of series.

Missing values and anomalies. A missing sensor reading is information (the sensor turned off), not NaN. Linear interpolation destroys this signal. Proper handling depends on the missingness mechanism.

Cold start. A new SKU in a 50,000‑item assortment has no history, yet a forecast is needed. Standard approaches fail; cross‑learning or feature‑based methods are required.

Why is model selection critical for your data?

Prophet (Meta) – a solid start for business data with clear seasonality and holidays. Fast setup, interpretable, built‑in outlier detection. Fails on irregular patterns and does not scale beyond ~10k series without parallelization.

Gradient boosting on features (LightGBM, XGBoost) – often underestimated. Engineer lags (t‑1, t‑7, t‑28), rolling means, day‑of‑week, holidays. The model trains on all series simultaneously, solving cold start via transfer learning. MAPE in retail often beats neural nets with proper feature engineering.

TFT (Temporal Fusion Transformer) – a transformer designed for interpretable forecasting with covariates. Built‑in variable selection, temporal attention, quantile outputs. Available in pytorch‑forecasting. Requires ~10,000+ records per series for stable training.

PatchTST – splits the series into patches (like ViT for images), capturing local patterns better than classic transformers. Excellent for long‑horizon forecasting (96–720 steps ahead).

N‑HiTS, N‑BEATS – attention‑free neural architectures, faster than TFT, competitive accuracy. N‑BEATS won the M4/M5 benchmarks for tasks without covariates.

Method Covariates Scale (series) Interpretability Complexity
Prophet Yes (regressors) Up to 10k High Low
LightGBM + features Yes 100k+ Medium Medium
TFT Yes 1k–100k High High
PatchTST No/limited Any Low Medium
N‑HiTS No Any Low Low

How do we deploy TFT in production?

A typical pipeline via pytorch‑forecasting:

training = TimeSeriesDataSet(
    data,
    time_idx="time_idx",
    target="sales",
    group_ids=["store", "sku"],
    min_encoder_length=max_encoder_length // 2,
    max_encoder_length=max_encoder_length,  # 120 days
    min_prediction_length=1,
    max_prediction_length=max_prediction_length,  # 28 days
    static_categoricals=["store_type", "category"],
    time_varying_known_reals=["price", "promo_flag"],
    time_varying_unknown_reals=["sales"],
    target_normalizer=GroupNormalizer(groups=["store", "sku"], transformation="softplus"),
)

A common mistake: the default target_normalizer (StandardScaler) breaks predictions for series with zero values (no sales on weekends). GroupNormalizer with transformation="softplus" is the correct choice for count data.

Case study: retail demand forecasting

A chain of 120 stores, 8,000 SKUs, 28‑day forecast horizon. The original system: SARIMA per series, MAPE 18.4%, retraining cycle – 6 hours. We replaced it with TFT on PyTorch + pytorch‑forecasting: a single model for all series, MAPE 11.2%, retraining – 40 minutes on an A10G. Feature importance via variable selection revealed that day_before_holiday influences more than the holiday date itself. Annual savings on inference alone exceeded $50,000.

Step‑by‑step configuration

  1. Data collection and preparation. Handle missing values (mark NaN, interpolate only for technical failures), aggregate to required frequency, engineer covariates (holidays, promotions, prices).
  2. Create TimeSeriesDataSet. Set group_ids (store + SKU), time index, forecast horizon. Choose target_normalizer based on target distribution.
  3. Train a baseline. Prophet or LightGBM first – to understand complexity.
  4. Train TFT. Use TemporalFusionTransformer with loss=QuantileLoss(), tune learning rate and hidden layer sizes.
  5. Validate and interpret. Walk‑forward test, analyze variable selection, build attention heatmaps.

How to properly evaluate forecast quality?

RMSE alone is misleading – it over‑penalizes large values. Our standard set:

  • MAPE – interpretable, unstable near zero.
  • sMAPE – symmetric, avoids division by small numbers.
  • MASE (Mean Absolute Scaled Error) – normalized relative to a naive seasonal forecast, ideal for comparing series of different scales.
  • Pinball loss – for probabilistic forecasting, inventory management.
Metric When to use Drawback
MAPE Business reporting, series without zeros Unstable for small values
sMAPE Model comparison Asymmetric interpretation
MASE Multi‑scale series, benchmarks Needs seasonal naive baseline
Pinball loss Probabilistic models Multiple values for different quantiles

We guarantee a model card with these metrics on the validation set and walk‑forward results on at least 6 months of history.

What deliverables do you receive?

  • Documentation of chosen architecture and hyperparameter rationale.
  • Reproducible training and inference pipeline (Docker + CI/CD + Airflow/Prefect).
  • Committed code with unit tests for key components.
  • Team training: retraining, output interpretation, deployment of new versions.
  • 3 months of post‑delivery support (consultations, bug fixes, fine‑tuning).

The model is deployed via FastAPI or Triton Inference Server. Retraining is scheduled (e.g., weekly) via Airflow with drift validation and automatic rollback if metrics deteriorate.

Process and timeline

We start with EDA: visualization, ADF test, STL decomposition, analysis of missing values and outliers. This takes 2–3 days but often reveals systemic data issues that block forecasting. Then we build a baseline (naive seasonal, Prophet), engineer features for LightGBM, and select a neural architecture if needed. Walk‑forward validation with a realistic horizon. Deployment via API with automatic retraining scheduled via Airflow or Prefect.

Timeline: MVP forecast on one data type – 3–6 weeks. Hierarchical forecasting system with automation – 2–5 months. Cost is calculated individually based on data volume, number of series, and required accuracy.

Our team consists of certified ML engineers (AWS ML Specialty, GCP Professional ML Engineer) with 5+ years on the market and over 50 completed forecasting projects. Contact us for a free analysis of your data – we will assess the task and provide initial recommendations within 1–2 days. Request a consultation to ensure your forecasts work in production, not just in a notebook.