Imagine: overnight, checkout conversion drops by 40%, and technical monitoring is silent — CPU normal, latency normal. Only in the morning does the CEO see the numbers in Google Analytics and lose revenue. Custom alerting rules on business metrics solve this: you learn about a drop in orders or a rise in payment errors within 5 minutes. We have set up such alerts for dozens of e-commerce, SaaS, and content projects. The approach is simple: instrument code with Prometheus metrics, set thresholds considering seasonality, and configure routing to Slack or PagerDuty. The result — reaction time to business-critical problems is reduced from hours to minutes. According to research, an hour of checkout downtime costs a large e-commerce an average of 500,000 rubles. In this article, we'll cover which metrics to monitor, how to implement alerts on Prometheus and CloudWatch, and how to avoid drowning in false positives.
Which business metrics to monitor?
Metrics depend on the product type. We identify three main categories.
E-commerce
- Number of completed orders per hour (sharp drop) — critical for revenue
- Conversion from cart to payment — if it falls below baseline by X%, it's a signal to check the payment gateway
- Revenue per rolling hour — helps quickly detect anomalies
- Number of payment errors — important for the technical team
SaaS
- New user registrations (zero in the last N hours) — signal of a sign-up process failure
- Active users online — unexpected drop may indicate an API issue
- API requests from key clients — anomalous growth or drop
Content projects
- Page views — sharp decrease = SEO or CDN issue
- Bounce rate — sharp increase may indicate slow loading
- Forms submitted — zero in N hours indicates a form bug
| Project type | Metric | Threshold | Action |
|---|---|---|---|
| E-commerce | Completed orders | 0 for 30 min during peak | Check payment gateway, CDN |
| E-commerce | Checkout conversion | < 30% (was 60%) | Check funnel, payment form bugs |
| SaaS | Registrations | 0 for 60 min | Check authentication process, database |
| Content | Page views | -50% in an hour | Check CDN, SEO traffic, server load |
How we implement alerts on business metrics
We use two main approaches: Prometheus (for self-hosted) and CloudWatch (for AWS). The choice depends on your infrastructure. From experience, Prometheus offers more flexibility in customization, while CloudWatch integrates better with Lambda and API Gateway.
Implementation via Prometheus
Instrument code with Prometheus client:
from prometheus_client import Counter, Histogram
orders_completed = Counter(
'orders_completed_total',
'Total completed orders',
['payment_method', 'product_category']
)
order_value = Histogram(
'order_value_rub',
'Order value in rubles',
buckets=[100, 500, 1000, 2500, 5000, 10000, 25000, 50000]
)
payment_errors = Counter(
'payment_errors_total',
'Payment processing errors',
['error_code', 'payment_provider']
)
Alerting rules for business metrics:
groups:
- name: business_alerts
rules:
- alert: NoOrdersReceived
expr: |
(
rate(orders_completed_total[30m]) == 0
and
hour() >= 9 and hour() <= 22
)
for: 5m
labels:
severity: critical
team: business
annotations:
summary: "No orders completed in last 30 minutes during business hours"
runbook_url: "https://wiki.company.com/runbooks/no-orders"
- alert: ConversionDropped
expr: |
rate(orders_completed_total[1h])
/
rate(cart_checkout_started_total[1h])
< 0.3
for: 15m
labels:
severity: warning
annotations:
summary: "Checkout conversion dropped to {{ $value | humanizePercentage }}"
- alert: PaymentErrorRateHigh
expr: |
rate(payment_errors_total[5m]) > 0.5
for: 3m
labels:
severity: critical
annotations:
summary: "{{ $value }} payment errors/sec — potential payment gateway issue"
Implementation via CloudWatch (AWS)
For AWS projects, we use CloudWatch. The code is instrumented via SDK (boto3) — sending OrdersCompleted and OrderRevenue metrics with dimensions PaymentMethod and Environment. Alerting rules are defined in Terraform:
resource "aws_cloudwatch_metric_alarm" "no_orders" {
alarm_name = "no-orders-30min"
comparison_operator = "LessThanThreshold"
evaluation_periods = 1
metric_name = "OrdersCompleted"
namespace = "MyApp/Business"
period = 1800
statistic = "Sum"
threshold = 1
treat_missing_data = "breaching"
dimensions = {
Environment = "production"
}
alarm_description = "No orders in 30 minutes"
alarm_actions = [aws_sns_topic.critical_alerts.arn]
}
How to set up alerts for seasonal metrics?
Fixed thresholds work poorly for metrics with seasonality. On Friday evening, orders are 3 times higher than on Monday morning. CloudWatch Anomaly Detection builds a forecast based on historical data and triggers only when deviation exceeds the band. In our projects, this reduces false positives by 70%.
resource "aws_cloudwatch_metric_alarm" "orders_anomaly" {
alarm_name = "orders-anomaly"
comparison_operator = "LessThanLowerOrGreaterThanUpperThreshold"
evaluation_periods = 2
threshold_metric_id = "e1"
metric_query {
id = "e1"
expression = "ANOMALY_DETECTION_BAND(m1, 2)"
return_data = true
}
metric_query {
id = "m1"
return_data = false
metric {
metric_name = "OrdersCompleted"
namespace = "MyApp/Business"
period = 300
stat = "Sum"
}
}
}
Routing business alerts
Business alerts should not wake developers at night. We configure routing via Alertmanager: alerts with label team=business and severity=critical go to Slack, while PaymentErrorRateHigh goes to PagerDuty, as it is a technical issue.
routes:
- match:
team: business
severity: critical
receiver: business-slack
- match:
team: business
alertname: PaymentErrorRateHigh
receiver: pagerduty-oncall
Additional settings: notification groups and time windows
To prevent noise, we group identical alerts into one ticket with a 5-minute interval, and set time windows for notifications: during non-working hours, alerts with severity=warning are simply logged.
Why do business alerts produce false positives? How to fix it
False positives are the main headache of monitoring. They make you ignore warnings. A few rules:
- Use anomaly detection instead of fixed thresholds for metrics with seasonality.
- Add "for" time windows (e.g., 5 minutes) — short-term spikes won't trigger an alert.
- For non-critical alerts (conversion dropped by 10%), use low priority to avoid overloading the responsible persons.
- Regularly review alerts based on historical data and disable useless ones.
What is included in the work
- Instrumenting code with metrics (2–4 days)
- Setting up alerting rules and routing (1–2 days)
- Anomaly Detection for seasonal metrics (1 day)
- Business dashboard in Grafana (1–2 days)
- Documentation on metrics and alerts
- Training the team on business alerts
- Post-deploy support: 30 days of accompaniment
Process of work
- Analytics — define business metrics and their baseline, choose the tool (Prometheus, CloudWatch, or other)
- Instrumentation — add metrics to the code, configure export
- Alerting rules — write rules considering seasonality and business priorities
- Routing — configure notifications for different teams and times of day
- Dashboard — create a business dashboard in Grafana with read-only access for the CEO
- Test and deploy — test alerts on historical data, deploy to production
Implementation timelines
| Stage | Timeline |
|---|---|
| Project evaluation | Free in 2 days |
| Basic implementation (5–7 metrics) | from 5 business days |
| Comprehensive solution with anomaly detection | from 8 business days |
Contact us for a consultation. Order turnkey custom alerting implementation. We have experience setting up monitoring for e-commerce, SaaS, and content projects. We guarantee stable operation and adequate notifications without false positives.







