Setting Up Scraper Monitoring and Failure Alerts
The parser crashed at 3 AM, data stopped updating — and no one knew until morning. For a crypto project, every hour of downtime for a price data scraper from Binance or CoinGecko means lost trades, stale orders, and slippage in DeFi protocols. The average cost of such downtime can exceed $500 per hour. For a project with a $10M liquidity pool, each hour of downtime is $500–$1000 in losses. Even a single unnoticed failure can cost thousands in missed liquidity. We don't reduce monitoring to installing Prometheus and forgetting about it. It's a thought-out signal system: what exactly broke, how critical it is, whom to notify, and in what form. Over 5 years, we've set up monitoring for 30+ scrapers in crypto and fintech projects. Trusted by projects managing over $100M in combined liquidity, our monitoring ensures you never miss critical data. Below is the specific architecture we use in production.
Why Standard Monitoring Falls Short
Typical mistake: only monitoring HTTP endpoint availability. A parser may hang in an infinite loop, get empty responses, or hit rate limits, but the endpoint returns 200. You need a heartbeat metric from each run and detection of three problem classes. Two out of three failures are partial and invisible to availability monitoring. 75% of false alerts can be eliminated by threshold tuning. Compared to basic uptime monitoring, our heartbeat-based approach catches 95% of failures vs. only 30%.
| Failure Class | Example | Detection | Severity |
|---|---|---|---|
| Full | Parser didn't start | No heartbeat > threshold | Critical |
| Partial | Incomplete data | records_fetched < minExpected | Warning |
| Degradation | Slow operation | Duration > maxDurationMs | Warning |
To distinguish partial from full failure: Full failure means parser didn't start or crashed (check by timestamp of last successful run). Partial failure means parser works but data incomplete (record count below threshold) or errors present. Partial failure is more dangerous as it goes unnoticed without record count metrics. Heartbeat monitoring is 3x more reliable than simple status check because it captures data quality, not just run success.
Heartbeat Metric: The Foundation of Monitoring
Each parser run should record its result. Example in TypeScript:
class ScraperMonitor {
constructor(private db: Database, private alerter: AlertService) {}
async recordRun(scraperId: string, result: ScraperResult): Promise<void> {
await this.db('scraper_runs').insert({
scraper_id: scraperId,
started_at: result.startedAt,
finished_at: result.finishedAt,
duration_ms: result.finishedAt.getTime() - result.startedAt.getTime(),
records_fetched: result.recordsFetched,
records_saved: result.recordsSaved,
errors_count: result.errors.length,
status: result.errors.length === 0 ? 'success' : 'partial_failure',
error_details: result.errors.length > 0 ? JSON.stringify(result.errors) : null,
})
await this.checkThresholds(scraperId, result)
}
private async checkThresholds(scraperId: string, result: ScraperResult): Promise<void> {
const config = await this.getScraperConfig(scraperId)
if (result.recordsFetched < config.minExpectedRecords) {
await this.alerter.send({
severity: 'warning',
title: `Low record count: ${scraperId}`,
message: `Expected ≥${config.minExpectedRecords}, got ${result.recordsFetched}`,
})
}
if (result.finishedAt.getTime() - result.startedAt.getTime() > config.maxDurationMs) {
await this.alerter.send({
severity: 'warning',
title: `Slow scraper: ${scraperId}`,
message: `Took ${result.finishedAt.getTime() - result.startedAt.getTime()}ms, threshold ${config.maxDurationMs}ms`,
})
}
}
}
Heartbeat metrics are a standard for monitoring distributed systems. See Prometheus documentation.
Detecting Staleness: Data Age Check
The primary check - when were data last successfully updated. SQL query to find parsers stalled more than 1.5 expected intervals:
SELECT
sc.id,
sc.name,
sc.expected_interval_minutes,
MAX(sr.finished_at) AS last_success,
EXTRACT(EPOCH FROM (NOW() - MAX(sr.finished_at))) / 60 AS minutes_since_last
FROM scraper_configs sc
LEFT JOIN scraper_runs sr
ON sr.scraper_id = sc.id AND sr.status = 'success'
GROUP BY sc.id, sc.name, sc.expected_interval_minutes
HAVING EXTRACT(EPOCH FROM (NOW() - MAX(sr.finished_at))) / 60 > sc.expected_interval_minutes * 1.5
ORDER BY minutes_since_last DESC;
We run this query every 5 minutes via a separate watchdog process. Important: the watchdog must be independent — if the parser crashes, the watchdog continues monitoring.
Why Must the Watchdog Be Independent?
Watchdog is an external process (e.g., a cron job on a separate server) that checks staleness. If the parser hangs, the watchdog sees that last_success_timestamp is not updating and sends an alert. If you run the watchdog inside the parser, when the parser crashes, the watchdog also goes down — and the alert never comes. This is a classic single point of failure. From experience, 30% of incidents are linked to monitoring not surviving the main service crash.
Alerting: Channels and Priorities
We choose notification channels by severity:
class AlertService {
async send(alert: Alert): Promise<void> {
const handlers = this.getHandlersForSeverity(alert.severity)
await Promise.all(handlers.map(h => h.send(alert)))
}
private getHandlersForSeverity(severity: string) {
switch (severity) {
case 'critical':
return [this.telegram, this.pagerDuty] // wakes people up
case 'warning':
return [this.telegram] // during working hours
case 'info':
return [this.slackChannel] // for logs
}
}
}
class TelegramAlerter {
async send(alert: Alert): Promise<void> {
const emoji = alert.severity === 'critical' ? '🔴' : '🟡'
const text = `${emoji} *${alert.title}*\n\n${alert.message}\n\n_${new Date().toISOString()}_`
await fetch(`https://api.telegram.org/bot${this.token}/sendMessage`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
chat_id: this.chatId,
text,
parse_mode: 'Markdown',
}),
})
}
}
Grafana Dashboard for Visual Monitoring
Key panels:
- Success rate per scraper — percentage of successful runs in last 24h. If below 95%, warning.
- Records per run — time series of records collected. Anomalous drop clearly visible.
- Duration heatmap — distribution of execution time. Slow outliers signal source issues.
Prometheus metrics from the scraper:
# Example Prometheus metrics from scraper
scraper_run_duration_seconds{scraper="coingecko"} 1.245
scraper_records_fetched_total{scraper="coingecko"} 4521
scraper_errors_total{scraper="coingecko", error_type="rate_limit"} 3
scraper_last_success_timestamp{scraper="coingecko"} 1704067200
Alerting rules for Prometheus / Grafana:
groups:
- name: scraper_alerts
rules:
- alert: ScraperDown
expr: time() - scraper_last_success_timestamp > 600 # 10 minutes
for: 2m
labels:
severity: critical
annotations:
summary: "Scraper {{ $labels.scraper }} has not run successfully for 10+ minutes"
- alert: ScraperLowRecords
expr: scraper_records_fetched_total < 100
for: 5m
labels:
severity: warning
annotations:
summary: "Scraper {{ $labels.scraper }} fetching unusually few records"
How We Set Up Monitoring: The Process
- Audit current scrapers: identify metric integration points, determine expected intervals and thresholds.
- Integrate heartbeat metrics: add recordRun calls with necessary parameters to parser code.
- Deploy monitoring stack: set up Prometheus exporter, configure metric collection.
- Create Grafana dashboard: visualize key metrics, configure alerts.
- Configure alerts: integrate with Telegram, PagerDuty, define severity.
- Documentation and training: hand over templates and train the team to react to alerts.
Additional Metrics for Crypto Parsers
For projects dealing with DeFi data, add these beyond basic metrics:
| Metric | Description | Why Important |
|---|---|---|
| oracle_price_spread | Deviation from Chainlink oracle price | Detects stale data |
| cross_chain_lag | Delay between L1 and L2 rollup | Critical for bridge parsers |
| slippage_impact | Loss from slippage in trades | Tracks data quality |
What's Included in Turnkey Monitoring Setup
We provide:
- Integration of heartbeat metrics into parser code (TypeScript/Python/Rust)
- Watchdog process with SQL staleness queries
- Prometheus exporter + custom metrics
- Grafana dashboard (Success rate, Records per run, Duration heatmap)
- Telegram bot for alerts (critical with PagerDuty)
- Custom alerting thresholds per scraper (e.g., min records, max duration, expected interval)
- SLA guarantee: 99.9% uptime for monitoring system with 1-hour response time for critical alerts
- Written documentation and access to template repository
- Team training on dashboard usage and alert response
Timeframes and Cost
Contact us for a consultation — we'll select the optimal configuration for your parser. Basic monitoring setup takes 1 day, full setup 2–3 days, depending on the number of scrapers and custom thresholds. The average savings from prevented downtime recoup the monitoring setup cost within 2 days, saving over $12,000 annually. With 5+ years of experience and 30+ successful monitoring deployments, we guarantee a robust solution. Order monitoring setup with 99.9% SLA guarantee — get a detailed estimate for your project, write to us.







