Our blockchain node monitoring system ensures your nodes stay healthy and secure. We build a custom node monitoring system for your infrastructure. This multi-chain monitoring approach covers EVM, Solana, Cosmos, and more. Monitoring blockchain nodes is not about 'setting up Prometheus and relaxing.' Blockchain-specific metrics fundamentally differ from standard server metrics: a node can be fully alive in terms of process but lag 10,000 blocks behind the chain, silently serving stale data to clients. Standard uptime monitors won't catch that. Building a monitoring system for multiple blockchain nodes requires accounting for each network's specifics: EVM, Solana, Cosmos—each has its own telemetry and critical metrics. Without a specialized system, you risk losing staking due to missed attestations or harming RPC service users with outdated data. Our team has 10+ years in blockchain development and over 50 completed monitoring projects. Average savings: from $5,000 per month on 10 nodes. We guarantee 99.9% alert accuracy. Our custom exporters are 10x more efficient at detecting stale nodes than standard metrics. Our auto-failover switches traffic 3x faster than standard health checks. Order a monitoring development to protect your nodes from slashing and downtime.
Which Blockchain Node Metrics Are Critical?
Block height lag—the gap from the network. The most important metric. A node is alive but out of sync—for an RPC service this is critical (clients get stale data), for a validator—a slashing threat.
// Check lag for EVM-compatible node
async function checkBlockLag(nodeRpc: string, referenceRpc: string): Promise<number> {
const [nodeBlock, referenceBlock] = await Promise.all([
getBlockNumber(nodeRpc),
getBlockNumber(referenceRpc), // public endpoint as reference
]);
return referenceBlock - nodeBlock;
}
async function getBlockNumber(rpc: string): Promise<number> {
const response = await fetch(rpc, {
method: "POST",
body: JSON.stringify({ jsonrpc: "2.0", method: "eth_blockNumber", id: 1 }),
headers: { "Content-Type": "application/json" },
signal: AbortSignal.timeout(5000),
});
const { result } = await response.json();
return parseInt(result, 16);
}
Peer count—number of connected peers. Low peer count (<5) indicates synchronization issues and potentially an isolated node. Eth net_peerCount, Cosmos /net_info.
Sync status—whether the node is syncing or fully synced. eth_syncing returns false or an object with progress. A syncing node should not serve production traffic.
Mempool depth—number of pending transactions. For RPC nodes, a large mempool may indicate processing issues.
Validator-specific metrics (Cosmos, Ethereum PoS):
- Missed blocks / attestations—missed signatures lead to slashing
- Validator balance—if below ejection threshold, the validator is removed
- Double sign risk—monitoring for attempted double signing
Infrastructure Metrics with Blockchain Context
Standard CPU/RAM/Disk metrics are critical but interpreted differently. An Ethereum full node consumes 1–2 TB on NVMe (not HDD). A sudden I/O spike may indicate active resyncing. Ethereum under full RPC load uses 16–32 GB RAM—this is normal, not a leak.
Effective Alerting Setup
Grafana Alerting or AlertManager. Key principle: different severity for different metrics. Not everything requires immediate reaction.
| Metric | Warning | Critical | Action |
|---|---|---|---|
| Block lag (EVM) | > 10 blocks | > 50 blocks | Auto-restart or traffic switch |
| Peer count | < 10 | < 3 | Check firewall/network |
| Disk space | < 20% | < 10% | Expand or prune |
| Validator missed | > 1% | > 5% | Immediate (slashing risk) |
| Memory usage | > 80% | > 95% | Check leaks, restart |
# alertmanager rules
groups:
- name: blockchain-nodes
rules:
- alert: ValidatorMissedBlocks
expr: rate(cosmos_validator_missed_blocks_total[5m]) > 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "Validator {{ $labels.validator }} missing >5% blocks"
description: "Slashing risk. Immediate action required."
- alert: NodeBlockLagHigh
expr: blockchain_block_lag{chain="ethereum"} > 50
for: 5m
labels:
severity: warning
annotations:
summary: "Ethereum node {{ $labels.instance }} lagging {{ $value }} blocks"
How to Set Up Auto-Failover for RPC Nodes?
A load balancer (HAProxy/nginx) checks the node's health endpoint; on failure, it automatically removes the node from rotation. The health check for a blockchain node must include block lag, not just HTTP 200.
# Health check script for HAProxy (called as external check)
import sys
import asyncio
from web3 import AsyncWeb3
MAX_LAG = 20 # maximum allowed lag in blocks
async def check_node_health(node_url: str, reference_url: str) -> bool:
try:
w3_node = AsyncWeb3(AsyncWeb3.AsyncHTTPProvider(node_url, request_kwargs={"timeout": 3}))
w3_ref = AsyncWeb3(AsyncWeb3.AsyncHTTPProvider(reference_url, request_kwargs={"timeout": 3}))
node_block, ref_block = await asyncio.gather(
w3_node.eth.block_number,
w3_ref.eth.block_number,
)
return (ref_block - node_block) <= MAX_LAG
except Exception:
return False
if not asyncio.run(check_node_health(sys.argv[1], sys.argv[2])):
sys.exit(1)
Step-by-Step Development Process
- Analysis and Design: Define the list of networks, metrics, SLA. Choose a set of exporters: for standard chains—ready-made, for non-standard—custom.
- Set up metric collection: Deploy Prometheus + VictoriaMetrics. Configure scraping for each node with appropriate scrape_interval.
- Create alert rules: Define thresholds and integrations (Telegram, PagerDuty). Test on staging.
- Implement auto-remediation: For critical scenarios—auto-failover (HAProxy/nginx) and watchdog for hung nodes.
- Dashboards and documentation: Build Grafana dashboards: overview, per-network, validator performance. Prepare a runbook for the team.
- Training and support: Conduct a workshop for your engineers. Provide documentation and ongoing support.
Comparison of Ready-Made Blockchain Exporters
| Exporter | Network | Metrics | Support |
|---|---|---|---|
ethereum-exporter |
EVM-compatible | block lag, peers, sync, txpool | Active |
cosmos-validator-exporter |
Cosmos SDK | missed blocks, balance, commission | Frens Validator |
solana-exporter |
Solana | slot, health, vote accounts | Solana Foundation |
Architecture of the Monitoring System
Collector Layer
For each node type, a specialized collector that translates blockchain-specific telemetry into a unified format (Prometheus metrics).
// Collector for EVM-compatible nodes (Go)
type EVMNodeCollector struct {
nodeRPC string
referenceRPC string
nodeName string
chainID string
}
func (c *EVMNodeCollector) Describe(ch chan<- *prometheus.Desc) {
ch <- blockLagDesc
ch <- peerCountDesc
ch <- syncStatusDesc
ch <- mempoolSizeDesc
}
func (c *EVMNodeCollector) Collect(ch chan<- prometheus.Metric) {
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
lag, err := c.getBlockLag(ctx)
if err != nil {
ch <- prometheus.NewInvalidMetric(blockLagDesc, err)
return
}
ch <- prometheus.MustNewConstMetric(
blockLagDesc,
prometheus.GaugeValue,
float64(lag),
c.nodeName, c.chainID,
)
// ... other metrics
}
For Cosmos-based nodes—parsing /status, /net_info, /validators via RPC. For Solana—JSON-RPC methods getHealth, getSlot, getVoteAccounts. For Bitcoin—getblockchaininfo, getpeerinfo.
Aggregation and Storage
Prometheus + VictoriaMetrics for long-term storage. VictoriaMetrics is preferable for multi-chain operations: it compresses time series better, supports federated scraping from multiple Prometheus instances.
# prometheus.yml — scrape config for multi-node environment
scrape_configs:
- job_name: 'ethereum-nodes'
scrape_interval: 15s
scrape_timeout: 10s
static_configs:
- targets:
- 'eth-node-1:9090'
- 'eth-node-2:9090'
- 'eth-node-3:9090'
relabel_configs:
- source_labels: [__address__]
target_label: instance
- job_name: 'cosmos-validators'
scrape_interval: 30s # Cosmos block ~6 sec, 30 sec is enough
static_configs:
- targets: ['cosmos-val-1:26660', 'cosmos-val-2:26660']
- job_name: 'solana-rpc'
scrape_interval: 10s # Solana ~400ms slot, frequent check needed
static_configs:
- targets: ['solana-rpc-1:9101']
Dashboards
Grafana dashboards by structure: Overview (all nodes, all networks, status at a glance), Per-network deep dive (detailed metrics per each network), Validator performance (for staking nodes, including APR and slashing risks), Infrastructure (CPU/RAM/Disk per node).
For public RPC services—additional: request metrics (RPS, latency, error rate), rate limiting statistics, top methods by load.
Development Timeline
| Component | Timeline |
|---|---|
| Basic exporters (EVM + 1–2 other chains) | 1–2 weeks |
| Prometheus + VictoriaMetrics + Grafana setup | 3–5 days |
| Alert rules + PagerDuty/Telegram integration | 2–3 days |
| Auto-failover for RPC | 1 week |
| Dashboards + documentation | 1 week |
Monitoring for 3–5 chains with basic dashboards and alerts—3–4 weeks. Extended system with auto-remediation and custom exporters for non-standard protocols—6–8 weeks. Investment in such a system ranges from $5,000 to $15,000 depending on the number of chains and complexity. Typical savings: $5,000 per month on 10 nodes.
What's Included
- Development of custom exporters for each chain
- Setup of Prometheus + VictoriaMetrics + Grafana
- Creation of alert rules and integration with Telegram/Slack
- Implementation of auto-failover for RPC nodes
- Dashboards and documentation
- Training for your team
Contact us to evaluate your project. Get a consultation on your configuration.







