Blockchain Node Monitoring: Prevent Slashing & Outages

We design and develop full-cycle blockchain solutions: from smart contract architecture to launching DeFi protocols, NFT marketplaces and crypto exchanges. Security audits, tokenomics, integration with existing infrastructure.
Showing 1 of 1All 1305 services
Blockchain Node Monitoring: Prevent Slashing & Outages
Medium
~1-2 weeks
Frequently Asked Questions

Blockchain Development Services

Blockchain Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1357
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1249
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    954
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1187
  • image_logo-advance_0.webp
    B2B Advance company logo design
    645
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    926

Our blockchain node monitoring system ensures your nodes stay healthy and secure. We build a custom node monitoring system for your infrastructure. This multi-chain monitoring approach covers EVM, Solana, Cosmos, and more. Monitoring blockchain nodes is not about 'setting up Prometheus and relaxing.' Blockchain-specific metrics fundamentally differ from standard server metrics: a node can be fully alive in terms of process but lag 10,000 blocks behind the chain, silently serving stale data to clients. Standard uptime monitors won't catch that. Building a monitoring system for multiple blockchain nodes requires accounting for each network's specifics: EVM, Solana, Cosmos—each has its own telemetry and critical metrics. Without a specialized system, you risk losing staking due to missed attestations or harming RPC service users with outdated data. Our team has 10+ years in blockchain development and over 50 completed monitoring projects. Average savings: from $5,000 per month on 10 nodes. We guarantee 99.9% alert accuracy. Our custom exporters are 10x more efficient at detecting stale nodes than standard metrics. Our auto-failover switches traffic 3x faster than standard health checks. Order a monitoring development to protect your nodes from slashing and downtime.

Which Blockchain Node Metrics Are Critical?

Block height lag—the gap from the network. The most important metric. A node is alive but out of sync—for an RPC service this is critical (clients get stale data), for a validator—a slashing threat.

// Check lag for EVM-compatible node
async function checkBlockLag(nodeRpc: string, referenceRpc: string): Promise<number> {
    const [nodeBlock, referenceBlock] = await Promise.all([
        getBlockNumber(nodeRpc),
        getBlockNumber(referenceRpc),  // public endpoint as reference
    ]);
    return referenceBlock - nodeBlock;
}

async function getBlockNumber(rpc: string): Promise<number> {
    const response = await fetch(rpc, {
        method: "POST",
        body: JSON.stringify({ jsonrpc: "2.0", method: "eth_blockNumber", id: 1 }),
        headers: { "Content-Type": "application/json" },
        signal: AbortSignal.timeout(5000),
    });
    const { result } = await response.json();
    return parseInt(result, 16);
}

Peer count—number of connected peers. Low peer count (<5) indicates synchronization issues and potentially an isolated node. Eth net_peerCount, Cosmos /net_info.

Sync status—whether the node is syncing or fully synced. eth_syncing returns false or an object with progress. A syncing node should not serve production traffic.

Mempool depth—number of pending transactions. For RPC nodes, a large mempool may indicate processing issues.

Validator-specific metrics (Cosmos, Ethereum PoS):

  • Missed blocks / attestations—missed signatures lead to slashing
  • Validator balance—if below ejection threshold, the validator is removed
  • Double sign risk—monitoring for attempted double signing

Infrastructure Metrics with Blockchain Context

Standard CPU/RAM/Disk metrics are critical but interpreted differently. An Ethereum full node consumes 1–2 TB on NVMe (not HDD). A sudden I/O spike may indicate active resyncing. Ethereum under full RPC load uses 16–32 GB RAM—this is normal, not a leak.

Effective Alerting Setup

Grafana Alerting or AlertManager. Key principle: different severity for different metrics. Not everything requires immediate reaction.

Metric Warning Critical Action
Block lag (EVM) > 10 blocks > 50 blocks Auto-restart or traffic switch
Peer count < 10 < 3 Check firewall/network
Disk space < 20% < 10% Expand or prune
Validator missed > 1% > 5% Immediate (slashing risk)
Memory usage > 80% > 95% Check leaks, restart
# alertmanager rules
groups:
  - name: blockchain-nodes
    rules:
      - alert: ValidatorMissedBlocks
        expr: rate(cosmos_validator_missed_blocks_total[5m]) > 0.05
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Validator {{ $labels.validator }} missing >5% blocks"
          description: "Slashing risk. Immediate action required."

      - alert: NodeBlockLagHigh
        expr: blockchain_block_lag{chain="ethereum"} > 50
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Ethereum node {{ $labels.instance }} lagging {{ $value }} blocks"

How to Set Up Auto-Failover for RPC Nodes?

A load balancer (HAProxy/nginx) checks the node's health endpoint; on failure, it automatically removes the node from rotation. The health check for a blockchain node must include block lag, not just HTTP 200.

# Health check script for HAProxy (called as external check)
import sys
import asyncio
from web3 import AsyncWeb3

MAX_LAG = 20  # maximum allowed lag in blocks

async def check_node_health(node_url: str, reference_url: str) -> bool:
    try:
        w3_node = AsyncWeb3(AsyncWeb3.AsyncHTTPProvider(node_url, request_kwargs={"timeout": 3}))
        w3_ref = AsyncWeb3(AsyncWeb3.AsyncHTTPProvider(reference_url, request_kwargs={"timeout": 3}))

        node_block, ref_block = await asyncio.gather(
            w3_node.eth.block_number,
            w3_ref.eth.block_number,
        )
        return (ref_block - node_block) <= MAX_LAG
    except Exception:
        return False

if not asyncio.run(check_node_health(sys.argv[1], sys.argv[2])):
    sys.exit(1)

Step-by-Step Development Process

  1. Analysis and Design: Define the list of networks, metrics, SLA. Choose a set of exporters: for standard chains—ready-made, for non-standard—custom.
  2. Set up metric collection: Deploy Prometheus + VictoriaMetrics. Configure scraping for each node with appropriate scrape_interval.
  3. Create alert rules: Define thresholds and integrations (Telegram, PagerDuty). Test on staging.
  4. Implement auto-remediation: For critical scenarios—auto-failover (HAProxy/nginx) and watchdog for hung nodes.
  5. Dashboards and documentation: Build Grafana dashboards: overview, per-network, validator performance. Prepare a runbook for the team.
  6. Training and support: Conduct a workshop for your engineers. Provide documentation and ongoing support.
Comparison of Ready-Made Blockchain Exporters
Exporter Network Metrics Support
ethereum-exporter EVM-compatible block lag, peers, sync, txpool Active
cosmos-validator-exporter Cosmos SDK missed blocks, balance, commission Frens Validator
solana-exporter Solana slot, health, vote accounts Solana Foundation

Architecture of the Monitoring System

Collector Layer

For each node type, a specialized collector that translates blockchain-specific telemetry into a unified format (Prometheus metrics).

// Collector for EVM-compatible nodes (Go)
type EVMNodeCollector struct {
    nodeRPC      string
    referenceRPC string
    nodeName     string
    chainID      string
}

func (c *EVMNodeCollector) Describe(ch chan<- *prometheus.Desc) {
    ch <- blockLagDesc
    ch <- peerCountDesc
    ch <- syncStatusDesc
    ch <- mempoolSizeDesc
}

func (c *EVMNodeCollector) Collect(ch chan<- prometheus.Metric) {
    ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
    defer cancel()

    lag, err := c.getBlockLag(ctx)
    if err != nil {
        ch <- prometheus.NewInvalidMetric(blockLagDesc, err)
        return
    }

    ch <- prometheus.MustNewConstMetric(
        blockLagDesc,
        prometheus.GaugeValue,
        float64(lag),
        c.nodeName, c.chainID,
    )
    // ... other metrics
}

For Cosmos-based nodes—parsing /status, /net_info, /validators via RPC. For Solana—JSON-RPC methods getHealth, getSlot, getVoteAccounts. For Bitcoin—getblockchaininfo, getpeerinfo.

Aggregation and Storage

Prometheus + VictoriaMetrics for long-term storage. VictoriaMetrics is preferable for multi-chain operations: it compresses time series better, supports federated scraping from multiple Prometheus instances.

# prometheus.yml — scrape config for multi-node environment
scrape_configs:
  - job_name: 'ethereum-nodes'
    scrape_interval: 15s
    scrape_timeout: 10s
    static_configs:
      - targets:
          - 'eth-node-1:9090'
          - 'eth-node-2:9090'
          - 'eth-node-3:9090'
    relabel_configs:
      - source_labels: [__address__]
        target_label: instance

  - job_name: 'cosmos-validators'
    scrape_interval: 30s  # Cosmos block ~6 sec, 30 sec is enough
    static_configs:
      - targets: ['cosmos-val-1:26660', 'cosmos-val-2:26660']

  - job_name: 'solana-rpc'
    scrape_interval: 10s  # Solana ~400ms slot, frequent check needed
    static_configs:
      - targets: ['solana-rpc-1:9101']

Dashboards

Grafana dashboards by structure: Overview (all nodes, all networks, status at a glance), Per-network deep dive (detailed metrics per each network), Validator performance (for staking nodes, including APR and slashing risks), Infrastructure (CPU/RAM/Disk per node).

For public RPC services—additional: request metrics (RPS, latency, error rate), rate limiting statistics, top methods by load.

Development Timeline

Component Timeline
Basic exporters (EVM + 1–2 other chains) 1–2 weeks
Prometheus + VictoriaMetrics + Grafana setup 3–5 days
Alert rules + PagerDuty/Telegram integration 2–3 days
Auto-failover for RPC 1 week
Dashboards + documentation 1 week

Monitoring for 3–5 chains with basic dashboards and alerts—3–4 weeks. Extended system with auto-remediation and custom exporters for non-standard protocols—6–8 weeks. Investment in such a system ranges from $5,000 to $15,000 depending on the number of chains and complexity. Typical savings: $5,000 per month on 10 nodes.

What's Included

  • Development of custom exporters for each chain
  • Setup of Prometheus + VictoriaMetrics + Grafana
  • Creation of alert rules and integration with Telegram/Slack
  • Implementation of auto-failover for RPC nodes
  • Dashboards and documentation
  • Training for your team

Contact us to evaluate your project. Get a consultation on your configuration.

Ethereum

Blockchain Infrastructure Deployment: Nodes, RPC, Indexing

Subgraph fell at 3:47 AM. By morning users saw outdated balances, transactions "hung" in the UI, support received 47 tickets in an hour. Cause: the handler in the subgraph failed on a transaction with a non-standard event log — and the entire index stopped. We have encountered such situations dozens of times. Our experience shows: blockchain infrastructure does not forgive gaps in observability. Guaranteeing uptime without multi-layered monitoring and fault-tolerant architecture is impossible. Over 8 years working with Ethereum, Polygon, and Solana, we have developed an approach that allows predictable deployment of infrastructure of any scale — from a single node to a multichain grid with dozens of subgraphs.

RPC Layer Architecture

Every dApp interaction with the blockchain goes through RPC — the JSON-RPC API provided by a node. Three options:

Managed providers — Alchemy, QuickNode, Infura, Ankr. Minimal operational costs, SLA, built-in monitoring. Limits: rate limits (Alchemy Free: 300 RU/sec), vendor lock, potential downtime during provider incidents. For most projects — the right choice at the start.

Self-owned nodes — full control, no rate limits, no third-party dependence. Cost: archive Ethereum node requires 2.5–3TB SSD, a strong server, and DevOps support. Sync from scratch on Ethereum via Geth/Nethermind — 3–7 days. Justified under high load or latency requirements.

Hybrid — self-owned node as primary, managed provider as fallback. Standard for protocols with high TVL. Proper load balancing can reduce costs by 20–30% compared to pure managed setup. Under high monthly request volume, hybrid saves significantly.

Provider Strength Limitation
Alchemy Supernode, Enhanced APIs, webhooks Expensive on high-volume
QuickNode Low latency, multi-chain More expensive than Alchemy on basic plan
Infura Historical reliability Rate limits on free, one major incident halted half of DeFi
Ankr Cheap, 40+ chains Less stable

How to Set Up an RPC Layer Without a Single Point of Failure?

At least two providers, DNS round-robin with health check every 5 seconds, automatic fallback when latency >500 ms. In practice, this gives 99.99% availability during any provider failure. For protocols with high TVL, we recommend a custom HA-proxy (nginx or Envoy) in front of two managed providers.

Why Is a Hybrid RPC Scheme More Cost-Effective Than Pure Managed?

At high request volumes, managed providers can be very expensive; a hybrid using a self-owned node as primary and a managed fallback cuts costs significantly without losing SLA.

Ethereum Node Clients

Execution clients: Geth (most used), Nethermind (C#, fast sync), Besu (Java, enterprise), Erigon (fastest sync, efficient archive mode ~2TB instead of 3TB).

Consensus clients (post-Merge): Lighthouse (Rust), Prysm (Go), Teku (Java), Nimbus (Nim). Each node after The Merge requires a pair of execution + consensus clients.

For DevOps: eth-docker — Docker Compose configurations for all client combinations. Setting up monitoring via Grafana + Prometheus is mandatory; a standard dashboard is available in each client's repository.

The Graph: Event Indexing

The Graph Protocol — decentralized indexing. A subgraph describes which events from which contracts to index and how to transform them into a GraphQL schema.

Subgraph structure:

  • subgraph.yaml — manifest: contract addresses, startBlock, events to handle
  • schema.graphql — GraphQL schema of entities
  • src/mapping.ts — AssemblyScript event handlers
dataSources:
  - kind: ethereum
    name: UniswapV3Pool
    network: mainnet
    source:
      address: "0x88e6A0c2dDD26FEEb64F039a2c41296FcB3f5640"
      abi: UniswapV3Pool
      startBlock: 12370624
    mapping:
      eventHandlers:
        - event: Swap(indexed address,indexed address,int256,int256,uint160,uint128,int24)
          handler: handleSwap

AssemblyScript handlers — not TypeScript. No nullable types, no closures, no many standard APIs. An error in the handler stops the subgraph indexing on that transaction. Important: add try-catch for operations that can fail (e.g., store.get() for an entity that may not exist).

How to Avoid Subgraph Indexing Stops?

Graph Node logs are monitored in real-time; on hasIndexingErrors = true an alert fires and an automatic node restart (via systemd or Kubernetes). Typical downtime on error — 150–300 seconds to recover. Additionally, for production we set up a watchdog that restarts Graph Node if subgraph lag exceeds 50 blocks.

Choosing Between Hosted Service and Decentralized Network

Graph Hosted Service (free, centralized) is deprecated in favor of Subgraph Studio + Graph Network. For production: deploy on Graph Network with GRT curation signal — the subgraph gets indexers proportional to curation.

Alternatives to The Graph: Ponder (TypeScript, self-hosted, easier to debug), Envio (ultra-fast indexer, supports EVM + non-EVM), Subsquid (TypeScript, own network), Moralis Streams (managed, webhook-based). Our experience shows: for high-load projects with unique logic, Ponder or Envio are more effective — they give full control over the process and do not require GRT tokenomics.

Webhooks and Real-Time Notifications

Alchemy Webhooks and QuickNode Streams allow receiving events in real-time via HTTP webhook or WebSocket. For monitoring addresses, new transactions, mints — this is faster than polling RPC.

Tenderly — platform for monitoring and alerts. You can set up an alert for a specific contract event, balance change, function call with certain parameters. Transaction simulation via Tenderly API is invaluable for debugging.

Monitoring and Observability

Minimum monitoring stack for a protocol:

On-chain: OpenZeppelin Defender Sentinel — watches contract events, triggers webhook or Autotask when conditions are met. Forta Network — community-maintained bots detect anomalies (large withdrawals, flash loans, governance attacks).

Infrastructure: Grafana + Prometheus for nodes, Datadog or Grafana Cloud for managed metrics. Alerts on: node is 10+ blocks behind, RPC latency >500ms, subgraph lag >100 blocks.

Uptime: Better Uptime or PagerDuty on RPC endpoint and subgraph health endpoint (The Graph provides _meta { hasIndexingErrors, block { number } }).

Why Is Monitoring Without Tenderly Insufficient?

Tenderly provides transaction simulation and detailed traces — critical for debugging subgraph and smart contract errors. Forta focuses on network anomalies, not your infrastructure. The combination of Tenderly plus a custom Grafana dashboard covers 90% of incident scenarios.

Multichain Infrastructure

A protocol on 5 chains = 5 separate RPC endpoints, 5 subgraphs, 5 monitoring configs. Manageable but requires deployment automation.

For subgraph multi-network deployment: graph deploy --network mainnet, graph deploy --network arbitrum-one etc. with a unified codebase and network-specific addresses in separate config files.

Chainlink CCIP and LayerZero for cross-chain messaging require monitoring of both chains and transactions on intermediate relayers. A reorg on the source chain after a confirmed mint on the target chain is a classic bridge problem. Solution: wait for finality (on Ethereum ~15 minutes after Merge for economic finality) before confirming on the target chain.

Infrastructure Setup Process

  1. Audit current stack — determine chains, request volume, latency and availability requirements.
  2. Architecture design — select providers, load balancing, redundancy.
  3. Subgraph development — manifest → schema → handlers → testing on local Graph Node → deploy to testnet → mainnet.
  4. Monitoring configuration — Tenderly alerts, Grafana dashboard, PagerDuty integration.
  5. Documentation and runbook — what to do when: subgraph falls behind, RPC downtime, node desync.
  6. Handover to operations — team training, access transfer, first month support.

What's Included

  • Deployment of managed or self-hosted Ethereum, Polygon, BNB Chain nodes
  • RPC layer setup with primary/fallback and load balancing
  • Subgraph development and deployment for your protocol
  • Monitoring connection (Tenderly, Grafana, alerts)
  • Runbook and operations documentation
  • Team training (up to 4 hours online)
  • 30-day support after delivery

Timeline

Task Duration
RPC and basic monitoring setup 1–2 weeks
Subgraph for one protocol 2–4 weeks
Self-hosted node with monitoring 2–3 weeks
Full infrastructure (multi-chain, monitoring, runbooks) 6–10 weeks

All projects are managed in a GitHub/GitLab repository with CI/CD; configuration code stays with you. Order infrastructure deployment — we'll show how to cut costs by 20–30% without losing reliability. Get a consultation — we'll demonstrate how we deployed infrastructure for a protocol with large TVL on Ethereum and Arbitrum. Contact us.