Collecting On-Chain Data: Methods and Tools

We design and develop full-cycle blockchain solutions: from smart contract architecture to launching DeFi protocols, NFT marketplaces and crypto exchanges. Security audits, tokenomics, integration with existing infrastructure.
Showing 1 of 1All 1305 services
Collecting On-Chain Data: Methods and Tools
Medium
~2-3 days
Frequently Asked Questions

Blockchain Development Services

Blockchain Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1360
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

When exporting history from blockchain explorers, we faced harsh limitations: Etherscan returns a maximum of 10,000 records per request, 5 requests per second on the free plan, no streaming. If you need to collect all transactions of the USDT contract over the last three years – that's over 30 million records. For DeFi strategy analysis or whale transaction tracking, millions of records are needed – standard APIs fall short. A simple solution won't work: pagination via page won't give more than 10,000. You need a strategy, and we found it: block-based pagination with subsequent deduplication.

In this article, we'll break down the main approaches: from direct Etherscan API to using Alchemy and Moralis, as well as HTML parsing and working with your own node. Each method differs in speed, completeness, and cost. For example, Alchemy allows you to get all transactions of an address with a single request without a record limit, and a self-hosted node gives access to internal calls.

Alchemy is 5 times better than Etherscan for rate limits, has no record limit thanks to pageKey, supports WebSocket streaming, and works across multiple chains. For projects requiring maximum performance, Alchemy is the go-to choice.

If you need to gather historical data for analysis – contact us, we'll find the optimal tool within one day. Typical project costs range from $5,000 to $20,000, and clients often save 30-50% on infrastructure by using our solutions.

Bypassing the Etherscan API 10,000 record limit

Etherscan provides startblock and endblock parameters. Block-based pagination lets you pull all transactions. Here are the steps:

  1. Set startblock to 0 (or the earliest block).
  2. Call the API with offset=10000 and page=1.
  3. Process the returned transactions.
  4. If the number of transactions is 10000, set startblock to the last block number + 1 and repeat from step 2.
  5. If a single block contains more than 10000 transactions, use nested pagination by incrementing page within that block.
  6. Add a delay between requests to respect rate limits.

The code below iterates over blocks, incrementing startblock to the last encountered one:

import httpx
import asyncio
from typing import AsyncGenerator

async def get_all_transactions(
    address: str, 
    api_key: str,
    start_block: int = 0
) -> AsyncGenerator[dict, None]:
    """Export ALL transactions of an address via block-based pagination"""
    
    base_url = "https://api.etherscan.io/api"
    current_block = start_block
    
    while True:
        async with httpx.AsyncClient() as client:
            resp = await client.get(base_url, params={
                "module": "account",
                "action": "txlist",
                "address": address,
                "startblock": current_block,
                "endblock": 99999999,
                "sort": "asc",
                "apikey": api_key,
                "offset": 10000,
                "page": 1,
            })
        
        data = resp.json()
        if data["status"] != "1" or not data["result"]:
            break
            
        txs = data["result"]
        for tx in txs:
            yield tx
        
        if len(txs) < 10000:
            break
        
        current_block = int(txs[-1]["blockNumber"]) + 1
        await asyncio.sleep(0.2)

Important nuance: if a single block contains >10,000 transactions for that address, the loop will hang. For such cases, nested pagination with the page parameter inside the block is needed.

Advantages of Alchemy and Moralis over Etherscan

Alchemy and Moralis remove most of Etherscan's limitations. Alchemy has a rate limit of 25 req/s, which is 5 times higher than Etherscan's (5 req/s). Here's a comparison:

Parameter Etherscan (free) Alchemy (free tier) Moralis (free tier)
Record limit 10000 None (pageKey) None (cursor)
Rate limit 5 req/s 25 req/s 100 req/min
Streaming No WebSocket (Enhanced API) WebSocket (Real-time)
Cross-chain No Yes Yes

Alchemy getAssetTransfers returns ETH + ERC-20 + ERC-721 in one call.

import { Alchemy, Network } from 'alchemy-sdk';

const alchemy = new Alchemy({ apiKey: process.env.ALCHEMY_KEY, network: Network.ETH_MAINNET });

const transfers = await alchemy.core.getAssetTransfers({
  fromAddress: '0x...',
  category: ['external', 'erc20', 'erc721', 'erc1155'],
  withMetadata: true,
  maxCount: 1000,
});

if (transfers.pageKey) {
  const more = await alchemy.core.getAssetTransfers({
    pageKey: transfers.pageKey,
  });
}

Alchemy Asset Transfers API recommends using pageKey for pagination. Moralis additionally provides internal transactions via getWalletTransactionsVerbose. For a quick estimation, contact us – we'll prepare a prototype within 2 days.

Justification for HTML page scraping

If the API doesn't return token holders or verified contracts, we use HTML scraping. The following example collects token holders from Etherscan:

import httpx
from bs4 import BeautifulSoup
import asyncio

async def get_token_holders(token_address: str, pages: int = 10) -> list[dict]:
    headers = {
        "User-Agent": "Mozilla/5.0",
    }
    holders = []
    
    async with httpx.AsyncClient(headers=headers) as client:
        for page in range(1, pages + 1):
            resp = await client.get(
                f"https://etherscan.io/token/{token_address}",
                params={"a": "#holders", "p": page}
            )
            soup = BeautifulSoup(resp.text, 'html.parser')
            table = soup.find('table', {'id': 'holdersTable'})
            if not table:
                break
            for row in table.find_all('tr')[1:]:
                cols = row.find_all('td')
                if len(cols) >= 3:
                    holders.append({
                        'rank': cols[0].text.strip(),
                        'address': cols[1].find('a')['href'].split('/')[-1],
                        'quantity': cols[2].text.strip(),
                    })
            await asyncio.sleep(2)
    return holders

Etherscan is protected by Cloudflare – for large-scale collection, you need residential proxies or the official API. Scraping BscScan and Solscan works similarly.

Avoiding duplicates during parallel collection

When collecting in parallel from multiple sources, duplicates are inevitable. Use ON CONFLICT in PostgreSQL:

CREATE TABLE eth_transactions (
    tx_hash      CHAR(66) PRIMARY KEY,
    block_number BIGINT NOT NULL,
    from_address CHAR(42) NOT NULL,
    to_address   CHAR(42),
    value        NUMERIC(38) DEFAULT 0,
    gas_used     BIGINT,
    status       SMALLINT,
    ts           TIMESTAMPTZ
);

INSERT INTO eth_transactions VALUES (...)
ON CONFLICT (tx_hash) DO NOTHING;

For events, the unique key is (tx_hash, log_index).

Working directly with a node

For maximum completeness (internal transactions, storage slots, MEV), run your own node with Erigon and --tracing. This gives:

  • All internal transactions with no limits
  • Call tracing
  • Storage data
curl -X POST $ETH_RPC_URL \
  -H "Content-Type: application/json" \
  -d '{"jsonrpc":"2.0","method":"eth_getBlockReceipts","params":["0x1234567"],"id":1}'

eth_getBlockReceipts EIP-1559 returns all receipts of a block in one request – an alternative to N separate eth_getTransactionReceipt calls.

Data collection stages

Stage Description Duration
Source analysis Determine required data and its location (API, HTML, node) 1-2 days
Tool selection Etherscan API, Alchemy, Moralis or a self-hosted node 1 day
Parser development Asynchronous collection with rate limit and pagination 2-5 days
Deduplication & normalization Merging data into a unified schema 1-2 days
Testing on real volume Millions of transactions, check for gaps 2-3 days
Deployment Containerization, monitoring, automatic restart 1-2 days

Our process takes from 3 to 10 days depending on complexity.

What's included in the result

  • Documentation on the data schema and collection methods.
  • Ready-to-use Python/TypeScript code with async support.
  • Pagination configuration for each explorer.
  • Deployment on a server or cloud (AWS Lambda, Kubernetes).
  • Team training on using the system.
  • One month of support after delivery.

Our experience spans over 30 projects in on-chain data collection and analysis. Typical project costs start at $5,000 for simple exports and can go up to $20,000 for complex multi-chain setups. Get a consultation – we'll help you choose the optimal tool for your task. Place a request, and we'll set up data collection of any complexity.

Blockchain Infrastructure Deployment: Nodes, RPC, Indexing

Subgraph fell at 3:47 AM. By morning users saw outdated balances, transactions "hung" in the UI, support received 47 tickets in an hour. Cause: the handler in the subgraph failed on a transaction with a non-standard event log — and the entire index stopped. We have encountered such situations dozens of times. Our experience shows: blockchain infrastructure does not forgive gaps in observability. Guaranteeing uptime without multi-layered monitoring and fault-tolerant architecture is impossible. Over 8 years working with Ethereum, Polygon, and Solana, we have developed an approach that allows predictable deployment of infrastructure of any scale — from a single node to a multichain grid with dozens of subgraphs.

RPC Layer Architecture

Every dApp interaction with the blockchain goes through RPC — the JSON-RPC API provided by a node. Three options:

Managed providers — Alchemy, QuickNode, Infura, Ankr. Minimal operational costs, SLA, built-in monitoring. Limits: rate limits (Alchemy Free: 300 RU/sec), vendor lock, potential downtime during provider incidents. For most projects — the right choice at the start.

Self-owned nodes — full control, no rate limits, no third-party dependence. Cost: archive Ethereum node requires 2.5–3TB SSD, a strong server, and DevOps support. Sync from scratch on Ethereum via Geth/Nethermind — 3–7 days. Justified under high load or latency requirements.

Hybrid — self-owned node as primary, managed provider as fallback. Standard for protocols with high TVL. Proper load balancing can reduce costs by 20–30% compared to pure managed setup. Under high monthly request volume, hybrid saves significantly.

Provider Strength Limitation
Alchemy Supernode, Enhanced APIs, webhooks Expensive on high-volume
QuickNode Low latency, multi-chain More expensive than Alchemy on basic plan
Infura Historical reliability Rate limits on free, one major incident halted half of DeFi
Ankr Cheap, 40+ chains Less stable

How to Set Up an RPC Layer Without a Single Point of Failure?

At least two providers, DNS round-robin with health check every 5 seconds, automatic fallback when latency >500 ms. In practice, this gives 99.99% availability during any provider failure. For protocols with high TVL, we recommend a custom HA-proxy (nginx or Envoy) in front of two managed providers.

Why Is a Hybrid RPC Scheme More Cost-Effective Than Pure Managed?

At high request volumes, managed providers can be very expensive; a hybrid using a self-owned node as primary and a managed fallback cuts costs significantly without losing SLA.

Ethereum Node Clients

Execution clients: Geth (most used), Nethermind (C#, fast sync), Besu (Java, enterprise), Erigon (fastest sync, efficient archive mode ~2TB instead of 3TB).

Consensus clients (post-Merge): Lighthouse (Rust), Prysm (Go), Teku (Java), Nimbus (Nim). Each node after The Merge requires a pair of execution + consensus clients.

For DevOps: eth-docker — Docker Compose configurations for all client combinations. Setting up monitoring via Grafana + Prometheus is mandatory; a standard dashboard is available in each client's repository.

The Graph: Event Indexing

The Graph Protocol — decentralized indexing. A subgraph describes which events from which contracts to index and how to transform them into a GraphQL schema.

Subgraph structure:

  • subgraph.yaml — manifest: contract addresses, startBlock, events to handle
  • schema.graphql — GraphQL schema of entities
  • src/mapping.ts — AssemblyScript event handlers
dataSources:
  - kind: ethereum
    name: UniswapV3Pool
    network: mainnet
    source:
      address: "0x88e6A0c2dDD26FEEb64F039a2c41296FcB3f5640"
      abi: UniswapV3Pool
      startBlock: 12370624
    mapping:
      eventHandlers:
        - event: Swap(indexed address,indexed address,int256,int256,uint160,uint128,int24)
          handler: handleSwap

AssemblyScript handlers — not TypeScript. No nullable types, no closures, no many standard APIs. An error in the handler stops the subgraph indexing on that transaction. Important: add try-catch for operations that can fail (e.g., store.get() for an entity that may not exist).

How to Avoid Subgraph Indexing Stops?

Graph Node logs are monitored in real-time; on hasIndexingErrors = true an alert fires and an automatic node restart (via systemd or Kubernetes). Typical downtime on error — 150–300 seconds to recover. Additionally, for production we set up a watchdog that restarts Graph Node if subgraph lag exceeds 50 blocks.

Choosing Between Hosted Service and Decentralized Network

Graph Hosted Service (free, centralized) is deprecated in favor of Subgraph Studio + Graph Network. For production: deploy on Graph Network with GRT curation signal — the subgraph gets indexers proportional to curation.

Alternatives to The Graph: Ponder (TypeScript, self-hosted, easier to debug), Envio (ultra-fast indexer, supports EVM + non-EVM), Subsquid (TypeScript, own network), Moralis Streams (managed, webhook-based). Our experience shows: for high-load projects with unique logic, Ponder or Envio are more effective — they give full control over the process and do not require GRT tokenomics.

Webhooks and Real-Time Notifications

Alchemy Webhooks and QuickNode Streams allow receiving events in real-time via HTTP webhook or WebSocket. For monitoring addresses, new transactions, mints — this is faster than polling RPC.

Tenderly — platform for monitoring and alerts. You can set up an alert for a specific contract event, balance change, function call with certain parameters. Transaction simulation via Tenderly API is invaluable for debugging.

Monitoring and Observability

Minimum monitoring stack for a protocol:

On-chain: OpenZeppelin Defender Sentinel — watches contract events, triggers webhook or Autotask when conditions are met. Forta Network — community-maintained bots detect anomalies (large withdrawals, flash loans, governance attacks).

Infrastructure: Grafana + Prometheus for nodes, Datadog or Grafana Cloud for managed metrics. Alerts on: node is 10+ blocks behind, RPC latency >500ms, subgraph lag >100 blocks.

Uptime: Better Uptime or PagerDuty on RPC endpoint and subgraph health endpoint (The Graph provides _meta { hasIndexingErrors, block { number } }).

Why Is Monitoring Without Tenderly Insufficient?

Tenderly provides transaction simulation and detailed traces — critical for debugging subgraph and smart contract errors. Forta focuses on network anomalies, not your infrastructure. The combination of Tenderly plus a custom Grafana dashboard covers 90% of incident scenarios.

Multichain Infrastructure

A protocol on 5 chains = 5 separate RPC endpoints, 5 subgraphs, 5 monitoring configs. Manageable but requires deployment automation.

For subgraph multi-network deployment: graph deploy --network mainnet, graph deploy --network arbitrum-one etc. with a unified codebase and network-specific addresses in separate config files.

Chainlink CCIP and LayerZero for cross-chain messaging require monitoring of both chains and transactions on intermediate relayers. A reorg on the source chain after a confirmed mint on the target chain is a classic bridge problem. Solution: wait for finality (on Ethereum ~15 minutes after Merge for economic finality) before confirming on the target chain.

Infrastructure Setup Process

  1. Audit current stack — determine chains, request volume, latency and availability requirements.
  2. Architecture design — select providers, load balancing, redundancy.
  3. Subgraph development — manifest → schema → handlers → testing on local Graph Node → deploy to testnet → mainnet.
  4. Monitoring configuration — Tenderly alerts, Grafana dashboard, PagerDuty integration.
  5. Documentation and runbook — what to do when: subgraph falls behind, RPC downtime, node desync.
  6. Handover to operations — team training, access transfer, first month support.

What's Included

  • Deployment of managed or self-hosted Ethereum, Polygon, BNB Chain nodes
  • RPC layer setup with primary/fallback and load balancing
  • Subgraph development and deployment for your protocol
  • Monitoring connection (Tenderly, Grafana, alerts)
  • Runbook and operations documentation
  • Team training (up to 4 hours online)
  • 30-day support after delivery

Timeline

Task Duration
RPC and basic monitoring setup 1–2 weeks
Subgraph for one protocol 2–4 weeks
Self-hosted node with monitoring 2–3 weeks
Full infrastructure (multi-chain, monitoring, runbooks) 6–10 weeks

All projects are managed in a GitHub/GitLab repository with CI/CD; configuration code stays with you. Order infrastructure deployment — we'll show how to cut costs by 20–30% without losing reliability. Get a consultation — we'll demonstrate how we deployed infrastructure for a protocol with large TVL on Ethereum and Arbitrum. Contact us.