When exporting history from blockchain explorers, we faced harsh limitations: Etherscan returns a maximum of 10,000 records per request, 5 requests per second on the free plan, no streaming. If you need to collect all transactions of the USDT contract over the last three years – that's over 30 million records. For DeFi strategy analysis or whale transaction tracking, millions of records are needed – standard APIs fall short. A simple solution won't work: pagination via page won't give more than 10,000. You need a strategy, and we found it: block-based pagination with subsequent deduplication.
In this article, we'll break down the main approaches: from direct Etherscan API to using Alchemy and Moralis, as well as HTML parsing and working with your own node. Each method differs in speed, completeness, and cost. For example, Alchemy allows you to get all transactions of an address with a single request without a record limit, and a self-hosted node gives access to internal calls.
Alchemy is 5 times better than Etherscan for rate limits, has no record limit thanks to pageKey, supports WebSocket streaming, and works across multiple chains. For projects requiring maximum performance, Alchemy is the go-to choice.
If you need to gather historical data for analysis – contact us, we'll find the optimal tool within one day. Typical project costs range from $5,000 to $20,000, and clients often save 30-50% on infrastructure by using our solutions.
Bypassing the Etherscan API 10,000 record limit
Etherscan provides startblock and endblock parameters. Block-based pagination lets you pull all transactions. Here are the steps:
- Set startblock to 0 (or the earliest block).
- Call the API with
offset=10000andpage=1. - Process the returned transactions.
- If the number of transactions is 10000, set startblock to the last block number + 1 and repeat from step 2.
- If a single block contains more than 10000 transactions, use nested pagination by incrementing
pagewithin that block. - Add a delay between requests to respect rate limits.
The code below iterates over blocks, incrementing startblock to the last encountered one:
import httpx
import asyncio
from typing import AsyncGenerator
async def get_all_transactions(
address: str,
api_key: str,
start_block: int = 0
) -> AsyncGenerator[dict, None]:
"""Export ALL transactions of an address via block-based pagination"""
base_url = "https://api.etherscan.io/api"
current_block = start_block
while True:
async with httpx.AsyncClient() as client:
resp = await client.get(base_url, params={
"module": "account",
"action": "txlist",
"address": address,
"startblock": current_block,
"endblock": 99999999,
"sort": "asc",
"apikey": api_key,
"offset": 10000,
"page": 1,
})
data = resp.json()
if data["status"] != "1" or not data["result"]:
break
txs = data["result"]
for tx in txs:
yield tx
if len(txs) < 10000:
break
current_block = int(txs[-1]["blockNumber"]) + 1
await asyncio.sleep(0.2)
Important nuance: if a single block contains >10,000 transactions for that address, the loop will hang. For such cases, nested pagination with the page parameter inside the block is needed.
Advantages of Alchemy and Moralis over Etherscan
Alchemy and Moralis remove most of Etherscan's limitations. Alchemy has a rate limit of 25 req/s, which is 5 times higher than Etherscan's (5 req/s). Here's a comparison:
| Parameter | Etherscan (free) | Alchemy (free tier) | Moralis (free tier) |
|---|---|---|---|
| Record limit | 10000 | None (pageKey) | None (cursor) |
| Rate limit | 5 req/s | 25 req/s | 100 req/min |
| Streaming | No | WebSocket (Enhanced API) | WebSocket (Real-time) |
| Cross-chain | No | Yes | Yes |
Alchemy getAssetTransfers returns ETH + ERC-20 + ERC-721 in one call.
import { Alchemy, Network } from 'alchemy-sdk';
const alchemy = new Alchemy({ apiKey: process.env.ALCHEMY_KEY, network: Network.ETH_MAINNET });
const transfers = await alchemy.core.getAssetTransfers({
fromAddress: '0x...',
category: ['external', 'erc20', 'erc721', 'erc1155'],
withMetadata: true,
maxCount: 1000,
});
if (transfers.pageKey) {
const more = await alchemy.core.getAssetTransfers({
pageKey: transfers.pageKey,
});
}
Alchemy Asset Transfers API recommends using pageKey for pagination. Moralis additionally provides internal transactions via getWalletTransactionsVerbose. For a quick estimation, contact us – we'll prepare a prototype within 2 days.
Justification for HTML page scraping
If the API doesn't return token holders or verified contracts, we use HTML scraping. The following example collects token holders from Etherscan:
import httpx
from bs4 import BeautifulSoup
import asyncio
async def get_token_holders(token_address: str, pages: int = 10) -> list[dict]:
headers = {
"User-Agent": "Mozilla/5.0",
}
holders = []
async with httpx.AsyncClient(headers=headers) as client:
for page in range(1, pages + 1):
resp = await client.get(
f"https://etherscan.io/token/{token_address}",
params={"a": "#holders", "p": page}
)
soup = BeautifulSoup(resp.text, 'html.parser')
table = soup.find('table', {'id': 'holdersTable'})
if not table:
break
for row in table.find_all('tr')[1:]:
cols = row.find_all('td')
if len(cols) >= 3:
holders.append({
'rank': cols[0].text.strip(),
'address': cols[1].find('a')['href'].split('/')[-1],
'quantity': cols[2].text.strip(),
})
await asyncio.sleep(2)
return holders
Etherscan is protected by Cloudflare – for large-scale collection, you need residential proxies or the official API. Scraping BscScan and Solscan works similarly.
Avoiding duplicates during parallel collection
When collecting in parallel from multiple sources, duplicates are inevitable. Use ON CONFLICT in PostgreSQL:
CREATE TABLE eth_transactions (
tx_hash CHAR(66) PRIMARY KEY,
block_number BIGINT NOT NULL,
from_address CHAR(42) NOT NULL,
to_address CHAR(42),
value NUMERIC(38) DEFAULT 0,
gas_used BIGINT,
status SMALLINT,
ts TIMESTAMPTZ
);
INSERT INTO eth_transactions VALUES (...)
ON CONFLICT (tx_hash) DO NOTHING;
For events, the unique key is (tx_hash, log_index).
Working directly with a node
For maximum completeness (internal transactions, storage slots, MEV), run your own node with Erigon and --tracing. This gives:
- All internal transactions with no limits
- Call tracing
- Storage data
curl -X POST $ETH_RPC_URL \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","method":"eth_getBlockReceipts","params":["0x1234567"],"id":1}'
eth_getBlockReceipts EIP-1559 returns all receipts of a block in one request – an alternative to N separate eth_getTransactionReceipt calls.
Data collection stages
| Stage | Description | Duration |
|---|---|---|
| Source analysis | Determine required data and its location (API, HTML, node) | 1-2 days |
| Tool selection | Etherscan API, Alchemy, Moralis or a self-hosted node | 1 day |
| Parser development | Asynchronous collection with rate limit and pagination | 2-5 days |
| Deduplication & normalization | Merging data into a unified schema | 1-2 days |
| Testing on real volume | Millions of transactions, check for gaps | 2-3 days |
| Deployment | Containerization, monitoring, automatic restart | 1-2 days |
Our process takes from 3 to 10 days depending on complexity.
What's included in the result
- Documentation on the data schema and collection methods.
- Ready-to-use Python/TypeScript code with async support.
- Pagination configuration for each explorer.
- Deployment on a server or cloud (AWS Lambda, Kubernetes).
- Team training on using the system.
- One month of support after delivery.
Our experience spans over 30 projects in on-chain data collection and analysis. Typical project costs start at $5,000 for simple exports and can go up to $20,000 for complex multi-chain setups. Get a consultation – we'll help you choose the optimal tool for your task. Place a request, and we'll set up data collection of any complexity.







