Parsing New Tokens: Detection, Analysis, and Risks

We design and develop full-cycle blockchain solutions: from smart contract architecture to launching DeFi protocols, NFT marketplaces and crypto exchanges. Security audits, tokenomics, integration with existing infrastructure.
Showing 1 of 1All 1305 services
Parsing New Tokens: Detection, Analysis, and Risks
Medium
~2-3 days
Frequently Asked Questions

Blockchain Development Services

Blockchain Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1360
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

You launch a bot for MEV or screening new tokens and face a problem: how to reliably detect contracts in real time? We solve this task through several parallel mechanisms — from monitoring factory contracts to analyzing deployment transactions. Our team has 5+ years of experience in blockchain development and has implemented scanners for 15+ projects in DeFi. We guarantee stable operation and low latency, even at peak loads: Ethereum generates up to 15 blocks per minute, each with hundreds of transactions.

One of our clients, a hedge fund, used the scanner for early token detection on Arbitrum and achieved a latency reduction from 15 seconds to 2. This allowed them to be first into liquidity pools, increasing ROI by 30% per quarter.

How We Parse New Token Data

How to Detect a New Token On-Chain?

Method 1: Monitoring Factory Contracts

Most tokens are deployed through factories: Uniswap V2/V3 factory when creating a pool, Token Factory contracts, or direct deployment with an event. Uniswap V2 factory emits PairCreated when a new pool is created — this is the most reliable signal of a new tradable token:

from web3 import Web3
import asyncio

UNISWAP_V2_FACTORY = "0x5C69bEe701ef814a2B6a3EDD4B1652CB9cc5aA6f"
PAIR_CREATED_TOPIC = "0x0d3648bd0f6ba80134a33ba9275ac585d9d315f0ad8355cddefde31afa28d0e9"

FACTORY_ABI = [{
    "name": "PairCreated",
    "type": "event",
    "inputs": [
        {"name": "token0", "type": "address", "indexed": True},
        {"name": "token1", "type": "address", "indexed": True},
        {"name": "pair", "type": "address", "indexed": False},
        {"name": "", "type": "uint256", "indexed": False}
    ]
}]

async def watch_new_pairs(w3: Web3, callback):
    factory = w3.eth.contract(address=UNISWAP_V2_FACTORY, abi=FACTORY_ABI)
    
    # Subscribe via WebSocket to new events
    event_filter = await w3.eth.filter({
        "address": UNISWAP_V2_FACTORY,
        "topics": [PAIR_CREATED_TOPIC]
    })
    
    while True:
        events = await event_filter.get_new_entries()
        for event_log in events:
            decoded = factory.events.PairCreated().process_log(event_log)
            await callback({
                "token0": decoded.args.token0,
                "token1": decoded.args.token1,
                "pair": decoded.args.pair,
                "block": event_log.blockNumber,
                "tx_hash": event_log.transactionHash.hex()
            })
        await asyncio.sleep(3)

For Uniswap V3 — similarly, listen for PoolCreated on 0x1F98431c8aD98523631AE4a59f267346ea31F984.

Method 2: Detecting ERC-20 Contract Deployment

Direct ERC-20 deployment does not emit standard events. We detect it by analyzing transaction receipts — if contractAddress is non-empty, it's a contract deployment:

async def scan_block_for_deployments(block_number: int, w3: Web3) -> list[dict]:
    block = w3.eth.get_block(block_number, full_transactions=True)
    deployments = []
    
    for tx in block.transactions:
        if tx.to is None:  # tx without to = contract deployment
            receipt = w3.eth.get_transaction_receipt(tx.hash)
            if receipt.contractAddress:
                # Check if it's ERC-20
                token_info = await check_if_erc20(receipt.contractAddress, w3)
                if token_info:
                    deployments.append({
                        "contract": receipt.contractAddress,
                        "deployer": tx["from"],
                        "block": block_number,
                        "tx_hash": tx.hash.hex(),
                        **token_info
                    })
    
    return deployments

async def check_if_erc20(address: str, w3: Web3) -> dict | None:
    """Check for mandatory ERC-20 methods"""
    minimal_abi = [
        {"name": "totalSupply", "type": "function", "inputs": [], "outputs": [{"type": "uint256"}]},
        {"name": "decimals", "type": "function", "inputs": [], "outputs": [{"type": "uint8"}]},
        {"name": "symbol", "type": "function", "inputs": [], "outputs": [{"type": "string"}]},
        {"name": "name", "type": "function", "inputs": [], "outputs": [{"type": "string"}]},
    ]
    try:
        contract = w3.eth.contract(address=address, abil=minimal_abi)
        return {
            "name": contract.functions.name().call(),
            "symbol": contract.functions.symbol().call(),
            "decimals": contract.functions.decimals().call(),
            "total_supply": contract.functions.totalSupply().call()
        }
    except Exception:
        return None  # not ERC-20 or reverting contract

Scanning every block is high load on RPC. On Ethereum mainnet ~15 blocks/minute, each may have hundreds of transactions. You need a dedicated Alchemy/QuickNode plan or your own node. Factory monitoring works 10 times faster than full block scanning.

Data Enrichment After Detection

A bare contract address is not very informative. Immediately after detection, we enrich:

async def enrich_new_token(contract_address: str, w3: Web3) -> dict:
    tasks = await asyncio.gather(
        get_contract_source_code(contract_address),   # Etherscan API
        get_lp_info(contract_address),                 # existing pools
        get_social_links(contract_address),            # from contract or Etherscan
        run_honeypot_check(contract_address),          # sell tax, tradability
        return_exceptions=True
    )
    
    source, lp_info, socials, honeypot = tasks
    return {
        "verified_source": bool(source and not isinstance(source, Exception)),
        "has_liquidity": bool(lp_info and not isinstance(lp_info, Exception)),
        "honeypot_risk": honeypot if not isinstance(honeypot, Exception) else "unknown",
        **socials if not isinstance(socials, Exception) else {}
    }

Automatic Token Risk Analysis

For a scanner with alerts — automatic bytecode and behavior analysis:

SCAM_PATTERNS = {
    "mint_function": "0x40c10f19",    # bytes4 selector for mint(address, uint256)
    "ownership_transfer": "0xf2fde38b",
    "blacklist_function": "0x44337ea1",
}

def check_bytecode_risks(bytecode: str) -> list[str]:
    risks = []
    if len(bytecode) < 100:
        risks.append("minimal_bytecode")  # proxy or placeholder
    
    for name, selector in SCAM_PATTERNS.items():
        if selector[2:] in bytecode:  # remove 0x
            risks.append(name)
    
    return risks

A real honeypot check requires simulating buy and sell transactions via eth_call — this determines buy/sell tax and whether you can sell the token at all. Services: honeypot.is API, GoPlus Security API.

Storage and Indexing

CREATE TABLE new_tokens (
    id              BIGSERIAL PRIMARY KEY,
    chain_id        INTEGER NOT NULL,
    contract        TEXT NOT NULL,
    name            TEXT,
    symbol          TEXT,
    decimals        SMALLINT,
    total_supply    NUMERIC,
    deployer        TEXT NOT NULL,
    deploy_block    INTEGER NOT NULL,
    deploy_tx       TEXT NOT NULL,
    deploy_time     TIMESTAMPTZ NOT NULL,
    verified        BOOLEAN DEFAULT FALSE,
    has_liquidity   BOOLEAN DEFAULT FALSE,
    risk_flags      TEXT[] DEFAULT '{}',
    enriched_at     TIMESTAMPTZ,
    UNIQUE(chain_id, contract)
);

CREATE INDEX ON new_tokens(deploy_time DESC);
CREATE INDEX ON new_tokens(chain_id, symbol);
CREATE INDEX ON new_tokens USING gin(risk_flags);

What Risks Does the Scanner Detect?

We categorize risks into three levels: critical (honeypot, reentrancy), high (mint, blacklist), and medium (high fee, low liquidity). For each token, a summary with flags and recommendations is generated. This allows immediate rejection of 90% of scam tokens at the detection stage.

Comparison of Detection Methods

Method Latency Reliability RPC Load
Factory (Uniswap/PancakeSwap) 1-2 blocks High (guaranteed event) Low (topic filter)
Direct ERC-20 deployment 1 block Medium (not all contracts are ERC-20) High (analyze every transaction)
Non-EVM (Solana) ~1 slot High (InitializeMint) Medium

Why Factory Monitoring Is Faster Than Block Scanning

Factory contracts emit an event when a pool is created — this is the only signal that needs to be tracked. Full block scanning requires processing every transaction, increasing RPC load by 10-15 times and slowing detection. Therefore, for production systems, we recommend combining both methods but prioritizing factory events.

Example of detailed bytecode analysis When a `mint` function with sell restrictions is detected, the token is marked as high-risk. A check via `eth_call` with different amounts reveals hidden fees of up to 99%.

What's Included in Scanner Development

Component Description
Solution architecture Stack selection, load distribution, DB schema
Parsing code Implementation of factory monitoring, direct deployments, enrichment
Risk analysis Bytecode analysis, honeypot check, API integration
REST API Endpoints for data retrieval, filtering, WebSocket
Documentation README, endpoint descriptions, request examples
Support 3 months post-deployment, bug fixes, consultations

Our Process

  1. Analytics — we study your requirements, select networks and methods.
  2. Design — create architecture and data schema.
  3. Implementation — write parsing, enrichment, and API code.
  4. Testing — simulate load, verify on real data.
  5. Deployment — deploy on your infrastructure or ours.
  6. Support — train your team, hand over documentation.

Timeline and Cost

Development time for a basic scanner for 3-4 EVM networks: from 3 to 5 weeks. Cost is calculated individually based on complexity. Get a consultation — we will evaluate your project and offer an optimal solution. Order scanner development for early token detection and reduce monitoring costs.

Blockchain Infrastructure Deployment: Nodes, RPC, Indexing

Subgraph fell at 3:47 AM. By morning users saw outdated balances, transactions "hung" in the UI, support received 47 tickets in an hour. Cause: the handler in the subgraph failed on a transaction with a non-standard event log — and the entire index stopped. We have encountered such situations dozens of times. Our experience shows: blockchain infrastructure does not forgive gaps in observability. Guaranteeing uptime without multi-layered monitoring and fault-tolerant architecture is impossible. Over 8 years working with Ethereum, Polygon, and Solana, we have developed an approach that allows predictable deployment of infrastructure of any scale — from a single node to a multichain grid with dozens of subgraphs.

RPC Layer Architecture

Every dApp interaction with the blockchain goes through RPC — the JSON-RPC API provided by a node. Three options:

Managed providers — Alchemy, QuickNode, Infura, Ankr. Minimal operational costs, SLA, built-in monitoring. Limits: rate limits (Alchemy Free: 300 RU/sec), vendor lock, potential downtime during provider incidents. For most projects — the right choice at the start.

Self-owned nodes — full control, no rate limits, no third-party dependence. Cost: archive Ethereum node requires 2.5–3TB SSD, a strong server, and DevOps support. Sync from scratch on Ethereum via Geth/Nethermind — 3–7 days. Justified under high load or latency requirements.

Hybrid — self-owned node as primary, managed provider as fallback. Standard for protocols with high TVL. Proper load balancing can reduce costs by 20–30% compared to pure managed setup. Under high monthly request volume, hybrid saves significantly.

Provider Strength Limitation
Alchemy Supernode, Enhanced APIs, webhooks Expensive on high-volume
QuickNode Low latency, multi-chain More expensive than Alchemy on basic plan
Infura Historical reliability Rate limits on free, one major incident halted half of DeFi
Ankr Cheap, 40+ chains Less stable

How to Set Up an RPC Layer Without a Single Point of Failure?

At least two providers, DNS round-robin with health check every 5 seconds, automatic fallback when latency >500 ms. In practice, this gives 99.99% availability during any provider failure. For protocols with high TVL, we recommend a custom HA-proxy (nginx or Envoy) in front of two managed providers.

Why Is a Hybrid RPC Scheme More Cost-Effective Than Pure Managed?

At high request volumes, managed providers can be very expensive; a hybrid using a self-owned node as primary and a managed fallback cuts costs significantly without losing SLA.

Ethereum Node Clients

Execution clients: Geth (most used), Nethermind (C#, fast sync), Besu (Java, enterprise), Erigon (fastest sync, efficient archive mode ~2TB instead of 3TB).

Consensus clients (post-Merge): Lighthouse (Rust), Prysm (Go), Teku (Java), Nimbus (Nim). Each node after The Merge requires a pair of execution + consensus clients.

For DevOps: eth-docker — Docker Compose configurations for all client combinations. Setting up monitoring via Grafana + Prometheus is mandatory; a standard dashboard is available in each client's repository.

The Graph: Event Indexing

The Graph Protocol — decentralized indexing. A subgraph describes which events from which contracts to index and how to transform them into a GraphQL schema.

Subgraph structure:

  • subgraph.yaml — manifest: contract addresses, startBlock, events to handle
  • schema.graphql — GraphQL schema of entities
  • src/mapping.ts — AssemblyScript event handlers
dataSources:
  - kind: ethereum
    name: UniswapV3Pool
    network: mainnet
    source:
      address: "0x88e6A0c2dDD26FEEb64F039a2c41296FcB3f5640"
      abi: UniswapV3Pool
      startBlock: 12370624
    mapping:
      eventHandlers:
        - event: Swap(indexed address,indexed address,int256,int256,uint160,uint128,int24)
          handler: handleSwap

AssemblyScript handlers — not TypeScript. No nullable types, no closures, no many standard APIs. An error in the handler stops the subgraph indexing on that transaction. Important: add try-catch for operations that can fail (e.g., store.get() for an entity that may not exist).

How to Avoid Subgraph Indexing Stops?

Graph Node logs are monitored in real-time; on hasIndexingErrors = true an alert fires and an automatic node restart (via systemd or Kubernetes). Typical downtime on error — 150–300 seconds to recover. Additionally, for production we set up a watchdog that restarts Graph Node if subgraph lag exceeds 50 blocks.

Choosing Between Hosted Service and Decentralized Network

Graph Hosted Service (free, centralized) is deprecated in favor of Subgraph Studio + Graph Network. For production: deploy on Graph Network with GRT curation signal — the subgraph gets indexers proportional to curation.

Alternatives to The Graph: Ponder (TypeScript, self-hosted, easier to debug), Envio (ultra-fast indexer, supports EVM + non-EVM), Subsquid (TypeScript, own network), Moralis Streams (managed, webhook-based). Our experience shows: for high-load projects with unique logic, Ponder or Envio are more effective — they give full control over the process and do not require GRT tokenomics.

Webhooks and Real-Time Notifications

Alchemy Webhooks and QuickNode Streams allow receiving events in real-time via HTTP webhook or WebSocket. For monitoring addresses, new transactions, mints — this is faster than polling RPC.

Tenderly — platform for monitoring and alerts. You can set up an alert for a specific contract event, balance change, function call with certain parameters. Transaction simulation via Tenderly API is invaluable for debugging.

Monitoring and Observability

Minimum monitoring stack for a protocol:

On-chain: OpenZeppelin Defender Sentinel — watches contract events, triggers webhook or Autotask when conditions are met. Forta Network — community-maintained bots detect anomalies (large withdrawals, flash loans, governance attacks).

Infrastructure: Grafana + Prometheus for nodes, Datadog or Grafana Cloud for managed metrics. Alerts on: node is 10+ blocks behind, RPC latency >500ms, subgraph lag >100 blocks.

Uptime: Better Uptime or PagerDuty on RPC endpoint and subgraph health endpoint (The Graph provides _meta { hasIndexingErrors, block { number } }).

Why Is Monitoring Without Tenderly Insufficient?

Tenderly provides transaction simulation and detailed traces — critical for debugging subgraph and smart contract errors. Forta focuses on network anomalies, not your infrastructure. The combination of Tenderly plus a custom Grafana dashboard covers 90% of incident scenarios.

Multichain Infrastructure

A protocol on 5 chains = 5 separate RPC endpoints, 5 subgraphs, 5 monitoring configs. Manageable but requires deployment automation.

For subgraph multi-network deployment: graph deploy --network mainnet, graph deploy --network arbitrum-one etc. with a unified codebase and network-specific addresses in separate config files.

Chainlink CCIP and LayerZero for cross-chain messaging require monitoring of both chains and transactions on intermediate relayers. A reorg on the source chain after a confirmed mint on the target chain is a classic bridge problem. Solution: wait for finality (on Ethereum ~15 minutes after Merge for economic finality) before confirming on the target chain.

Infrastructure Setup Process

  1. Audit current stack — determine chains, request volume, latency and availability requirements.
  2. Architecture design — select providers, load balancing, redundancy.
  3. Subgraph development — manifest → schema → handlers → testing on local Graph Node → deploy to testnet → mainnet.
  4. Monitoring configuration — Tenderly alerts, Grafana dashboard, PagerDuty integration.
  5. Documentation and runbook — what to do when: subgraph falls behind, RPC downtime, node desync.
  6. Handover to operations — team training, access transfer, first month support.

What's Included

  • Deployment of managed or self-hosted Ethereum, Polygon, BNB Chain nodes
  • RPC layer setup with primary/fallback and load balancing
  • Subgraph development and deployment for your protocol
  • Monitoring connection (Tenderly, Grafana, alerts)
  • Runbook and operations documentation
  • Team training (up to 4 hours online)
  • 30-day support after delivery

Timeline

Task Duration
RPC and basic monitoring setup 1–2 weeks
Subgraph for one protocol 2–4 weeks
Self-hosted node with monitoring 2–3 weeks
Full infrastructure (multi-chain, monitoring, runbooks) 6–10 weeks

All projects are managed in a GitHub/GitLab repository with CI/CD; configuration code stays with you. Order infrastructure deployment — we'll show how to cut costs by 20–30% without losing reliability. Get a consultation — we'll demonstrate how we deployed infrastructure for a protocol with large TVL on Ethereum and Arbitrum. Contact us.