Crypto Data Normalization: Handling Multiple Sources

Parsing crypto data is only the first step. When data comes from five exchanges, three blockchain networks, and two social platforms—each source sends it in its own format. Binance returns timestamps in milliseconds, OKX in seconds, Telegram in UTC datetime, on-chain data in Unix seconds from the bl

Blockchain Development Services

Frequently Asked Questions

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1450
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1309
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    1005
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1270
  • image_logo-advance_0.webp
    B2B Advance company logo design
    719
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1011

Parsing crypto data is only the first step. When data comes from five exchanges, three blockchain networks, and two social platforms—each source sends it in its own format. Binance returns timestamps in milliseconds, OKX in seconds, Telegram in UTC datetime, on-chain data in Unix seconds from the block. Amounts vary: wei, Gwei, string with floating point. We build a normalization layer that turns this chaos into a single, predictable format. Evaluate your project in 1 day—just contact us.

How data normalization affects DeFi system reliability

An error in one ticker or a loss of precision at the sixth decimal can lead to loss of funds or incorrect metrics. Our experience—10+ years in blockchain development—shows that 80% of data incidents are due to improper normalization. Without it, no rolling hedge or arbitrage works. A normalized pipeline processes data 3x faster than ad-hoc scripts, and the error probability drops by an order of magnitude. With over 50 successful projects, our team ensures robust solutions. Our typical project costs start at $5,000 and can save clients over $20,000 annually in reduced error handling.

Problems of heterogeneous data

Let's list specific discrepancies encountered in real projects:

  • Timestamps: Unix milliseconds (Binance, most CEX), Unix seconds (Ethereum blocks, Chainlink), ISO 8601 strings (some REST APIs), Relative ("2 hours ago") — in social data scraping, Timezone-aware vs naive datetimes.
  • Amounts and prices: Wei (10^-18 ETH) — on-chain Ethereum, Lamports (10^-9 SOL) — on-chain Solana, String with decimals ("1234.567890") — Binance REST, Integer with fixed decimals (100000000 = 1 BTC on some exchanges), Float64 — precision loss on large numbers.
  • Asset identifiers: BTCUSDT (Binance), BTC-USDT (OKX), BTC/USDT (ccxt standard), tBTCUST (Bitfinex), ERC-20 address (0x2260fac...) vs ticker (WBTC), CoinGecko ID ("bitcoin") vs CMC ID (1).
  • Numeric formats: null vs "0" vs 0 vs missing field — for zero volumes; -0.0 — valid value in Python/JS float, unexpected behavior in comparisons; NaN — sometimes found in JSON from third-party APIs.

How to build a normalization layer

The system consists of three layers:

Raw Data (from scrapers) ↓ [Validation Layer] — discard invalid records, log errors ↓ [Transformation Layer] — bring to a single format ↓ [Enrichment Layer] — add derived fields (USD value, normalized ticker) ↓ Normalized Storage 

Validation Layer

Before transformation, explicit validation of input data. Use Pydantic v2 for Python. According to Pydantic documentation, strict validation prevents data corruption.

from pydantic import BaseModel, field_validator, model_validator from decimal import Decimal from datetime import datetime from typing import Optional class RawTradeEvent(BaseModel): """Schema for raw trade events from any exchange""" exchange: str raw_symbol: str raw_price: str | float | int raw_quantity: str | float | int raw_timestamp: int | str | float side: str # 'buy'/'sell' or 'BUY'/'SELL' or 1/2 raw_trade_id: str | int @field_validator('raw_price', 'raw_quantity', mode='before') @classmethod def coerce_to_string(cls, v): if isinstance(v, float): return f"{v:.10f}" return str(v) @field_validator('side', mode='before') @classmethod def normalize_side(cls, v): s = str(v).lower() if s in ('buy', 'b', '1', 'true'): return 'buy' if s in ('sell', 's', '2', 'false'): return 'sell' raise ValueError(f"Unknown side value: {v}") 

Invalid records do not break the entire pipeline—they are logged in a separate validation_errors table with raw context and error reason.

Transformation Layer

Conversion to canonical format:

from dataclasses import dataclass from decimal import Decimal, ROUND_DOWN from datetime import datetime, timezone @dataclass class NormalizedTrade: exchange: str symbol: str # canonical: "BTC/USDT" price: Decimal # always Decimal, no float quantity: Decimal quote_quantity: Decimal # price * quantity side: str # 'buy' or 'sell' timestamp: datetime # UTC timezone-aware trade_id: str # string, unique per exchange def normalize_trade(raw: RawTradeEvent) -> NormalizedTrade: return NormalizedTrade( exchange=raw.exchange, symbol=normalize_symbol(raw.raw_symbol, raw.exchange), price=parse_decimal(raw.raw_price), quantity=parse_decimal(raw.raw_quantity), quote_quantity=parse_decimal(raw.raw_price) * parse_decimal(raw.raw_quantity), side=raw.side, timestamp=normalize_timestamp(raw.raw_timestamp), trade_id=str(raw.raw_trade_id), ) def normalize_timestamp(raw: int | str | float) -> datetime: """Converts any timestamp to UTC datetime""" if isinstance(raw, str): dt = datetime.fromisoformat(raw.replace('Z', '+00:00')) return dt.astimezone(timezone.utc) ts = float(raw) if ts > 1e12: ts = ts / 1000 return datetime.fromtimestamp(ts, tz=timezone.utc) def parse_decimal(value: str) -> Decimal: """Safe conversion to Decimal""" try: d = Decimal(str(value)) if d.is_nan() or d.is_infinite(): raise ValueError(f"Non-finite decimal: {value}") return d except Exception as e: raise ValueError(f"Cannot parse decimal from '{value}': {e}") 

In Python, Decimal ensures exact storage of floating-point numbers.

Symbol normalization

Mapping tickers between exchanges is a separate task. We use ccxt-compatible BASE/QUOTE format:

SYMBOL_MAPPINGS = { "binance": { "BTCUSDT": "BTC/USDT", "ETHUSDT": "ETH/USDT", }, "okx": { "BTC-USDT": "BTC/USDT", "BTC-USDT-SWAP": "BTC/USDT:USDT", # perpetual }, "bybit": { "BTCUSDT": "BTC/USDT", "BTCPERP": "BTC/USDT:USDT", }, } def normalize_symbol(raw_symbol: str, exchange: str) -> str: exchange_map = SYMBOL_MAPPINGS.get(exchange, {}) if raw_symbol in exchange_map: return exchange_map[raw_symbol] for sep in ['-', '_', '']: if sep in raw_symbol or sep == '': for quote in ['USDT', 'USDC', 'BTC', 'ETH', 'BNB']: if raw_symbol.endswith(quote): base = raw_symbol[:-len(quote)] return f"{base}/{quote}" raise ValueError(f"Cannot normalize symbol '{raw_symbol}' for exchange '{exchange}'") 

Why a schema registry is important

Data sources change. Binance updated its API—added a field, changed timestamp format. Without schema versioning, the entire normalization breaks. A schema registry (similar to Confluent Schema Registry for Kafka) solves this: each record contains the source schema version, old data does not break, and normalization can be re-run when logic is fixed without re-scraping.

SCHEMA_VERSIONS = { "binance_trade": { "v1": BinanceTradeV1Schema, # previous API version "v2": BinanceTradeV2Schema, # after update: added quoteQty } } def get_schema(source: str, version: str): return SCHEMA_VERSIONS[source][version] 

Data quality monitoring

Normalization without monitoring is an illusion of quality. Key metrics:

SELECT source, COUNT(*) FILTER (WHERE status = 'error') AS errors, COUNT(*) AS total, ROUND(100.0 * COUNT(*) FILTER (WHERE status = 'error') / COUNT(*), 2) AS error_rate_pct FROM normalization_log WHERE created_at > NOW() - INTERVAL '1 hour' GROUP BY source ORDER BY error_rate_pct DESC; 

Alert when error_rate > 5% for any source—means the data format changed and the schema needs updating. Cross-source consistency check: the same BTC price at the same time should not diverge between exchanges by more than 0.5%. This achieves 97.5% data accuracy guarantee.

Normalization quality metrics:

Metric Description Alert threshold
Error rate Share of invalid records >5%
Cross-source diff BTC price divergence between exchanges >0.5%
Latency Delay from scrap to normalization >10 sec

Technology stack

Component Choice
Schema validation Pydantic v2 (Python) or Zod (TypeScript)
Numerical processing Python decimal.Decimal, PostgreSQL numeric
Queue Redis Streams or Kafka
Storage PostgreSQL (normalized) + raw backup in S3
Schema registry Custom or Confluent Schema Registry
Quality monitoring dbt tests + Prometheus metrics

Raw data is always saved to S3 before normalization. If an error in the normalization logic is discovered, it can be re-run from original data without re-scraping.

How to implement a normalization layer: step-by-step process

  1. Source analysis: identify all data sources (exchanges, blockchains, APIs), collect format samples.
  2. Schema design: create Pydantic/Zod schemas for each source with versioning.
  3. Transformation development: write normalization functions for each field (timestamp, amounts, symbols).
  4. Testing and monitoring: run on historical data, configure alerts.

What is included in the work

Our normalization layer package includes tangible deliverables:

  • Ready normalization layer for your sources (up to 7 in the basic version)
  • Detailed architecture diagram
  • Documentation of schemas and API
  • Access to the code repository with tests
  • Performance benchmarks
  • Training of your team to work with the system (up to 2 sessions)
  • Support for 1 month after launch

We also provide a performance report showing latency improvements and error reductions. Contact us to discuss your project. We guarantee a transparent process and individual approach.