We develop tick-data processing pipelines — recording every trade with price, volume, and side. Standard OHLCV candles lose market microstructure: liquidity imbalance, large trades, buy/sell flow. Without a quality pipeline, an ML model trains on noise. For example, in one project for Binance (from our practice), the load reached 300,000 ticks per second — ClickHouse handled it, while PostgreSQL crashed at 10,000. Our over five years of experience guarantees reliability. Contact us — we are ready to design and implement a pipeline for your tasks.
Why Tick Data Matters More Than OHLCV for ML
When aggregating into 1-minute candles, up to 80% of information is lost: you do not see how trades are distributed within the interval, whether there was a volume spike, or who was the aggressor. Volume bars, dollar bars, and imbalance bars preserve these signals. ML models trained on ticks show 15–20% higher accuracy in price direction prediction tasks.
Problems We Solve
- High loads. Exchanges generate up to 500,000 ticks per second. Standard databases cannot handle such insertion rates.
- Latency. For HFT strategies, the delay from tick receipt to signal must not exceed 10 ms.
- Storage. Tick data for a year amounts to tens of terabytes. Partitioning, TTL, and efficient compression are necessary. ClickHouse infrastructure savings can reach 50% compared to traditional relational databases.
- Bar diversity. Time bars are uneven during low activity periods. Volume/dollar/imbalance bars adapt to market activity.
How We Do It: Stack and Case Study
In one project for Binance (from our practice), we built a pipeline that collects aggregated trades via WebSocket, buffers them in memory, and asynchronously inserts into ClickHouse.
import asyncio import websockets import json from datetime import datetime import asyncpg class TickDataCollector: def __init__(self, symbol, db_pool): self.symbol = symbol self.db_pool = db_pool self.buffer = [] self.buffer_size = 1000 async def connect_binance_trades(self): url = f"wss://stream.binance.com:9443/ws/{self.symbol.lower()}@aggTrade" async with websockets.connect(url, ping_interval=20) as ws: async for msg in ws: trade = json.loads(msg) tick = { 'symbol': self.symbol, 'timestamp': datetime.fromtimestamp(trade['T'] / 1000), 'price': float(trade['p']), 'quantity': float(trade['q']), 'is_buyer_maker': trade['m'], 'trade_id': trade['a'] } self.buffer.append(tick) if len(self.buffer) >= self.buffer_size: await self.flush_to_db() async def flush_to_db(self): async with self.db_pool.acquire() as conn: await conn.executemany( """INSERT INTO trades (symbol, timestamp, price, quantity, is_buyer_maker, trade_id) VALUES ($1, $2, $3, $4, $5, $6)""", [(t['symbol'], t['timestamp'], t['price'], t['quantity'], t['is_buyer_maker'], t['trade_id']) for t in self.buffer] ) self.buffer.clear() Storage is organized in ClickHouse with the MergeTree engine, daily partitioning, and a TTL of 365 days. This provides efficient compression (10x compared to CSV) and high insertion speed.
CREATE TABLE trades ( timestamp DateTime64(3), symbol LowCardinality(String), price Float64, quantity Float32, is_buyer_maker UInt8, trade_id UInt64 ) ENGINE = MergeTree() PARTITION BY toYYYYMMDD(timestamp) ORDER BY (symbol, timestamp) TTL timestamp + INTERVAL 365 DAY SETTINGS index_granularity = 8192; ClickHouse inserts 500K+ rows/sec — 50 times faster than PostgreSQL for such loads. Monthly aggregations complete in seconds. We guarantee your pipeline will handle any market activity.
How to Build Volume Bars from Ticks: Step by Step
- Connect to the exchange WebSocket to receive aggregated trades.
- Accumulate ticks in a buffer (e.g., 1000 records).
- When the specified volume is reached, close the bar and save it to ClickHouse.
- Use the
create_volume_barsfunction from the example below.
Volume bars close when a given volume accumulates, not at a fixed time interval. This yields a uniform number of observations regardless of market activity.
def create_volume_bars(ticks_df, bar_volume=10): """Each bar = bar_volume units of the asset""" bars = [] current_bar = {'open': None, 'high': -np.inf, 'low': np.inf, 'close': None, 'volume': 0, 'start_time': None} for _, tick in ticks_df.iterrows(): if current_bar['open'] is None: current_bar['open'] = tick['price'] current_bar['start_time'] = tick['timestamp'] current_bar['high'] = max(current_bar['high'], tick['price']) current_bar['low'] = min(current_bar['low'], tick['price']) current_bar['close'] = tick['price'] current_bar['volume'] += tick['quantity'] if current_bar['volume'] >= bar_volume: bars.append(current_bar.copy()) current_bar = {'open': None, 'high': -np.inf, 'low': np.inf, 'close': None, 'volume': 0, 'start_time': None} return pd.DataFrame(bars) Similarly, dollar bars (by USD volume) and imbalance bars (by buy/sell imbalance) are constructed.
| Bar Type | Closing Criterion | When to Use |
|---|---|---|
| Time | Time interval | High liquidity, uniform activity |
| Volume | Accumulated volume | Adapt to volatility spikes |
| Dollar | Accumulated USD volume | Invariant to asset price |
| Imbalance | Buy/sell imbalance | Find reversal points |
What Benefits Does Tick Feature Engineering Provide?
Features extracted from ticks improve ML model quality: flow imbalance, trade frequency, VWAP deviation, large trade ratio. In a real-time streaming ML pipeline, these features are calculated on sliding windows.
def create_tick_features(ticks_df, window_ticks=[50, 200, 1000]): features = [] for i in range(max(window_ticks), len(ticks_df)): row_features = {} for window in window_ticks: window_data = ticks_df.iloc[i-window:i] buy_vol = window_data[~window_data['is_buyer_maker']]['quantity'].sum() sell_vol = window_data[window_data['is_buyer_maker']]['quantity'].sum() row_features[f'flow_imbalance_{window}'] = ( (buy_vol - sell_vol) / (buy_vol + sell_vol + 1e-8) ) row_features[f'trade_frequency_{window}'] = ( window / (window_data['timestamp'].max() - window_data['timestamp'].min()).total_seconds() + 1e-8 ) row_features[f'avg_trade_size_{window}'] = window_data['quantity'].mean() row_features[f'large_trade_ratio_{window}'] = ( (window_data['quantity'] > window_data['quantity'].quantile(0.9)).mean() ) vwap = (window_data['price'] * window_data['quantity']).sum() / window_data['quantity'].sum() row_features[f'vwap_deviation_{window}'] = ( ticks_df.iloc[i]['price'] - vwap ) / vwap features.append(row_features) return pd.DataFrame(features) Large trades (above the 99th percentile) often indicate institutional activity. Analyzing their direction provides an additional signal.
How to Ensure Latency <10 ms?
Real streaming architecture:
Binance WebSocket → asyncio consumer → buffer → ClickHouse batch insert → Redis sorted set (last 10k ticks) → Feature calculator (sliding window) → ML inference → Signal output Latency from tick to signal is under 10 ms. Achieved through asynchronous I/O, Redis buffering, and precomputed features on time windows. Savings on ClickHouse cluster compared to traditional databases can reach 50%.
According to ClickHouse documentation, insertion speed reaches 500,000 rows per second ClickHouse Documentation.
Process of Work
| Stage | Duration | Outcome |
|---|---|---|
| Analytics | 1–2 days | Requirements document and data schema |
| Design | 2–3 days | Stack selection, DB schema design, bar type definition |
| Implementation | 1–2 weeks | Collector, aggregators, feature engineering, ML pipeline integration |
| Testing | 3–5 days | Validation on historical data, speed stress test |
| Deployment | 2–3 days | Deployment in your cluster (Docker/K8s), monitoring |
Timeline and Deliverables
Base version (single symbol, ClickHouse, Redis) — from 2 weeks. Full pipeline with volume/dollar/imbalance bars, feature engineering, and real-time inference — from 4 weeks. Project cost varies; exact pricing is determined after analysis. Investment in a quality pipeline pays off through improved ML model accuracy for trading.
What is included:
- Architecture documentation.
- Pipeline source code with comments.
- ClickHouse, Redis, and queue setup.
- Integration with your ML infrastructure.
- Team training (2–3 calls).
- 2 months of support after deployment.
Checklist for Pipeline Verification
- Check insertion speed: ClickHouse must insert at least 100K rows/sec on a single core.
- Ensure TTL is set — without it, the disk fills up within a month.
- Set up latency monitoring for each stage.
- Test automatic WebSocket reconnection on disconnect.
- Validate aggregations on historical data — compare with reference bars.
Common Mistakes
- Too fine partitioning (by hour): large number of partitions degrades ClickHouse performance. Optimal is by day.
- Ignoring TTL: without automatic data cleanup, the disk fills up within a month.
- Using time bars for low-liquidity assets: most candles will be empty.
We have been developing turnkey tick-data pipelines for over five years, implementing 30+ projects for crypto trading. Get a free analysis of your data and recommendations for pipeline optimization. Contact us — we will evaluate your project and propose the best solution.







