We develop tick-data processing pipelines — recording every trade with price, volume, and side. Standard OHLCV candles lose market microstructure: liquidity imbalance, large trades, buy/sell flow. Without a quality pipeline, an ML model trains on noise. For example, in one project for Binance (from our practice), the load reached 300,000 ticks per second — ClickHouse handled it, while PostgreSQL crashed at 10,000. Our over five years of experience guarantees reliability. Contact us — we are ready to design and implement a pipeline for your tasks.
Why Tick Data Matters More Than OHLCV for ML
When aggregating into 1-minute candles, up to 80% of information is lost: you do not see how trades are distributed within the interval, whether there was a volume spike, or who was the aggressor. Volume bars, dollar bars, and imbalance bars preserve these signals. ML models trained on ticks show 15–20% higher accuracy in price direction prediction tasks.
Problems We Solve
- High loads. Exchanges generate up to 500,000 ticks per second. Standard databases cannot handle such insertion rates.
- Latency. For HFT strategies, the delay from tick receipt to signal must not exceed 10 ms.
- Storage. Tick data for a year amounts to tens of terabytes. Partitioning, TTL, and efficient compression are necessary. ClickHouse infrastructure savings can reach 50% compared to traditional relational databases.
- Bar diversity. Time bars are uneven during low activity periods. Volume/dollar/imbalance bars adapt to market activity.
How We Do It: Stack and Case Study
In one project for Binance (from our practice), we built a pipeline that collects aggregated trades via WebSocket, buffers them in memory, and asynchronously inserts into ClickHouse.
import asyncio
import websockets
import json
from datetime import datetime
import asyncpg
class TickDataCollector:
def __init__(self, symbol, db_pool):
self.symbol = symbol
self.db_pool = db_pool
self.buffer = []
self.buffer_size = 1000
async def connect_binance_trades(self):
url = f"wss://stream.binance.com:9443/ws/{self.symbol.lower()}@aggTrade"
async with websockets.connect(url, ping_interval=20) as ws:
async for msg in ws:
trade = json.loads(msg)
tick = {
'symbol': self.symbol,
'timestamp': datetime.fromtimestamp(trade['T'] / 1000),
'price': float(trade['p']),
'quantity': float(trade['q']),
'is_buyer_maker': trade['m'],
'trade_id': trade['a']
}
self.buffer.append(tick)
if len(self.buffer) >= self.buffer_size:
await self.flush_to_db()
async def flush_to_db(self):
async with self.db_pool.acquire() as conn:
await conn.executemany(
"""INSERT INTO trades (symbol, timestamp, price, quantity, is_buyer_maker, trade_id)
VALUES ($1, $2, $3, $4, $5, $6)""",
[(t['symbol'], t['timestamp'], t['price'], t['quantity'],
t['is_buyer_maker'], t['trade_id']) for t in self.buffer]
)
self.buffer.clear()
Storage is organized in ClickHouse with the MergeTree engine, daily partitioning, and a TTL of 365 days. This provides efficient compression (10x compared to CSV) and high insertion speed.
CREATE TABLE trades (
timestamp DateTime64(3),
symbol LowCardinality(String),
price Float64,
quantity Float32,
is_buyer_maker UInt8,
trade_id UInt64
) ENGINE = MergeTree()
PARTITION BY toYYYYMMDD(timestamp)
ORDER BY (symbol, timestamp)
TTL timestamp + INTERVAL 365 DAY
SETTINGS index_granularity = 8192;
ClickHouse inserts 500K+ rows/sec — 50 times faster than PostgreSQL for such loads. Monthly aggregations complete in seconds. We guarantee your pipeline will handle any market activity.
How to Build Volume Bars from Ticks: Step by Step
- Connect to the exchange WebSocket to receive aggregated trades.
- Accumulate ticks in a buffer (e.g., 1000 records).
- When the specified volume is reached, close the bar and save it to ClickHouse.
- Use the
create_volume_barsfunction from the example below.
Volume bars close when a given volume accumulates, not at a fixed time interval. This yields a uniform number of observations regardless of market activity.
def create_volume_bars(ticks_df, bar_volume=10):
"""Each bar = bar_volume units of the asset"""
bars = []
current_bar = {'open': None, 'high': -np.inf, 'low': np.inf,
'close': None, 'volume': 0, 'start_time': None}
for _, tick in ticks_df.iterrows():
if current_bar['open'] is None:
current_bar['open'] = tick['price']
current_bar['start_time'] = tick['timestamp']
current_bar['high'] = max(current_bar['high'], tick['price'])
current_bar['low'] = min(current_bar['low'], tick['price'])
current_bar['close'] = tick['price']
current_bar['volume'] += tick['quantity']
if current_bar['volume'] >= bar_volume:
bars.append(current_bar.copy())
current_bar = {'open': None, 'high': -np.inf, 'low': np.inf,
'close': None, 'volume': 0, 'start_time': None}
return pd.DataFrame(bars)
Similarly, dollar bars (by USD volume) and imbalance bars (by buy/sell imbalance) are constructed.
| Bar Type | Closing Criterion | When to Use |
|---|---|---|
| Time | Time interval | High liquidity, uniform activity |
| Volume | Accumulated volume | Adapt to volatility spikes |
| Dollar | Accumulated USD volume | Invariant to asset price |
| Imbalance | Buy/sell imbalance | Find reversal points |
What Benefits Does Tick Feature Engineering Provide?
Features extracted from ticks improve ML model quality: flow imbalance, trade frequency, VWAP deviation, large trade ratio. In a real-time streaming ML pipeline, these features are calculated on sliding windows.
def create_tick_features(ticks_df, window_ticks=[50, 200, 1000]):
features = []
for i in range(max(window_ticks), len(ticks_df)):
row_features = {}
for window in window_ticks:
window_data = ticks_df.iloc[i-window:i]
buy_vol = window_data[~window_data['is_buyer_maker']]['quantity'].sum()
sell_vol = window_data[window_data['is_buyer_maker']]['quantity'].sum()
row_features[f'flow_imbalance_{window}'] = (
(buy_vol - sell_vol) / (buy_vol + sell_vol + 1e-8)
)
row_features[f'trade_frequency_{window}'] = (
window / (window_data['timestamp'].max() -
window_data['timestamp'].min()).total_seconds() + 1e-8
)
row_features[f'avg_trade_size_{window}'] = window_data['quantity'].mean()
row_features[f'large_trade_ratio_{window}'] = (
(window_data['quantity'] > window_data['quantity'].quantile(0.9)).mean()
)
vwap = (window_data['price'] * window_data['quantity']).sum() / window_data['quantity'].sum()
row_features[f'vwap_deviation_{window}'] = (
ticks_df.iloc[i]['price'] - vwap
) / vwap
features.append(row_features)
return pd.DataFrame(features)
Large trades (above the 99th percentile) often indicate institutional activity. Analyzing their direction provides an additional signal.
How to Ensure Latency <10 ms?
Real streaming architecture:
Binance WebSocket → asyncio consumer → buffer → ClickHouse batch insert
→ Redis sorted set (last 10k ticks)
→ Feature calculator (sliding window)
→ ML inference
→ Signal output
Latency from tick to signal is under 10 ms. Achieved through asynchronous I/O, Redis buffering, and precomputed features on time windows. Savings on ClickHouse cluster compared to traditional databases can reach 50%.
According to ClickHouse documentation, insertion speed reaches 500,000 rows per second ClickHouse Documentation.
Process of Work
| Stage | Duration | Outcome |
|---|---|---|
| Analytics | 1–2 days | Requirements document and data schema |
| Design | 2–3 days | Stack selection, DB schema design, bar type definition |
| Implementation | 1–2 weeks | Collector, aggregators, feature engineering, ML pipeline integration |
| Testing | 3–5 days | Validation on historical data, speed stress test |
| Deployment | 2–3 days | Deployment in your cluster (Docker/K8s), monitoring |
Timeline and Deliverables
Base version (single symbol, ClickHouse, Redis) — from 2 weeks. Full pipeline with volume/dollar/imbalance bars, feature engineering, and real-time inference — from 4 weeks. Project cost varies; exact pricing is determined after analysis. Investment in a quality pipeline pays off through improved ML model accuracy for trading.
What is included:
- Architecture documentation.
- Pipeline source code with comments.
- ClickHouse, Redis, and queue setup.
- Integration with your ML infrastructure.
- Team training (2–3 calls).
- 2 months of support after deployment.
Checklist for Pipeline Verification
- Check insertion speed: ClickHouse must insert at least 100K rows/sec on a single core.
- Ensure TTL is set — without it, the disk fills up within a month.
- Set up latency monitoring for each stage.
- Test automatic WebSocket reconnection on disconnect.
- Validate aggregations on historical data — compare with reference bars.
Common Mistakes
- Too fine partitioning (by hour): large number of partitions degrades ClickHouse performance. Optimal is by day.
- Ignoring TTL: without automatic data cleanup, the disk fills up within a month.
- Using time bars for low-liquidity assets: most candles will be empty.
We have been developing turnkey tick-data pipelines for over five years, implementing 30+ projects for crypto trading. Get a free analysis of your data and recommendations for pipeline optimization. Contact us — we will evaluate your project and propose the best solution.







