Development of Market Making Algorithm from Scratch
We develop market making algorithms from scratch — from liquidity analysis to server deployment. In crypto, market making remains highly profitable for niche assets: mid-cap altcoins, perpetual futures with low liquidity. With proper tuning, average daily return is 0.15–0.3% of inventory, and fill rate reaches 70–85%. We use the Avellaneda-Stoikov model and dynamic spread to minimize inventory risk and maximize P&L.
Our guarantee: 95%+ quote uptime and resilience to inventory risks even under high volatility. 5+ years of experience and 20+ projects in this field.
How do we ensure stability?
We use a fault-tolerant architecture with hot standby, latency monitoring (<10ms), and automatic alerts when risk limits are exceeded. Each algorithm undergoes stress-testing on historical data and is customized for the specific asset. Certified engineers provide 24/7 support.Basic Market Making Model
Naive market making — place a bid X% below mid-price and an ask X% above. Problem: inventory risk. If the price moves sharply in one direction, the market maker accumulates an unfavorable position. Losses can reach 50% of capital in a single session if risks are not managed.
Avellaneda-Stoikov model — an mathematically optimal market making strategy. It accounts for inventory risk and time horizon:
bid_price = mid - δ/2 - γσ²(T-t)q ask_price = mid + δ/2 - γσ²(T-t)q where: δ = spread (optimal) γ = risk aversion coefficient σ = asset volatility q = current inventory (in asset units) T = end of trading period t = current time Key point: with positive inventory (many assets accumulated), the algorithm shifts quotes downward to sell surplus faster. With negative inventory, it shifts upward to buy.
How to configure the Avellaneda-Stoikov model?
- Collect historical data: prices, volumes, spread, volatility (σ).
- Choose risk aversion coefficient (γ) — in practice 0.01–0.1.
- Optimize target inventory (q_target) and time horizon (T).
- Run backtest on the last 30 days of data.
- Configure hard/soft limits: e.g., max inventory = 10% of capital.
We tune parameters individually using genetic algorithms and grid search.
How does the Avellaneda-Stoikov model work?
The Avellaneda-Stoikov model is a stochastic approach that dynamically adjusts quotes based on current inventory and remaining time. The risk aversion coefficient γ determines how aggressively the algorithm closes positions. In practice, we tune γ on historical data to balance spread profitability and inventory risk.
What is inventory risk and how to minimize it?
Inventory risk is the main enemy of a market maker. If the position exceeds allowed limits, we apply several methods:
- Hard limit: when inventory > MAX_INVENTORY — stop placing orders on that side. Wait for fills.
- Soft limit with skewing: gradually shift quotes against the direction of accumulated inventory. The larger the inventory, the stronger the shift.
- Hedging: open a hedge position on another exchange or in perpetual futures. If we accumulate a lot of BTC spot, we sell BTC-PERP.
For each project, we choose a combination of methods based on asset volatility and trading volume. We guarantee that drawdown from inventory risk does not exceed a predefined threshold (usually 5% of capital).
Spread Management
The spread should not be fixed — it adapts to market conditions:
- Volatility-based spread:
spread = base_spread × (current_volatility / mean_volatility). When volatility is high, the spread widens — inventory risk increases. - Order book depth: if liquidity in the order book is low, adverse selection risk is higher, spread widens.
- Time of day: during low activity periods, spread widens.
- Toxic flow: if the last N trades were predominantly on one side, it may indicate informed trading. The algorithm widens the spread or temporarily removes quotes.
Multi-Level Quotes
Instead of a single pair of orders (1 bid + 1 ask), we place multiple levels:
Bid 3: mid - 0.5% × 1000 USDT Bid 2: mid - 0.3% × 500 USDT Bid 1: mid - 0.15% × 200 USDT --- MID PRICE --- Ask 1: mid + 0.15% × 200 USDT Ask 2: mid + 0.3% × 500 USDT Ask 3: mid + 0.5% × 1000 USDT Orders close to the mid price fill more often and earn exchange rebates. Distant orders protect against sharp moves.
Order Cancellation and Re-Quoting
Orders need to be updated regularly as the mid-price changes:
- Threshold-based re-quoting: if the mid price shifts by more than N%, cancel old orders and place new ones.
- Time-based re-quoting: forced update every T seconds.
- Event-based: re-quote on every change in the best bid/ask in the order book.
Frequent order cancellations consume API request quota. Exchanges impose rate limits. For Binance: 1200 requests/min HTTP, separate limits for WebSocket. Optimizing update frequency is crucial.
Exchange Market Making Programs
Major exchanges pay for providing liquidity:
| Exchange | Program | Conditions |
|---|---|---|
| Binance | Liquidity Provider | Rebate up to -0.005% |
| Bybit | Market Maker | Zero or negative maker fee |
| OKX | Market Maker | Special fee conditions |
| Kraken | Market Maker | Maker rebate upon request |
To qualify for these conditions, you must maintain minimum quote uptime (>80% of the time bid/ask within a certain range from mid) and minimum volume.
Development Stages
| Stage | Duration | Result |
|---|---|---|
| Liquidity analysis | 2–5 days | Report with optimal model and risk parameters |
| Algorithm implementation | 2–4 weeks | Modules: pricing, order management, risk control |
| Exchange integration | 3–7 days | Stable WebSocket + REST connection |
| Testing (backtest + paper) | 1–2 weeks | Report on Sharpe, drawdown, fill rate |
| Deployment and monitoring | 3–5 days | Server with Grafana, alerts in Telegram |
Monitoring and Metrics
P&L breakdown: spread income - inventory risk losses - fees.
Fill rate: percentage of orders filled. Too low (<50%) → spread too wide. Too high (>90%) → spread too narrow, excessive adverse selection.
Inventory exposure: current position in USD, maximum per session, average. Uptime: percentage of time quotes are placed (target >99.5%). Latency: time from receiving market update to placing/updating orders (target <10ms).
Tech Stack
Language: Python (asyncio + aiohttp/websockets) for strategies with latency > 50ms. C++ or Rust for latency-critical components.
Exchange connectors: CCXT Pro (Python) provides a unified API for WebSocket. For production, we build custom connectors for each exchange.
Storage: PostgreSQL for trades, orders, positions. InfluxDB or TimescaleDB for performance metrics.
Monitoring: Grafana dashboards for real-time P&L, inventory, latency. Alerts in Telegram when risk limits are exceeded.
Contact us to evaluate your project within 2 days. Get an architect consultation — discuss the details.







