Setting Up a Docker Container for a Trading Bot
Running a trading bot directly on a VPS without isolation means dependency issues, unpredictable restarts on failure, and lack of reproducibility. API keys are stored in plain sight, package versions conflict, and recovery takes hours after a crash. We've encountered this in nearly every other project. Docker solves all of this in a couple of hours of configuration: environment isolation, automatic restarts, secure key handling, logging, and updates without downtime. We see clients saving up to 40% on VPS costs after switching to Docker. We offer turnkey deployment with guaranteed stability and resource savings.
Why Docker is Better Than Bare-Metal for a Trading Bot
Running on a bare VPS without a container means manual dependency management, risk of version conflicts, and slow recovery after a failure. Docker isolates the environment and provides reproducibility: the bot image works identically on any server. As noted in the Wikipedia article about Docker, containerization ensures isolation and reproducibility. Let's compare key metrics:
| Parameter |
Bare-metal |
Docker |
| Time to deploy a new instance |
30-60 minutes |
5-10 minutes |
| Recovery after failure |
15-30 minutes |
30 seconds (healthcheck) |
| Dependency isolation |
manual |
automatic |
| Memory usage |
~300 MB (average) |
~200 MB (slim image) |
Docker reduces recovery time from hours to seconds. This is critical for frequent redeploys. Additionally, containers consume fewer resources, lowering VPS costs — savings of $50–100 per month.
Healthcheck: How to Set Up Automatic Restart
Healthcheck is a check to see if the bot is alive. If it hangs, Docker restarts the container. We write a heartbeat file inside the container and check its freshness:
healthcheck:
test: ["CMD", "bash", "-c", "[ $(($(date +%s) - $(date +%s -r /app/data/heartbeat))) -lt 60 ]"]
interval: 30s
timeout: 10s
retries: 3
start_period: 10s
The bot writes a heartbeat every 10 seconds (threading + pathlib). If the file is not updated for longer than 60 seconds, it's time to restart. Parameters: interval 30 seconds, timeout 10 seconds, retries 3, start_period 10 seconds.
Dockerfile and docker-compose
Dockerfile for a Python Bot
FROM python:3.12-slim
RUN apt-get update && apt-get install -y --no-install-recommends build-essential && rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
RUN adduser --disabled-password --gecos '' botuser
USER botuser
CMD ["python", "-u", "bot.py"]
The -u flag disables stdout buffering — logs appear immediately.
docker-compose with Secrets and Auto-Start
services:
trading-bot:
build: .
container_name: trading_bot
restart: unless-stopped
environment:
- PYTHONUNBUFFERED=1
- TZ=UTC
env_file:
- .env.secrets
volumes:
- ./data:/app/data
- ./logs:/app/logs
deploy:
resources:
limits:
memory: 512M
cpus: '0.5'
healthcheck:
test: ["CMD", "python", "-c", "import os; os.path.exists('/app/data/heartbeat')"]
interval: 30s
timeout: 10s
retries: 3
start_period: 10s
logging:
driver: "json-file"
options:
max-size: "50m"
max-file: "5"
.env.secrets (never commit to git):
BINANCE_API_KEY=your_key_here
BINANCE_SECRET=your_secret_here
TELEGRAM_BOT_TOKEN=notification_token
Add .env.secrets to .gitignore. For higher security, we use Docker Secrets or HashiCorp Vault — this is included in our "Security" package.
Comparison of Secret Storage Methods
| Method |
Simplicity |
Security |
Recommendation |
| .env file |
High |
Low (keys on filesystem) |
For testing |
| Docker Secrets |
Medium |
High (secrets in memory only) |
Production |
| HashiCorp Vault |
Low |
Very high |
Large projects |
Process: Stages of Work
- Bot code analysis — check dependencies, failure points, memory leaks.
- Docker image design — optimize layers, minimize size.
- docker-compose configuration — healthcheck, logging, resource limits, volume mounts.
- Secret integration — connect secure storage for API keys.
- CI/CD — automated build and deployment on repository changes (GitHub Actions).
- Testing — testnet verification, stress test.
- Deployment and documentation — launch on VPS, operations and recovery instructions.
In one project, we migrated a bot from bare-metal to Docker — recovery time dropped from 20 minutes to 30 seconds, and VPS costs decreased by 40% due to memory optimization. The client got stable operation without nightly failures.
What's Included in the Work?
- Audit of existing code and dependencies
- Docker image build with layer optimization
- docker-compose setup with healthcheck, logs, and resource limits
- Integration of secure key storage (env_file or Docker Secrets)
- CI/CD creation for automated build and deploy
- Operations and recovery documentation
- 7-day support after launch
We'll assess your project for free — contact us.
Estimated Timelines
Setting up a single container: 1 to 3 days depending on bot complexity. For complex solutions with multiple services (bot, DB, queues): up to 5 days.
Update Without Full Downtime
docker-compose build trading-bot
docker-compose up -d --no-deps trading-bot
--no-deps leaves other services untouched. Downtime: 5-10 seconds. This is sufficient for most bots. If zero-downtime is required, we use a blue-green strategy (higher cost and complexity, but possible).
Our team has 10+ years of experience with Docker in production and over 40 successful projects. We guarantee 99.9% SLA for the container. If you have questions, reach out for a consultation.
Blockchain Infrastructure Deployment: Nodes, RPC, Indexing
Subgraph fell at 3:47 AM. By morning users saw outdated balances, transactions "hung" in the UI, support received 47 tickets in an hour. Cause: the handler in the subgraph failed on a transaction with a non-standard event log — and the entire index stopped. We have encountered such situations dozens of times. Our experience shows: blockchain infrastructure does not forgive gaps in observability. Guaranteeing uptime without multi-layered monitoring and fault-tolerant architecture is impossible. Over 8 years working with Ethereum, Polygon, and Solana, we have developed an approach that allows predictable deployment of infrastructure of any scale — from a single node to a multichain grid with dozens of subgraphs.
RPC Layer Architecture
Every dApp interaction with the blockchain goes through RPC — the JSON-RPC API provided by a node. Three options:
Managed providers — Alchemy, QuickNode, Infura, Ankr. Minimal operational costs, SLA, built-in monitoring. Limits: rate limits (Alchemy Free: 300 RU/sec), vendor lock, potential downtime during provider incidents. For most projects — the right choice at the start.
Self-owned nodes — full control, no rate limits, no third-party dependence. Cost: archive Ethereum node requires 2.5–3TB SSD, a strong server, and DevOps support. Sync from scratch on Ethereum via Geth/Nethermind — 3–7 days. Justified under high load or latency requirements.
Hybrid — self-owned node as primary, managed provider as fallback. Standard for protocols with high TVL. Proper load balancing can reduce costs by 20–30% compared to pure managed setup. Under high monthly request volume, hybrid saves significantly.
| Provider |
Strength |
Limitation |
| Alchemy |
Supernode, Enhanced APIs, webhooks |
Expensive on high-volume |
| QuickNode |
Low latency, multi-chain |
More expensive than Alchemy on basic plan |
| Infura |
Historical reliability |
Rate limits on free, one major incident halted half of DeFi |
| Ankr |
Cheap, 40+ chains |
Less stable |
How to Set Up an RPC Layer Without a Single Point of Failure?
At least two providers, DNS round-robin with health check every 5 seconds, automatic fallback when latency >500 ms. In practice, this gives 99.99% availability during any provider failure. For protocols with high TVL, we recommend a custom HA-proxy (nginx or Envoy) in front of two managed providers.
Why Is a Hybrid RPC Scheme More Cost-Effective Than Pure Managed?
At high request volumes, managed providers can be very expensive; a hybrid using a self-owned node as primary and a managed fallback cuts costs significantly without losing SLA.
Ethereum Node Clients
Execution clients: Geth (most used), Nethermind (C#, fast sync), Besu (Java, enterprise), Erigon (fastest sync, efficient archive mode ~2TB instead of 3TB).
Consensus clients (post-Merge): Lighthouse (Rust), Prysm (Go), Teku (Java), Nimbus (Nim). Each node after The Merge requires a pair of execution + consensus clients.
For DevOps: eth-docker — Docker Compose configurations for all client combinations. Setting up monitoring via Grafana + Prometheus is mandatory; a standard dashboard is available in each client's repository.
The Graph: Event Indexing
The Graph Protocol — decentralized indexing. A subgraph describes which events from which contracts to index and how to transform them into a GraphQL schema.
Subgraph structure:
-
subgraph.yaml — manifest: contract addresses, startBlock, events to handle
-
schema.graphql — GraphQL schema of entities
-
src/mapping.ts — AssemblyScript event handlers
dataSources:
- kind: ethereum
name: UniswapV3Pool
network: mainnet
source:
address: "0x88e6A0c2dDD26FEEb64F039a2c41296FcB3f5640"
abi: UniswapV3Pool
startBlock: 12370624
mapping:
eventHandlers:
- event: Swap(indexed address,indexed address,int256,int256,uint160,uint128,int24)
handler: handleSwap
AssemblyScript handlers — not TypeScript. No nullable types, no closures, no many standard APIs. An error in the handler stops the subgraph indexing on that transaction. Important: add try-catch for operations that can fail (e.g., store.get() for an entity that may not exist).
How to Avoid Subgraph Indexing Stops?
Graph Node logs are monitored in real-time; on hasIndexingErrors = true an alert fires and an automatic node restart (via systemd or Kubernetes). Typical downtime on error — 150–300 seconds to recover. Additionally, for production we set up a watchdog that restarts Graph Node if subgraph lag exceeds 50 blocks.
Choosing Between Hosted Service and Decentralized Network
Graph Hosted Service (free, centralized) is deprecated in favor of Subgraph Studio + Graph Network. For production: deploy on Graph Network with GRT curation signal — the subgraph gets indexers proportional to curation.
Alternatives to The Graph: Ponder (TypeScript, self-hosted, easier to debug), Envio (ultra-fast indexer, supports EVM + non-EVM), Subsquid (TypeScript, own network), Moralis Streams (managed, webhook-based). Our experience shows: for high-load projects with unique logic, Ponder or Envio are more effective — they give full control over the process and do not require GRT tokenomics.
Webhooks and Real-Time Notifications
Alchemy Webhooks and QuickNode Streams allow receiving events in real-time via HTTP webhook or WebSocket. For monitoring addresses, new transactions, mints — this is faster than polling RPC.
Tenderly — platform for monitoring and alerts. You can set up an alert for a specific contract event, balance change, function call with certain parameters. Transaction simulation via Tenderly API is invaluable for debugging.
Monitoring and Observability
Minimum monitoring stack for a protocol:
On-chain: OpenZeppelin Defender Sentinel — watches contract events, triggers webhook or Autotask when conditions are met. Forta Network — community-maintained bots detect anomalies (large withdrawals, flash loans, governance attacks).
Infrastructure: Grafana + Prometheus for nodes, Datadog or Grafana Cloud for managed metrics. Alerts on: node is 10+ blocks behind, RPC latency >500ms, subgraph lag >100 blocks.
Uptime: Better Uptime or PagerDuty on RPC endpoint and subgraph health endpoint (The Graph provides _meta { hasIndexingErrors, block { number } }).
Why Is Monitoring Without Tenderly Insufficient?
Tenderly provides transaction simulation and detailed traces — critical for debugging subgraph and smart contract errors. Forta focuses on network anomalies, not your infrastructure. The combination of Tenderly plus a custom Grafana dashboard covers 90% of incident scenarios.
Multichain Infrastructure
A protocol on 5 chains = 5 separate RPC endpoints, 5 subgraphs, 5 monitoring configs. Manageable but requires deployment automation.
For subgraph multi-network deployment: graph deploy --network mainnet, graph deploy --network arbitrum-one etc. with a unified codebase and network-specific addresses in separate config files.
Chainlink CCIP and LayerZero for cross-chain messaging require monitoring of both chains and transactions on intermediate relayers. A reorg on the source chain after a confirmed mint on the target chain is a classic bridge problem. Solution: wait for finality (on Ethereum ~15 minutes after Merge for economic finality) before confirming on the target chain.
Infrastructure Setup Process
- Audit current stack — determine chains, request volume, latency and availability requirements.
- Architecture design — select providers, load balancing, redundancy.
- Subgraph development — manifest → schema → handlers → testing on local Graph Node → deploy to testnet → mainnet.
- Monitoring configuration — Tenderly alerts, Grafana dashboard, PagerDuty integration.
- Documentation and runbook — what to do when: subgraph falls behind, RPC downtime, node desync.
- Handover to operations — team training, access transfer, first month support.
What's Included
- Deployment of managed or self-hosted Ethereum, Polygon, BNB Chain nodes
- RPC layer setup with primary/fallback and load balancing
- Subgraph development and deployment for your protocol
- Monitoring connection (Tenderly, Grafana, alerts)
- Runbook and operations documentation
- Team training (up to 4 hours online)
- 30-day support after delivery
Timeline
| Task |
Duration |
| RPC and basic monitoring setup |
1–2 weeks |
| Subgraph for one protocol |
2–4 weeks |
| Self-hosted node with monitoring |
2–3 weeks |
| Full infrastructure (multi-chain, monitoring, runbooks) |
6–10 weeks |
All projects are managed in a GitHub/GitLab repository with CI/CD; configuration code stays with you. Order infrastructure deployment — we'll show how to cut costs by 20–30% without losing reliability. Get a consultation — we'll demonstrate how we deployed infrastructure for a protocol with large TVL on Ethereum and Arbitrum. Contact us.