A trading bot that crashes at 3 AM isn't just about missed trades. It's about open positions without management, missed stop-losses, and losses of tens of percent of capital. Imagine: your bot earns 2% per day, and suddenly the server goes down for 2 hours. Loss—4% of the deposit. For a $100k account, that's $4k that could have been saved with a simple healthcheck. Every hour of downtime can be costly—setting up monitoring pays off in one losing trade that you prevented. Uptime monitoring isn't about pretty Grafana dashboards; it's a system that wakes you up before the market does it more painfully. We configure monitoring in 1–2 days, and it runs for years without failures. With over 50 monitoring automation projects under our belt, we have the experience to anticipate typical mistakes.
What to Monitor
Bot uptime is not just "the process is running." The process can be alive while the bot isn't trading. Three levels of checks:
- Process alive — the process is running, not hung.
- Application alive — the bot processes data and regularly writes the timestamp of the last activity. If the timestamp hasn't been updated for N minutes—something is wrong.
- Trading alive — the bot is not just working but actually trading: number of orders over a period, P&L, open positions match the strategy.
How to Set Up a Healthcheck Endpoint?
The simplest and most reliable approach is to add an HTTP endpoint directly to the bot. We use FastAPI:
from fastapi import FastAPI import asyncio import time app = FastAPI() last_heartbeat = time.time() bot_state = {"status": "running", "last_trade": None, "open_positions": 0} @app.get("/health") async def health(): age = time.time() - last_heartbeat if age > 60: # not updated for more than a minute return {"status": "stale", "heartbeat_age_seconds": age}, 503 return {"status": "ok", **bot_state} # In the main bot loop async def bot_loop(): global last_heartbeat while True: last_heartbeat = time.time() await run_strategy() await asyncio.sleep(5) The endpoint returns status 200 when healthy and 503 when the heartbeat is overdue. External monitoring catches the 503 and sends an alert.
Comparison of External Monitoring Tools
Uptime Kuma deploys 300 times faster than Prometheus and requires 10 times less server resources. For a single bot, it's the optimal choice.
| Tool | Monitoring Type | Deployment Time | Alerts | Reliability |
|---|---|---|---|---|
| Uptime Kuma | Self-hosted (Docker) | 5 minutes | Telegram, Discord, email | High (self-hosted) |
| Better Uptime | SaaS | 10 minutes | Slack, PagerDuty, SMS | High (SLA 99.9%) |
| Prometheus + Grafana | Self-hosted | 2–3 hours | Alertmanager, Telegram | Very high, but more complex |
Uptime Kuma is a self-hosted alternative to UptimeRobot. It checks the HTTP endpoint every N seconds and sends notifications when it's unavailable. Deploy in 5 minutes with Docker:
docker run -d --restart=always -p 3001:3001 \ -v uptime-kuma:/app/data louislam/uptime-kuma:1 For the bot: Monitor Type = HTTP, URL = http://your-bot-host:8080/health, interval = 30 seconds, expected status = 200.
Better Uptime / PagerDuty — if you need SLA guarantees and escalation policies. We will choose the option that fits your budget.
Why You Need a Watchdog
If the bot itself can't send an alert (process dead), you need an external watchdog. The simplest version is a bash script with cron:
#!/bin/bash # /usr/local/bin/bot-watchdog.sh HEALTH_URL="http://localhost:8080/health" TELEGRAM_TOKEN="..." CHAT_ID="..." response=$(curl -s -o /dev/null -w "%{http_code}" --max-time 10 "$HEALTH_URL") if [ "$response" != "200" ]; then curl -s -X POST "https://api.telegram.org/bot${TELEGRAM_TOKEN}/sendMessage" \ -d "chat_id=${CHAT_ID}" \ -d "text=ALERT: Trading bot health check failed (HTTP ${response})" fi How to Determine the Optimal Heartbeat Interval?
The interval depends on market volatility and reaction time. For high-frequency trading—10–30 seconds, for regular strategies—30–60 seconds. The main rule: the interval should be less than the time it takes for a missed trade to become critical. Also consider network latency and healthcheck processing time.
Common Mistakes and Their Solutions
One frequent mistake is a too-heavy healthcheck endpoint that causes timeouts and false alarms. Solution: make the endpoint as lightweight as possible, only checking for a heartbeat without deep logic. Another mistake is too frequent checks (every 5 seconds), creating load and noise. Optimal interval is 30–60 seconds. A third is the lack of crash loop protection, where the bot restarts infinitely. Use StartLimitBurst=3 in systemd or Docker restart policies. If you don't want to deal with this yourself, order a ready-made solution—we'll set up monitoring in one day.
Step-by-Step Monitoring Setup in 1 Day
- Add a healthcheck endpoint to the bot code (example above).
- Deploy Uptime Kuma on the server (
docker run). - Configure monitoring: URL =
http://your-bot:8080/health, interval = 30s. - Connect Telegram alert (BotFather + your chat_id).
- Install the watchdog script in cron (every minute).
- Configure systemd with
Restart=on-failureandStartLimitBurst=3. - Test: stop the bot—within 30 seconds an alert should arrive.
Automatic Restart via systemd
If the bot runs as a systemd service, specify:
[Unit] Description=Trading Bot After=network.target [Service] ExecStart=/usr/bin/python3 /opt/bot/main.py Restart=on-failure RestartSec=10 StartLimitIntervalSec=60 StartLimitBurst=3 [Install] WantedBy=multi-user.target Restart=on-failure — automatic restart on crash. StartLimitBurst=3 — no more than 3 restarts in 60 seconds (crash loop protection).
What's Included in the Setup?
We offer a comprehensive monitoring setup in 1–2 business days:
- adding a healthcheck endpoint to the bot (or adapting an existing one)
- deploying Uptime Kuma / configuring external monitoring
- setting up Telegram alerts
- watchdog script
- automatic restart via systemd/Docker
- basic Prometheus metrics, if analytics on trading activity is needed
If you don't have time for self-configuration, leave a request—we'll evaluate your project and offer a turnkey solution. We have been automating monitoring for over 5 years—dozens of configured bots that run without failures. Order reliable monitoring today and get a consultation before work begins.







