One day, a client of ours — an e‑commerce site with a catalog of 50,000 products — discovered that a competitor had copied their entire inventory in three hours using a simple Python script (aiohttp with a pool of 200 proxies). Server load doubled, legitimate users complained about delays. We implemented multi‑layer protection that blocks 99% of bots without slowing real customers. Load dropped by 40%, and API response time returned to its original 50ms.
According to Imperva research, bots account for over 40% of internet traffic; without protection, your API is vulnerable.
Why simple rate limiting doesn’t work
IP‑based rate limiting is easily bypassed by a proxy pool. Bots can mimic human request patterns. You need analysis of behavioral patterns, headers, and even TLS fingerprints. Only a combination of methods provides reliable protection. Behavioral analysis detects three times more bots than simple IP blocking.
How we solve it: multi‑layer protection
Each layer filters out a portion of traffic:
Client → WAF (IP reputation) → Rate Limiting → Bot Detection → API Logic
↓
Fingerprint + Behavioral Analysis + CAPTCHA
No single method is perfect — we combine them for maximum effectiveness. Get a consultation from an engineer with 10+ years of experience to evaluate your project.
What signals give a bot away?
| Signal | Weight | Description |
|---|---|---|
Missing User‑Agent or curl/python‑requests |
+40 | Typical automatic clients |
Missing Accept‑Language / Accept‑Encoding |
+20 | Browsers always send these headers |
| Requests exactly every N ms | +35 | A human cannot be that precise |
| Identical URL pattern (sequential crawl) | +30 | ?page=1, ?page=2... |
| No Referer during navigation | +15 | Browsers usually pass it |
| Multiple requests from the same IP range | +25 | Distributed bot |
| Atypical TLS fingerprint (JA3) | +30 | Node.js/Python TLS differs from browsers |
Comparison of protection methods
| Method | Advantages | Disadvantages |
|---|---|---|
| Rate limiting | Simplicity | Ineffective against IP pools |
| Behavioral analysis | High accuracy | Requires Redis, more resources |
| JA3 fingerprint | Hard to spoof | Not all proxies pass fingerprint |
| Honeypot | Zero load | Detected by mid‑level bots |
| CAPTCHA | High reliability | Hurts UX, not always applicable for APIs |
We combine behavioral analysis with JA3 and honeypot. This set gives the best balance of accuracy and performance.
Why behavioral analysis outperforms rate limiting?
Behavioral analysis considers not only frequency but also qualitative signs: missing browser headers, machine‑like timing precision, sequential URL crawling. Even if a bot uses distributed IPs, its behavior gives it away. In practice, this reduces false positives by 80% compared to pure rate limiting.
How to implement protection in 5 steps
- Audit your current API and identify critical endpoints.
- Choose the stack: Python + Redis + nginx for JA3 fingerprint.
- Implement a behavioral detector (see code below).
- Integrate honeypot and CAPTCHA (Cloudflare Turnstile).
- Monitor with Prometheus and Grafana.
We help at every stage, guaranteeing zero false positives.
How behavioral analysis works
The behavioral detector calculates a score from 0 to 100 based on several criteria. Here’s the key implementation:
import time
import statistics
from collections import defaultdict, deque
class BotDetector:
def __init__(self, redis_client):
self.r = redis_client
self.window = 300 # 5‑minute analysis window
def analyze_request(self, request) -> dict:
"""Returns score (0-100) and suspicion reasons"""
score = 0
reasons = []
# 1. Browser headers
headers = request.headers
ua = headers.get('User-Agent', '')
bot_uas = ['python-requests', 'curl', 'wget', 'Go-http-client',
'Java/', 'okhttp', 'axios', 'node-fetch']
for bot_ua in bot_uas:
if bot_ua.lower() in ua.lower():
score += 40
reasons.append(f'bot_useragent:{bot_ua}')
break
if not ua:
score += 40
reasons.append('no_useragent')
if not headers.get('Accept-Language'):
score += 20
reasons.append('no_accept_language')
if not headers.get('Accept-Encoding'):
score += 15
reasons.append('no_accept_encoding')
# 2. Request frequency analysis
ip = request.remote_addr
timing_score = self._analyze_timing(ip)
if timing_score > 0:
score += timing_score
reasons.append(f'suspicious_timing:{timing_score}')
# 3. URL pattern (sequential crawl)
path = request.path
pattern_score = self._analyze_url_pattern(ip, path)
if pattern_score > 0:
score += pattern_score
reasons.append(f'url_pattern:{pattern_score}')
# 4. JA3 TLS fingerprint (via nginx variable)
ja3 = headers.get('X-JA3-Fingerprint')
if ja3 and self._is_suspicious_ja3(ja3):
score += 30
reasons.append(f'suspicious_ja3:{ja3[:16]}')
return {
'score': min(score, 100),
'is_bot': score >= 60,
'reasons': reasons,
'action': self._get_action(score)
}
def _analyze_timing(self, ip: str) -> int:
"""Analysis of time intervals between requests"""
key = f"timing:{ip}"
now = time.time()
self.r.lpush(key, now)
self.r.ltrim(key, 0, 49)
self.r.expire(key, self.window)
timestamps = [float(t) for t in self.r.lrange(key, 0, -1)]
if len(timestamps) < 5:
return 0
timestamps.sort()
intervals = [timestamps[i+1] - timestamps[i] for i in range(len(timestamps)-1)]
if not intervals:
return 0
avg = statistics.mean(intervals)
stdev = statistics.stdev(intervals) if len(intervals) > 1 else 0
cv = stdev / avg if avg > 0 else 0
if cv < 0.05 and avg < 2.0:
return 35
if cv < 0.15 and avg < 1.0:
return 25
return 0
def _analyze_url_pattern(self, ip: str, path: str) -> int:
"""Detect sequential crawling"""
key = f"paths:{ip}"
self.r.lpush(key, path)
self.r.ltrim(key, 0, 19)
self.r.expire(key, self.window)
paths = self.r.lrange(key, 0, -1)
if len(paths) < 5:
return 0
import re
numeric_pattern = re.compile(r'/(\d+)$')
numbers = [int(m.group(1)) for p in paths if (m := numeric_pattern.search(p.decode()))]
if len(numbers) >= 5:
diffs = [numbers[i] - numbers[i+1] for i in range(len(numbers)-1)]
if all(d == diffs[0] for d in diffs) and abs(diffs[0]) in [1, -1]:
return 30
return 0
def _is_suspicious_ja3(self, ja3: str) -> bool:
SUSPICIOUS_JA3 = {
'e7d705a3286e19ea42f587b344ee6865', # Python requests
'b386946a5a44d1ddcc843bc75336dfce', # Scrapy
'6734f37431670b3ab4292b8f60f29984', # Go default
}
return ja3.lower() in SUSPICIOUS_JA3
def _get_action(self, score: int) -> str:
if score < 30: return 'allow'
if score < 60: return 'challenge'
if score < 80: return 'throttle'
return 'block'
Additional measures against CAPTCHA bypass
CAPTCHA bypass is rare, but we add honeypot and JA3 for extra strength. A honeypot is a hidden endpoint visited only by bots. We add /api/items/all (which doesn't exist for real users) and blacklist the IP on access. CAPTCHA is triggered only at score >= 60, so normal users aren't bothered. Integration with Cloudflare Turnstile is transparent for real customers.
@app.route('/api/search')
def search():
if g.get('bot_score', 0) >= 60:
token = request.headers.get('CF-Turnstile-Token')
if not verify_turnstile(token):
return jsonify({'error': 'CAPTCHA required', 'captcha_site_key': TURNSTILE_SITE_KEY}), 429
q = request.args.get('q', '')
return jsonify(search_service.query(q))
Example honeypot endpoint:
@app.route('/api/items/all')
def honeypot_endpoint():
ip = request.remote_addr
bot_detector.blacklist_ip(ip, duration=86400, reason='honeypot')
return jsonify({'items': [], 'total': 0})
What’s included in the work
- Audit of your current API and vulnerabilities
- Implementation of the behavioral detector (BotDetector)
- Integration with Redis and middleware (Flask/Django/Express)
- nginx configuration to pass JA3 fingerprint
- Creation of honeypot endpoints and CAPTCHA system
- Metrics dashboard (Prometheus + Grafana) for monitoring
- Documentation and team training
- 2 weeks of technical support after launch
Estimated timeline
Implementation of multi‑layer protection takes from 5 to 10 working days, depending on API complexity. The cost is determined individually — we’ll evaluate your project after analysis. Clients save significant resources after implementation; the investment typically pays for itself in 2–3 months.
Order an API protection audit today. Get a consultation from an engineer with 10+ years of experience in web development and high‑load systems.
Learn more about Web scraping.







