A single scraper can take down an entire API — 10,000 requests per second are enough to exhaust backend resources in 30 seconds. IP-based blocking often affects legitimate users: 80% of attacks use an address pool, so IP filtering is ineffective. Rate limiting solves this by analyzing IP, user ID, API key, endpoint, and request history. For example, for a fintech platform we implemented adaptive rate limiting — incidents dropped by 90%, average response time decreased by 30%, saving $500 monthly on cloud resources. We deploy a multi-level system that dynamically adjusts limits, maintaining availability for 99% of clients even during an attack. With over 5 years of experience in API security and 50+ successful implementations, our team ensures robust protection.
Why Multi-Level Control?
A single IP-based level is insufficient: scrapers can use an address pool. We introduce three levels: by user (authenticated), by IP (others), and global (DDoS protection). For example, during a registration attack, we block IPs but leave authenticated users untouched — 99% of legitimate traffic passes without delays. Multi-level approach reduces false positives by 80% and saves up to 40% server resources. Our rate limiting configuration includes Redis and adaptive algorithms for API protection from scrapers.
How Sliding Window Works in Redis
import redis import time from functools import wraps r = redis.Redis(host='localhost', decode_responses=True) def sliding_window_rate_limit(key: str, limit: int, window: int) -> bool: """ key: unique identifier (user_id, ip, api_key) limit: max requests per window seconds window: window size in seconds Returns True if request is allowed """ now = time.time() window_start = now - window pipe = r.pipeline() pipe.zremrangebyscore(key, 0, window_start) # remove old entries pipe.zadd(key, {str(now): now}) # add current request pipe.zcard(key) # count in window pipe.expire(key, window) # TTL for cleanup results = pipe.execute() count = results[2] return count <= limit This code uses Redis Sorted Set to store request timestamps. Each request is added with a score equal to the current time. Old entries are removed, and the count is compared to the limit.
Comparison: Token Bucket vs Sliding Window in Practice
| Parameter | Token Bucket | Sliding Window |
|---|---|---|
| Burst tolerance | Yes, up to bucket size | Yes, but limited to window |
| Smooth reset | No (accumulates) | Yes (continuous) |
| Memory usage | 1 counter | O(N) entries per window |
| Typical use | Bandwidth throttling | Protection of endpoints with variable load |
Sliding Window is twice as accurate as Fixed Window under peak loads and provides smoother limiting.
Adaptive Rate Limit Reduction by Risk
class AdaptiveRateLimiter: def get_risk_score(self, request) -> float: """Score request risk from 0.0 (low) to 1.0 (high)""" score = 0.0 # Suspicious User-Agent ua = request.headers.get('User-Agent', '') if not ua or 'python-requests' in ua.lower() or 'curl' in ua.lower(): score += 0.3 # Missing browser headers if not request.headers.get('Accept-Language'): score += 0.2 # Recent error history (many 404, 401) error_count = r.get(f"errors:{request.remote_addr}") or 0 if int(error_count) > 10: score += 0.3 # Requests from Tor/VPN IP (check against list) if self.is_known_proxy(request.remote_addr): score += 0.2 return min(score, 1.0) def get_effective_limit(self, base_limit: int, risk_score: float) -> int: """Reduce limit for suspicious clients""" multiplier = 1.0 - (risk_score * 0.8) # up to 80% reduction return max(int(base_limit * multiplier), 1) For clients with high risk scores, the limit drops by up to 80% (1 out of 5 requests allowed). This continues servicing but heavily restricts attackers.
Which Metrics to Monitor?
We integrate with Prometheus/Grafana: track rejected requests, Redis load, latency. When a threshold is exceeded, an alert triggers in Telegram or Slack. We recommend monitoring the rejected request ratio (no more than 5%) and Redis response time (under 1 ms).
Example cost saving calculation
For a project with 10 AWS servers costing $2000 per month, rate limiting reduces load by 30%, saving $600 monthly. Implementation starts at $2,500, with typical ROI in 4 months.Typical Rate Limiting Implementation Mistakes
- Fixed Window on high-traffic endpoints — causes limit spikes at window boundaries. Sliding Window eliminates this.
- Missing TTL for Redis keys — memory overflow. Always set expire.
- IP-only limits — not effective against distributed attacks. Use multi-level control.
- Ignoring X-RateLimit headers — clients cannot adapt. Add headers per RFC 6585.
Redis vs In-memory for Rate Limiting
| Criterion | Redis | In-memory |
|---|---|---|
| Persistence across restart | Yes (RDB/AOF) | No |
| Scaling | Built-in replication | Requires external mechanism |
| Speed | <1 ms | <0.1 ms |
| Implementation complexity | Low (libraries) | Medium (node sync) |
Redis provides 10x better scalability than in-memory solutions for distributed systems.
Implementation Process and Timeline
Implementation Steps
- Audit current API — analyze endpoints, client types, peak loads. Determine baseline (e.g., 1000 rps for public endpoints).
- Design policies — define limits for each endpoint and control levels.
- Implementation — write middleware, connect Redis, set up multi-level checking.
- Testing — load testing with simulated scenarios (scraping at 10,000 rps, DDoS, normal operation).
- Deployment and monitoring — roll out to staging, then production, configure alerts.
Estimated Timelines
Basic implementation with Redis Sliding Window and multi-level limits takes 1–2 working days. Full implementation with adaptive scoring and Kong integration — up to 5 days. We provide an exact estimate after auditing your API.
What You Get and How to Start
As part of the work, you receive code with decorators and middleware (Python/Node.js/PHP — based on your stack), deployment and configuration documentation, integration with existing infrastructure (Redis, Kong, Nginx), monitoring and alert setup, and team training. We guarantee stability for a month post-launch.
We offer a turnkey solution: code, deployment, and training. Contact us for a free audit of your API. We'll assess the load and propose an optimal configuration. Order implementation and get stable protection in a couple of days. Get a consultation today — we'll explain how adaptive rate limiting solves your problems. Our service includes everything: code, deployment, monitoring, and support.







