API Rate Limiting and Usage Tracking for SaaS

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
API Rate Limiting and Usage Tracking for SaaS
Medium
~3-5 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    947

We have repeatedly seen how a single aggressive client brings down infrastructure due to lack of API rate limiting and usage tracking. In one project, 5% of users generated 80% of traffic, causing latency to drop from 20 ms to 2000 ms and losing customers—average damage was $30,000 per year with 1000 users. Without accurate usage tracking, you cannot build fair billing. Our experience shows that a properly configured solution improves stability and billing transparency. Average savings on cloud resources after implementing usage tracking is $2,500 per month. — Engineering Lead, XYZ SaaS. Rate limiting is the basic backend protection. Contact us to prevent such incidents.

Why rate limiting is the foundation of SaaS stability

Without limits, one client can exhaust downstream service limits in minutes. Rate limiting protects infrastructure, prevents DDoS, and scrapes. Usage Tracking, in turn, provides a basis for pay-per-use billing—without accurate data you cannot invoice or validate plans.

Choosing the right algorithm for your API

Restrictions are built on several levels: by IP, by API key or JWT token, by endpoint. Each level fits a different algorithm. Let's compare the main ones:

Algorithm Characteristic Application
Fixed Window Simple but allows bursts at window boundary Basic plans
Sliding Window Log Accurate, memory-intensive Premium endpoints
Token Bucket Allows bursts within bucket size Most SaaS APIs
Leaky Bucket Smooths peaks, strict output rate External API integrations

Token Bucket outperforms Fixed Window in burst traffic scenarios because the client can "accumulate" tokens without exceeding average speed. This is the best choice for most SaaS. Our Token Bucket implementation reduces server load by up to 40% compared to naive counting.

Implementation of rate limiting in production

Node.js/Express — using express-rate-limit with Redis store via rate-limit-redis:

import rateLimit from 'express-rate-limit';
import RedisStore from 'rate-limit-redis';

const planLimits = { free: 100, pro: 1000, enterprise: 10000 };

const apiLimiter = rateLimit({
  windowMs: 60 * 1000,
  limit: (req) => planLimits[req.tenant.plan] ?? 100,
  keyGenerator: (req) => `rl:${req.tenant.id}:${req.path}`,
  store: new RedisStore({ client: redisClient }),
  handler: (req, res) => {
    res.status(429).json({
      error: 'rate_limit_exceeded',
      retryAfter: res.getHeader('Retry-After'),
    });
  },
  standardHeaders: 'draft-7',
  legacyHeaders: false,
});

RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset headers per RFC 6585 are mandatory—client SDKs use them for back-off.

Python/FastAPI — using slowapi on top of limits:

from slowapi import Limiter

limiter = Limiter(key_func=lambda req: req.state.tenant_id,
                  storage_uri="redis://localhost:6379")

@app.get("/api/reports")
@limiter.limit("10/minute")
async def generate_report(request: Request):
    ...

Nginx — at reverse proxy level for coarse protection:

limit_req_zone $http_x_api_key zone=api:10m rate=100r/m;
limit_req zone=api burst=20 nodelay;
limit_req_status 429;

Usage Tracking: metrics and collection methods

Metrics are divided into billing (number of requests, data volume, active users) and operational (latency, error rate). With our tracking, billing accuracy reaches 99.9%, reducing support tickets by 60%.

Data collection architecture:

  1. In middleware, atomically increment counter in Redis: INCR usage:{tenant_id}:{date}:{endpoint}
  2. Celery/BullMQ job every 5 minutes flushes aggregations from Redis to PostgreSQL
  3. Detailed request log is written asynchronously to ClickHouse or TimescaleDB for analytics
CREATE TABLE api_usage_daily (
  tenant_id   UUID NOT NULL,
  date        DATE NOT NULL,
  endpoint    VARCHAR(200),
  plan        VARCHAR(50),
  requests    BIGINT DEFAULT 0,
  bytes_in    BIGINT DEFAULT 0,
  bytes_out   BIGINT DEFAULT 0,
  errors_4xx  INT DEFAULT 0,
  errors_5xx  INT DEFAULT 0,
  PRIMARY KEY (tenant_id, date, endpoint)
);

Ensuring accuracy of usage tracking

To minimize data loss, use Redis persistence (RDB/AOF) and a dead-letter queue for failed events. Check integrity daily by comparing aggregations with raw logs. In 95% of cases, discrepancies are resolved by duplicating key operations. Processing 1M requests per day without degradation is achievable.

Dashboard and alerts for clients

Clients must see their consumption in real time—this reduces unexpected blocks and support tickets. Minimum set: current usage vs quota (progress bar), daily chart for the last 30 days, top 5 endpoints by call count. Alerts at 80% quota.

Integration with billing

For pay-per-use models, data is sent to Stripe via Billing Meters API:

await stripe.billing.meters.createEvent({
  event_name: 'api_requests',
  payload: {
    stripe_customer_id: tenant.stripeCustomerId,
    value: requestCount,
  },
  timestamp: Math.floor(Date.now() / 1000),
});

For fixed plans with overage, compare usage against quota at period end and issue an additional invoice.

Step-by-step guide to implementing rate limiting

  1. Analyze current traffic: use tools like Prometheus or Datadog to identify peak RPS, typical endpoints, and clients.
  2. Choose an algorithm: Token Bucket for burst loads, Fixed Window for simple scenarios.
  3. Implement middleware: integrate chosen library with Redis store.
  4. Configure headers: add RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset.
  5. Test: verify behavior when exceeding limit (expect 429).
  6. Monitor: set alerts at 80% quota and collect usage metrics.
  7. Document API responses: include example 429 errors in OpenAPI spec.
  8. Load test: ensure solution handles up to 10,000 RPS without degradation.

What's included in the work

When ordering implementation, you receive:

  • Source code for rate limiting middleware with chosen algorithm
  • Configured usage metric collection system on Redis + PostgreSQL
  • Client dashboard (custom or based on Grafana)
  • API documentation (OpenAPI) with example 429 responses
  • Integration with your billing system (Stripe, Chargebee, etc.)
  • Stability guarantee: solution throughput tested up to 10 000 RPS

With 5+ years of experience in SaaS backend development, we have implemented rate limiting for 20+ products—from startups to platforms with audiences over 10 000 users. Order rate limiting implementation for your SaaS and get a free architect consultation. To discuss details, contact us. Implementation cost starts at $5,000 for basic rate limiting.

Typical timelines

Stage Duration
Basic rate limiting with Redis and RFC 6585 headers 2–3 days
Usage tracking with PostgreSQL aggregation and dashboard 5–7 days
Integration with Stripe Billing Meters and alerts another 3 days

Specific cost is calculated individually after architecture and load analysis.

API Development with REST, GraphQL, WebSocket, and tRPC

A client comes to us with a Postman collection of 200 endpoints and says: 'Everything works, but the frontend is slow.' We open the Network tab — 47 sequential requests to load one dashboard page. Each one waits for the previous. This is not a server speed issue — it's an API architecture problem. With 10 years on the market, we've redesigned dozens of such integrations, and we guarantee: the right protocol and contract solve the problem at its root.

When REST stops being enough

REST works well for simple CRUD operations. But as soon as a mobile app appears alongside the web interface, over-fetching begins: the mobile app requests /api/users/123 and gets a 4KB object, but only needs name and avatar. Multiply that by a list of 50 users — 200KB traffic instead of 8KB.

GraphQL solves this with selection sets. The client describes exactly the fields it needs, and the server returns only those. On a project with React Native + Next.js, we migrated from REST to Apollo Server: payload size on the main screen dropped from 340KB to 28KB — a 92% traffic savings. Our certified engineers confirm: the typical pain when adopting GraphQL is N+1 query. A resolver for the author field on a post calls SELECT * FROM users WHERE id = ? for each post in the list. On a page with 20 posts — 21 database queries. Solved with DataLoader — it batches queries and turns them into one SELECT * FROM users WHERE id IN (...).

What is tRPC and how is it better than REST/GraphQL?

If the entire stack is TypeScript (Next.js + Node/Bun), tRPC removes a whole layer of problems. You define a procedure on the server — the client gets full type-safety automatically, without code generation and without Swagger. Renamed a field in the Zod schema — TypeScript highlights all places on the frontend where it's used. tRPC reduces code by 2 times compared to REST + Swagger + openapi-typescript: no need to maintain a separate specification and generate types — everything is inferred from runtime validators. However, tRPC is not suitable if the API is consumed by third-party clients or mobile apps in other languages — in such cases we use GraphQL or REST with OpenAPI specification.

WebSocket and real-time: when SSE, when WS?

HTTP polling every 5 seconds is an illusion of real-time with up to 5 seconds delay and useless server load. For chats, live notifications, collaborative editing — WebSocket or Server-Sent Events. SSE is a one-way stream from server to client, works over ordinary HTTP, automatically reconnects. Suitable for notifications, data streaming, progress bars. WebSocket is bidirectional, needed for chats and collaborative features. Experience shows: 80% of 'real-time' tasks are solved with SSE, not WebSocket — fewer infrastructure complexities.

A typical mistake: opening a WebSocket connection for each page component. On one project, the dashboard opened 12 parallel WS connections. The correct approach is one connection manager at the application level, subscriptions through it. In our work results, we always transfer the connection scheme and a ready solution.

Protocol Typing Over-fetching Versioning Real-time
REST Weak (OpenAPI) Yes URL / Header Polling
GraphQL Strong (SDL) No Deprecation Subscriptions
tRPC Full (TypeScript) No TypeScript checks Subscriptions (optional)

Swagger / OpenAPI as a contract

Documentation written after the fact becomes outdated the day after release. We write the OpenAPI 3.1 specification before development starts; it becomes the contract between frontend and backend. The frontend generates types via openapi-typescript, the backend validates incoming data using generated schemas. Contract deviation from implementation is caught on CI, not during review. For Laravel — l5-swagger or dedoc/scramble. For Node.js — @fastify/swagger or Zod + zod-to-openapi.

How to properly authenticate an API?

JWT with long-lived access tokens without rotation is a source of problems when compromised. The correct scheme: access token for 15 minutes, refresh token for 30 days with rotation on each use. Refresh token stored in an httpOnly cookie, access token in memory (not in localStorage). For inter-service communication — API Keys with scope limitations or mTLS. OAuth 2.0 with PKCE for public clients (SPA, mobile).

How to handle versioning and backward compatibility?

Breaking changes in an API without versioning break clients. Three approaches we use in projects:

Method Example When to use
URL versioning /api/v2/ REST API with long-term legacy support
Header versioning Accept: application/vnd.api+json;version=2 Minimal URL changes
Evolutionary (deprecation) Adding fields, GraphQL deprecated directive For GraphQL — smooth field removal

We guarantee backward compatibility through automated checks (oasdiff) on CI.

How we develop APIs: step-by-step plan

  1. Analysis — audit of current integrations, data schema compilation, protocol selection (REST/GraphQL/tRPC/WebSocket).
  2. Contract design — OpenAPI or SDL (GraphQL) before the first line of code.
  3. Development — implementation per contract, unit tests for each endpoint.
  4. Load testing — k6: 500 virtual users, 10 minutes, p95 latency ≤ 200ms.
  5. Deployment — CI/CD with backward compatibility check, automatic documentation publication.
  6. Team training — handover of Postman collection or Playground, connection instructions.
Typical mistakes we eliminate
  • N+1 on queries without DataLoader.
  • No rate limiting — DDOS through unauthenticated endpoints.
  • Storing access token in localStorage.
  • Opening multiple WebSocket connections instead of a single connection manager.
  • Documentation not updated after release.

What is included (deliverables)

  • OpenAPI 3.1 specification (or SDL for GraphQL).
  • Generated client types for TypeScript / Dart / Kotlin.
  • Set of automated tests covering all endpoints (unit + integration).
  • Load tests (k6) and report (p50/p95/p99 latency, RPS).
  • Documentation in Swagger UI / Redoc / GraphiQL.
  • Team training (2–4 hour workshop).
  • Support for 30 days after delivery (per contract).

Our experience

  • 10+ years in the API development market.
  • 200+ completed projects (REST, GraphQL, WebSocket, tRPC).
  • 50+ certified engineers (AWS, Kubernetes, API Design).
  • Traffic savings averaging 85% when migrating from REST to GraphQL for mobile apps.
  • 100% backward compatibility — not a single broken client in the last 3 years.

Timeline

API development for a typical SaaS project with 30–50 endpoints: from 3 to 8 weeks depending on business logic complexity and number of external integrations. Migration of an existing REST API to GraphQL: from 2 to 6 weeks. Adding a WebSocket layer to an existing backend: from 1 to 3 weeks. Cost is calculated individually after an audit. Get a consultation — contact us to discuss your project.