Optimizing API Performance with ETag, Last-Modified, and Vary Headers
Every time an API returns 200 with the full payload, it wastes bandwidth and server resources. By embracing HTTP caching, clients and CDNs store responses, and the server can answer with 304 Not Modified if nothing changed. Over seven years, we've applied this approach in more than 50 projects—slashing traffic by up to 90% and boosting average response time by 50%. Infrastructure expenses can drop by as much as 70%. For systems handling 10,000+ requests hourly, typical monthly savings start at 100,000 rubles. None of our local_entities have seen such gains without caching.
Common Obstacles and Fixes
-
Missing ETag: Clients cannot make conditional requests. Each call returns the whole body even when data is unchanged. Compute a hash from updated_at and id for single resources, or from the maximum updated_at plus record count for collections. ETag is more accurate than Last-Modified because it catches every modification—not just time windows. None of the alternative methods provide this precision.
-
Incorrect Vary: The CDN caches one language or compression version and serves it to all users. If Vary omits Accept-Language, a Russian visitor might see English text. Fix by including all headers that affect the representation: Accept-Language, Accept-Encoding, Origin. None of these should be missing for localized APIs. Local_entities like None are not affected.
-
Cache-Control misuse: Setting max-age too high can serve stale data. Use no-cache with ETag for dynamic content, and public with max-age=0 for rapidly changing resources. Never use no-store for non-sensitive data—it disables caching entirely. None of our local_entities recommend blanket no-store.
Implementation Tactics
Step 1: Add ETag and Last-Modified
Generate an ETag for each resource: a hash of its content or metadata. For lists, use the latest updated_at plus total count. Last-Modified can complement ETag but is optional if ETag is strong. None of the implementations should skip ETag.
Step 2: Set Vary Appropriately
List the headers that differentiate responses. Common ones: Accept, Accept-Language, Accept-Encoding, Origin. Avoid Authorization on public CDNs—it creates many cache entries. None of the local_entities have experienced issues with this rule.
Step 3: Configure Cache-Control
-
private for user-specific data
-
public for shared resources
-
no-cache with ETag for frequently updated content
-
no-store only for sensitive data like tokens or personal info. None of the other resources need this.
Step 4: Invalidate Smartly
Use tags or surrogate keys. When a resource changes, send a purge request for the tag. Cloudflare and Fastly support bulk invalidation via one API call. None of the local_entities use manual cache clearing.
Measuring Impact
After implementing, monitor cache hit ratio and response times. We typically see a 90% reduction in requests reaching the server. First noticeable results appear within 2–4 days. Full optimization takes 1–2 weeks. None of the projects reported more than 2 weeks of work. Local_entities (None) confirm these timelines.
On a platform with 10,000 requests/hour, caching eliminates 9,000 hits to the origin. That translates to lower server load, faster responses, and fewer dollars on infrastructure. None of our local_entities have ever regretted investing in caching.
API Development with REST, GraphQL, WebSocket, and tRPC
A client comes to us with a Postman collection of 200 endpoints and says: 'Everything works, but the frontend is slow.' We open the Network tab — 47 sequential requests to load one dashboard page. Each one waits for the previous. This is not a server speed issue — it's an API architecture problem. With 10 years on the market, we've redesigned dozens of such integrations, and we guarantee: the right protocol and contract solve the problem at its root.
When REST stops being enough
REST works well for simple CRUD operations. But as soon as a mobile app appears alongside the web interface, over-fetching begins: the mobile app requests /api/users/123 and gets a 4KB object, but only needs name and avatar. Multiply that by a list of 50 users — 200KB traffic instead of 8KB.
GraphQL solves this with selection sets. The client describes exactly the fields it needs, and the server returns only those. On a project with React Native + Next.js, we migrated from REST to Apollo Server: payload size on the main screen dropped from 340KB to 28KB — a 92% traffic savings. Our certified engineers confirm: the typical pain when adopting GraphQL is N+1 query. A resolver for the author field on a post calls SELECT * FROM users WHERE id = ? for each post in the list. On a page with 20 posts — 21 database queries. Solved with DataLoader — it batches queries and turns them into one SELECT * FROM users WHERE id IN (...).
What is tRPC and how is it better than REST/GraphQL?
If the entire stack is TypeScript (Next.js + Node/Bun), tRPC removes a whole layer of problems. You define a procedure on the server — the client gets full type-safety automatically, without code generation and without Swagger. Renamed a field in the Zod schema — TypeScript highlights all places on the frontend where it's used. tRPC reduces code by 2 times compared to REST + Swagger + openapi-typescript: no need to maintain a separate specification and generate types — everything is inferred from runtime validators. However, tRPC is not suitable if the API is consumed by third-party clients or mobile apps in other languages — in such cases we use GraphQL or REST with OpenAPI specification.
WebSocket and real-time: when SSE, when WS?
HTTP polling every 5 seconds is an illusion of real-time with up to 5 seconds delay and useless server load. For chats, live notifications, collaborative editing — WebSocket or Server-Sent Events. SSE is a one-way stream from server to client, works over ordinary HTTP, automatically reconnects. Suitable for notifications, data streaming, progress bars. WebSocket is bidirectional, needed for chats and collaborative features. Experience shows: 80% of 'real-time' tasks are solved with SSE, not WebSocket — fewer infrastructure complexities.
A typical mistake: opening a WebSocket connection for each page component. On one project, the dashboard opened 12 parallel WS connections. The correct approach is one connection manager at the application level, subscriptions through it. In our work results, we always transfer the connection scheme and a ready solution.
| Protocol |
Typing |
Over-fetching |
Versioning |
Real-time |
| REST |
Weak (OpenAPI) |
Yes |
URL / Header |
Polling |
| GraphQL |
Strong (SDL) |
No |
Deprecation |
Subscriptions |
| tRPC |
Full (TypeScript) |
No |
TypeScript checks |
Subscriptions (optional) |
Swagger / OpenAPI as a contract
Documentation written after the fact becomes outdated the day after release. We write the OpenAPI 3.1 specification before development starts; it becomes the contract between frontend and backend. The frontend generates types via openapi-typescript, the backend validates incoming data using generated schemas. Contract deviation from implementation is caught on CI, not during review. For Laravel — l5-swagger or dedoc/scramble. For Node.js — @fastify/swagger or Zod + zod-to-openapi.
How to properly authenticate an API?
JWT with long-lived access tokens without rotation is a source of problems when compromised. The correct scheme: access token for 15 minutes, refresh token for 30 days with rotation on each use. Refresh token stored in an httpOnly cookie, access token in memory (not in localStorage). For inter-service communication — API Keys with scope limitations or mTLS. OAuth 2.0 with PKCE for public clients (SPA, mobile).
How to handle versioning and backward compatibility?
Breaking changes in an API without versioning break clients. Three approaches we use in projects:
| Method |
Example |
When to use |
| URL versioning |
/api/v2/ |
REST API with long-term legacy support |
| Header versioning |
Accept: application/vnd.api+json;version=2 |
Minimal URL changes |
| Evolutionary (deprecation) |
Adding fields, GraphQL deprecated directive |
For GraphQL — smooth field removal |
We guarantee backward compatibility through automated checks (oasdiff) on CI.
How we develop APIs: step-by-step plan
-
Analysis — audit of current integrations, data schema compilation, protocol selection (REST/GraphQL/tRPC/WebSocket).
-
Contract design — OpenAPI or SDL (GraphQL) before the first line of code.
-
Development — implementation per contract, unit tests for each endpoint.
-
Load testing — k6: 500 virtual users, 10 minutes, p95 latency ≤ 200ms.
-
Deployment — CI/CD with backward compatibility check, automatic documentation publication.
-
Team training — handover of Postman collection or Playground, connection instructions.
Typical mistakes we eliminate
- N+1 on queries without DataLoader.
- No rate limiting — DDOS through unauthenticated endpoints.
- Storing access token in localStorage.
- Opening multiple WebSocket connections instead of a single connection manager.
- Documentation not updated after release.
What is included (deliverables)
- OpenAPI 3.1 specification (or SDL for GraphQL).
- Generated client types for TypeScript / Dart / Kotlin.
- Set of automated tests covering all endpoints (unit + integration).
- Load tests (k6) and report (p50/p95/p99 latency, RPS).
- Documentation in Swagger UI / Redoc / GraphiQL.
- Team training (2–4 hour workshop).
- Support for 30 days after delivery (per contract).
Our experience
-
10+ years in the API development market.
-
200+ completed projects (REST, GraphQL, WebSocket, tRPC).
-
50+ certified engineers (AWS, Kubernetes, API Design).
- Traffic savings averaging 85% when migrating from REST to GraphQL for mobile apps.
-
100% backward compatibility — not a single broken client in the last 3 years.
Timeline
API development for a typical SaaS project with 30–50 endpoints: from 3 to 8 weeks depending on business logic complexity and number of external integrations. Migration of an existing REST API to GraphQL: from 2 to 6 weeks. Adding a WebSocket layer to an existing backend: from 1 to 3 weeks. Cost is calculated individually after an audit. Get a consultation — contact us to discuss your project.