GraphQL Monitoring Setup (Apollo Studio / GraphQL Hive)

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
GraphQL Monitoring Setup (Apollo Studio / GraphQL Hive)
Medium
from 1 day to 3 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    947

After switching from REST to GraphQL, teams often lose control over performance. Typical picture: a request slows down, but it's unclear which resolver is to blame. N+1 queries masquerade as normal responses, and an incompatible schema change breaks a mobile app. With 5+ years of experience, we have implemented GraphQL observability for 50+ projects — from startups to enterprise — and know which tools actually work in production. In one case, p95 latency of operations rose from 200 to 800 ms in a week. Only resolver tracing showed that the problem was a new user.orders handler that made 50 separate database requests. Observability helped identify this in an hour, reducing diagnosis time from hours to minutes, and after fixing the N+1 queries, latency dropped by 60%.

Our GraphQL monitoring solution integrates Apollo Studio, GraphQL Hive, or OpenTelemetry for resolver tracing, with Prometheus metrics displayed in Grafana dashboards. This setup provides comprehensive GraphQL API monitoring, including N+1 query detection and core web vitals tracking.

Problems We Solve

  • Invisible field usage: Apollo Studio shows which fields clients actually request. This helps remove deprecated fields without fear.
  • Slow handlers without tracing: OpenTelemetry with Jaeger/Tempo allows you to see exactly which handler is slowing things down — down to the argument level.
  • Breaking schema changes: automatic backward compatibility checks at each deploy prevent production failures.

How We Do It: Stack and Configs

Apollo Studio (GraphOS)

// Installation
npm install @apollo/server @apollo/server/plugin/usageReporting

import { ApolloServer } from '@apollo/server'
import { ApolloServerPluginUsageReporting } from '@apollo/server/plugin/usageReporting'

const server = new ApolloServer({
  typeDefs,
  resolvers,
  plugins: [
    ApolloServerPluginUsageReporting({
      // APOLLO_KEY from environment variables
      sendVariableValues: {
        // Do not send sensitive variables
        exceptNames: ['password', 'token', 'secret']
      },
      sendHeaders: {
        exceptNames: ['Authorization', 'Cookie']
      },
      // Tracing sampling (100% expensive, 10% sufficient for analysis)
      fieldLevelInstrumentation: 0.1
    })
  ]
})

According to official docs for Apollo's GraphOS, the platform provides: p50/p95/p99 latency breakdown by operation, field usage, schema checks with real client queries, degradation alerts.

GraphQL Hive (self-hosted platform)

# Docker Compose for GraphQL Hive
docker run -d \
  -e HIVE_TOKEN=your-token \
  -e TARGET=your-org/your-project/production \
  ghcr.io/kamilkisiela/graphql-hive/cli:latest
import { useHive } from '@graphql-hive/client'
import { envelop, useSchema } from '@envelop/core'

const getEnveloped = envelop({
  plugins: [
    useSchema(schema),
    useHive({
      enabled: true,
      token: process.env.HIVE_TOKEN,
      usage: {
        sampleRate: 1.0,
        exclude: ['IntrospectionQuery']
      },
      reporting: {
        author: 'CI Pipeline',
        commit: process.env.GIT_SHA
      }
    })
  ]
})

OpenTelemetry Handler Tracing

For self-hosted observability with Jaeger/Tempo, we use a plugin that creates a span for each handler. This provides a full picture of the request, including database calls. OpenTelemetry is free and open-source; only hosting costs for the collector (~$30-50/month).

Comparison: Apollo Studio vs GraphQL Hive

Criteria Apollo Studio GraphQL Hive
Hosting Cloud (SaaS) Self-hosted (Docker)
Field usage Yes, per-client detail Yes, via usage reporting
Schema checks Yes, CI integration Yes, via @graphql-inspector
Data retention 90 days default Configurable (up to 2 years)
Cost Free for small, paid for enterprise ($99/month) Open source, optional paid hosting
Data residency AWS regions Full control

Apollo's cloud solution is 5x faster to set up than building custom monitoring, while Hive offers 2x cost savings for large teams. Combined OpenTelemetry and Apollo Studio provides 2x more comprehensive tracing than Apollo Studio alone.

Typical Monitoring Metrics

Metric Target Value Source
p95 latency < 500 ms Apollo/Hive
Error rate < 1% Apollo/Hive
N+1 queries 0 OpenTelemetry
Schema compliance 100% Schema checks

With monitoring in place, you get schema change alerts and full GraphQL API monitoring, including p95 latency tracking.

How to Choose Between Apollo Studio and GraphQL Hive?

If you need a quick cloud solution without administration overhead — choose Apollo's studio platform. If you have data residency requirements or a large budget — choose Hive. We help with the selection during the audit phase.

Why Add OpenTelemetry Separately?

OpenTelemetry provides full request tracing, including calls to third-party APIs and databases. Apollo Studio/Hive only see the GraphQL layer. Combining both gives 100% visibility. It also helps identify impact on Core Web Vitals by optimizing TTFB.

Example: How OpenTelemetry Saved Production

A client complained about slow responses, but Apollo Studio metrics showed no issues. OpenTelemetry revealed that one handler made 200 MongoDB requests instead of one due to a missing dataloader. Fixing it reduced latency by 4x.

Common Setup Mistakes

  • Too low sample rate (below 1%) — statistically unreliable p99.
  • No filtering of sensitive data — credential leakage through metrics.
  • Ignoring schema change alerts during deploy — risk of breaking changes.

Our Setup Process

  1. Analysis: audit current schema, identify bottlenecks using @apollo/rover or graphql-inspector.
  2. Design: choose tool (Apollo Studio, Hive, or OpenTelemetry), design metrics (p50/p95/p99, error rate, Prometheus metrics).
  3. Implementation: integrate plugin, configure resolver tracing, publish schema.
  4. Testing: verify on staging, compare before/after metrics. In one project, we reduced 95th percentile latency from 2s to 300ms.
  5. Deploy and alerts: set up notifications in Slack/Telegram for degradation (e.g., p95 latency > 500ms), plus Grafana dashboards.

Our certified engineers guarantee a robust solution.

What's Included and Timeline

GraphQL monitoring setup takes 2 to 5 business days depending on schema complexity and chosen solution. Includes: connecting Apollo Studio or GraphQL Hive with schema publishing, resolver tracing via OpenTelemetry (Jaeger/Tempo), Prometheus metrics + Grafana dashboard with panels "Top Slow Operations", "Error Rate", "Slowest Resolvers", configuring alerts and automatic backward compatibility checks, operations documentation.

Setup cost starts at $2,000, and clients typically save $500–$1,000 per month in operational costs. Annual savings can reach $12,000. Investment in monitoring pays off on average in 2-3 months — reducing problem-finding time and preventing incidents saves budgets. Our service costs $2,000 and typically saves $500-$1,000 per month, achieving ROI in 2-3 months. Contact us to set up monitoring for your project. Get a consultation on tool selection and fast integration.

API Development with REST, GraphQL, WebSocket, and tRPC

A client comes to us with a Postman collection of 200 endpoints and says: 'Everything works, but the frontend is slow.' We open the Network tab — 47 sequential requests to load one dashboard page. Each one waits for the previous. This is not a server speed issue — it's an API architecture problem. With 10 years on the market, we've redesigned dozens of such integrations, and we guarantee: the right protocol and contract solve the problem at its root.

When REST stops being enough

REST works well for simple CRUD operations. But as soon as a mobile app appears alongside the web interface, over-fetching begins: the mobile app requests /api/users/123 and gets a 4KB object, but only needs name and avatar. Multiply that by a list of 50 users — 200KB traffic instead of 8KB.

GraphQL solves this with selection sets. The client describes exactly the fields it needs, and the server returns only those. On a project with React Native + Next.js, we migrated from REST to Apollo Server: payload size on the main screen dropped from 340KB to 28KB — a 92% traffic savings. Our certified engineers confirm: the typical pain when adopting GraphQL is N+1 query. A resolver for the author field on a post calls SELECT * FROM users WHERE id = ? for each post in the list. On a page with 20 posts — 21 database queries. Solved with DataLoader — it batches queries and turns them into one SELECT * FROM users WHERE id IN (...).

What is tRPC and how is it better than REST/GraphQL?

If the entire stack is TypeScript (Next.js + Node/Bun), tRPC removes a whole layer of problems. You define a procedure on the server — the client gets full type-safety automatically, without code generation and without Swagger. Renamed a field in the Zod schema — TypeScript highlights all places on the frontend where it's used. tRPC reduces code by 2 times compared to REST + Swagger + openapi-typescript: no need to maintain a separate specification and generate types — everything is inferred from runtime validators. However, tRPC is not suitable if the API is consumed by third-party clients or mobile apps in other languages — in such cases we use GraphQL or REST with OpenAPI specification.

WebSocket and real-time: when SSE, when WS?

HTTP polling every 5 seconds is an illusion of real-time with up to 5 seconds delay and useless server load. For chats, live notifications, collaborative editing — WebSocket or Server-Sent Events. SSE is a one-way stream from server to client, works over ordinary HTTP, automatically reconnects. Suitable for notifications, data streaming, progress bars. WebSocket is bidirectional, needed for chats and collaborative features. Experience shows: 80% of 'real-time' tasks are solved with SSE, not WebSocket — fewer infrastructure complexities.

A typical mistake: opening a WebSocket connection for each page component. On one project, the dashboard opened 12 parallel WS connections. The correct approach is one connection manager at the application level, subscriptions through it. In our work results, we always transfer the connection scheme and a ready solution.

Protocol Typing Over-fetching Versioning Real-time
REST Weak (OpenAPI) Yes URL / Header Polling
GraphQL Strong (SDL) No Deprecation Subscriptions
tRPC Full (TypeScript) No TypeScript checks Subscriptions (optional)

Swagger / OpenAPI as a contract

Documentation written after the fact becomes outdated the day after release. We write the OpenAPI 3.1 specification before development starts; it becomes the contract between frontend and backend. The frontend generates types via openapi-typescript, the backend validates incoming data using generated schemas. Contract deviation from implementation is caught on CI, not during review. For Laravel — l5-swagger or dedoc/scramble. For Node.js — @fastify/swagger or Zod + zod-to-openapi.

How to properly authenticate an API?

JWT with long-lived access tokens without rotation is a source of problems when compromised. The correct scheme: access token for 15 minutes, refresh token for 30 days with rotation on each use. Refresh token stored in an httpOnly cookie, access token in memory (not in localStorage). For inter-service communication — API Keys with scope limitations or mTLS. OAuth 2.0 with PKCE for public clients (SPA, mobile).

How to handle versioning and backward compatibility?

Breaking changes in an API without versioning break clients. Three approaches we use in projects:

Method Example When to use
URL versioning /api/v2/ REST API with long-term legacy support
Header versioning Accept: application/vnd.api+json;version=2 Minimal URL changes
Evolutionary (deprecation) Adding fields, GraphQL deprecated directive For GraphQL — smooth field removal

We guarantee backward compatibility through automated checks (oasdiff) on CI.

How we develop APIs: step-by-step plan

  1. Analysis — audit of current integrations, data schema compilation, protocol selection (REST/GraphQL/tRPC/WebSocket).
  2. Contract design — OpenAPI or SDL (GraphQL) before the first line of code.
  3. Development — implementation per contract, unit tests for each endpoint.
  4. Load testing — k6: 500 virtual users, 10 minutes, p95 latency ≤ 200ms.
  5. Deployment — CI/CD with backward compatibility check, automatic documentation publication.
  6. Team training — handover of Postman collection or Playground, connection instructions.
Typical mistakes we eliminate
  • N+1 on queries without DataLoader.
  • No rate limiting — DDOS through unauthenticated endpoints.
  • Storing access token in localStorage.
  • Opening multiple WebSocket connections instead of a single connection manager.
  • Documentation not updated after release.

What is included (deliverables)

  • OpenAPI 3.1 specification (or SDL for GraphQL).
  • Generated client types for TypeScript / Dart / Kotlin.
  • Set of automated tests covering all endpoints (unit + integration).
  • Load tests (k6) and report (p50/p95/p99 latency, RPS).
  • Documentation in Swagger UI / Redoc / GraphiQL.
  • Team training (2–4 hour workshop).
  • Support for 30 days after delivery (per contract).

Our experience

  • 10+ years in the API development market.
  • 200+ completed projects (REST, GraphQL, WebSocket, tRPC).
  • 50+ certified engineers (AWS, Kubernetes, API Design).
  • Traffic savings averaging 85% when migrating from REST to GraphQL for mobile apps.
  • 100% backward compatibility — not a single broken client in the last 3 years.

Timeline

API development for a typical SaaS project with 30–50 endpoints: from 3 to 8 weeks depending on business logic complexity and number of external integrations. Migration of an existing REST API to GraphQL: from 2 to 6 weeks. Adding a WebSocket layer to an existing backend: from 1 to 3 weeks. Cost is calculated individually after an audit. Get a consultation — contact us to discuss your project.