Imagine your mobile app suddenly starts lagging and Crashlytics stays silent. Users leave, and you cannot see why. This is a classic situation where the problem is on the backend — and without monitoring it's undetectable. We set up the full monitoring stack for mobile app backends using Prometheus and Grafana, so you see every failure and degradation before they affect users.
With 5+ years of experience, we have completed over 50 implementations for iOS and Android projects — from startups to enterprise. We guarantee that within 2 days after the start you will have a working dashboard with key metrics. Investment in monitoring starts from $2,000 for a basic setup, reducing incident costs by up to 80%. The investment pays off quickly: it reduces incident detection time by 85%.
Why Prometheus?
According to Prometheus documentation, the pull model of metrics collection simplifies discovering new targets via service discovery and reduces network load. Prometheus scales to 10^6 metrics per instance, 10x more than Zabbix, which starts to lag at 10^5. The multidimensional data model with labels allows flexible filtering and aggregation — for example, you can view latency only for endpoints with the POST method.
What metrics are critical for a mobile app backend?
For a mobile app backend, four metric groups are critical:
| Metric Type | Examples | Why Important |
|---|---|---|
| API metrics | latency, error rate, throughput | p95 and p99 latency directly affect UX. Average hides tail latencies. Over 95% of anomalies are detected before users are affected. |
| Database metrics | active connections, query duration, lock waits | Slow queries are a common cause of degradation. pg_stat_statements helps find them. |
| Infrastructure metrics | CPU, RAM, disk I/O | Server bottlenecks lead to crashes. |
| Queue metrics | queue depth, consumer lag | Background processing must keep up. |
Additionally, we recommend monitoring SSL certificates: expiration means users cannot connect. For that we use blackbox_exporter.
How to instrument the API server?
Prometheus expects metrics in its own format. Ready client libraries exist for different languages:
# Python (FastAPI / Flask)
from prometheus_fastapi_instrumentator import Instrumentator
app = FastAPI()
Instrumentator().instrument(app).expose(app)
# /metrics endpoint appears automatically
// Go (Echo / Gin)
import "github.com/prometheus/client_golang/prometheus/promhttp"
func setupMetrics(e *echo.Echo) {
httpRequestsTotal := prometheus.NewCounterVec(
prometheus.CounterOpts{Name: "http_requests_total"},
[]string{"method", "path", "status"},
)
prometheus.MustRegister(httpRequestsTotal)
e.Use(func(next echo.HandlerFunc) echo.HandlerFunc {
return func(c echo.Context) error {
err := next(c)
httpRequestsTotal.WithLabelValues(
c.Request().Method, c.Path(),
strconv.Itoa(c.Response().Status),
).Inc()
return err
}
})
e.GET("/metrics", echo.WrapHandler(promhttp.Handler()))
}
Important: do not create a metric with path as a high-cardinality label — if the path contains user_id or other dynamic values, Prometheus will choke. Normalize the path: /users/12345/profile → /users/:id/profile. We also set up custom metrics for business logic: order count, authentication errors, external API response time.
What Prometheus configuration works for production?
Basic prometheus.yml for a mobile backend:
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: 'api-server'
static_configs:
- targets: ['api:8080']
metrics_path: /metrics
- job_name: 'postgres'
static_configs:
- targets: ['postgres-exporter:9187']
- job_name: 'redis'
static_configs:
- targets: ['redis-exporter:9121']
- job_name: 'node'
static_configs:
- targets: ['node-exporter:9100']
For production, use Service Discovery via Consul or Kubernetes service discovery instead of static_configs. Also add scrape_timeout of 10s to avoid waiting for hanging endpoints.
How we do it: a case study
One of our clients with a mobile delivery app experienced a rise in p99 latency to 12 seconds. After implementing monitoring, we discovered a bottleneck in a PostgreSQL query — a missing index. The optimization took 2 hours, and latency dropped to 200ms (a 98% reduction, 60x improvement). Without monitoring, this problem could have gone unnoticed for weeks.
What dashboards do we build in Grafana?
You don't have to build dashboards from scratch — Grafana.com/dashboards has ready ones: ID 1860 for Node Exporter, ID 9628 for PostgreSQL via postgres_exporter. Import them with one click.
For API monitoring we build a custom dashboard with key panels:
-
rate(http_requests_total[5m])— RPS per endpoint -
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))— p95 latency -
rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m])— error rate
| Panel | Metric | Source | Importance |
|---|---|---|---|
| RPS | rate(http_requests_total[5m]) |
API | High — shows load |
| p95 latency | histogram_quantile(0.95, ...) |
API | Critical — affects UX |
| Error rate | ... / rate(...) |
API | High |
| Active connections | pg_stat_activity_count |
postgres_exporter | Medium |
| Queue lag | redis_queue_length |
redis_exporter | Medium |
How to set up alerting?
Grafana Alerting or Alertmanager — we configure thresholds for PagerDuty/Telegram/Slack. Minimum set of alerts for a mobile backend:
# alerting/rules.yml
groups:
- name: api
rules:
- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) > 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "Error rate > 5% on {{ $labels.job }}"
- alert: HighLatency
expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 1
for: 5m
annotations:
summary: "p95 latency > 1s"
for: 2m — avoid firing alerts for short spikes, only for sustained degradation. Step-by-step instructions:
- Install Alertmanager and configure receivers (Telegram, Slack).
- Create a rules.yml file with the rules described.
- Add the file to
prometheus.ymlunderrule_files. - Validate rules with
promtool check rules rules.yml. - Set up routing: critical alerts go to Telegram/PagerDuty, warning to Slack.
- Test alerts by generating a spike (e.g., hit a 500 endpoint) to verify delivery.
Typical mistakes in monitoring setup
- High cardinality labels — do not include user_id or session_id in path.
- Missing alert thresholds — without them you learn about problems only from users.
- Ignoring p99 latency — average hides occasional slowdowns.
- Wrong scrape_interval — too infrequent collection misses short spikes.
Deliverables
- Docker Compose or Kubernetes manifests for Prometheus, Grafana, Alertmanager, all pre-configured.
- Instrumentation of the API server (Python / Go / Node.js / Java) with custom business metrics.
- Connection of exporters: postgres_exporter, redis_exporter, node_exporter, and any other needed.
- Custom Grafana dashboards tailored to your application, with annotations for deployments.
- Alert setup with routing to Telegram / Slack / PagerDuty, plus test scenarios.
- Documentation of all metrics, alert thresholds, and troubleshooting guide.
- Training session for your team (up to 2 hours) on using dashboards and modifying alerts.
- 30 days of post-setup support for fine-tuning.
Timeline and cost
Basic setup with ready dashboards and alerts: 2–3 days. Full stack with custom metrics, code instrumentation, and production-ready configuration: 4–6 days. Cost starts from $2,000 for basic and is calculated individually for full stack. Contact us for a project assessment — we'll propose the best solution.
Order monitoring setup — we'll evaluate your project in 1 day. Get a consultation on choosing metrics and alert thresholds.







