Mobile Backend Monitoring: Prometheus + Grafana Setup

TRUETECH is engaged in the development, support and maintenance of iOS, Android, PWA mobile applications. We have extensive experience and expertise in publishing mobile applications in popular markets like Google Play, App Store, Amazon, AppGallery and others.

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Mobile Backend Monitoring: Prometheus + Grafana Setup
Medium
~2-3 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    860
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    746
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1163
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1035
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    970
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    564

Imagine your mobile app suddenly starts lagging and Crashlytics stays silent. Users leave, and you cannot see why. This is a classic situation where the problem is on the backend — and without monitoring it's undetectable. We set up the full monitoring stack for mobile app backends using Prometheus and Grafana, so you see every failure and degradation before they affect users.

With 5+ years of experience, we have completed over 50 implementations for iOS and Android projects — from startups to enterprise. We guarantee that within 2 days after the start you will have a working dashboard with key metrics. Investment in monitoring starts from $2,000 for a basic setup, reducing incident costs by up to 80%. The investment pays off quickly: it reduces incident detection time by 85%.

Why Prometheus?

According to Prometheus documentation, the pull model of metrics collection simplifies discovering new targets via service discovery and reduces network load. Prometheus scales to 10^6 metrics per instance, 10x more than Zabbix, which starts to lag at 10^5. The multidimensional data model with labels allows flexible filtering and aggregation — for example, you can view latency only for endpoints with the POST method.

What metrics are critical for a mobile app backend?

For a mobile app backend, four metric groups are critical:

Metric Type Examples Why Important
API metrics latency, error rate, throughput p95 and p99 latency directly affect UX. Average hides tail latencies. Over 95% of anomalies are detected before users are affected.
Database metrics active connections, query duration, lock waits Slow queries are a common cause of degradation. pg_stat_statements helps find them.
Infrastructure metrics CPU, RAM, disk I/O Server bottlenecks lead to crashes.
Queue metrics queue depth, consumer lag Background processing must keep up.

Additionally, we recommend monitoring SSL certificates: expiration means users cannot connect. For that we use blackbox_exporter.

How to instrument the API server?

Prometheus expects metrics in its own format. Ready client libraries exist for different languages:

# Python (FastAPI / Flask)
from prometheus_fastapi_instrumentator import Instrumentator

app = FastAPI()
Instrumentator().instrument(app).expose(app)
# /metrics endpoint appears automatically
// Go (Echo / Gin)
import "github.com/prometheus/client_golang/prometheus/promhttp"

func setupMetrics(e *echo.Echo) {
    httpRequestsTotal := prometheus.NewCounterVec(
        prometheus.CounterOpts{Name: "http_requests_total"},
        []string{"method", "path", "status"},
    )
    prometheus.MustRegister(httpRequestsTotal)

    e.Use(func(next echo.HandlerFunc) echo.HandlerFunc {
        return func(c echo.Context) error {
            err := next(c)
            httpRequestsTotal.WithLabelValues(
                c.Request().Method, c.Path(),
                strconv.Itoa(c.Response().Status),
            ).Inc()
            return err
        }
    })
    e.GET("/metrics", echo.WrapHandler(promhttp.Handler()))
}

Important: do not create a metric with path as a high-cardinality label — if the path contains user_id or other dynamic values, Prometheus will choke. Normalize the path: /users/12345/profile/users/:id/profile. We also set up custom metrics for business logic: order count, authentication errors, external API response time.

What Prometheus configuration works for production?

Basic prometheus.yml for a mobile backend:

global:
  scrape_interval: 15s
  evaluation_interval: 15s

scrape_configs:
  - job_name: 'api-server'
    static_configs:
      - targets: ['api:8080']
    metrics_path: /metrics

  - job_name: 'postgres'
    static_configs:
      - targets: ['postgres-exporter:9187']

  - job_name: 'redis'
    static_configs:
      - targets: ['redis-exporter:9121']

  - job_name: 'node'
    static_configs:
      - targets: ['node-exporter:9100']

For production, use Service Discovery via Consul or Kubernetes service discovery instead of static_configs. Also add scrape_timeout of 10s to avoid waiting for hanging endpoints.

How we do it: a case study

One of our clients with a mobile delivery app experienced a rise in p99 latency to 12 seconds. After implementing monitoring, we discovered a bottleneck in a PostgreSQL query — a missing index. The optimization took 2 hours, and latency dropped to 200ms (a 98% reduction, 60x improvement). Without monitoring, this problem could have gone unnoticed for weeks.

What dashboards do we build in Grafana?

You don't have to build dashboards from scratch — Grafana.com/dashboards has ready ones: ID 1860 for Node Exporter, ID 9628 for PostgreSQL via postgres_exporter. Import them with one click.

For API monitoring we build a custom dashboard with key panels:

  • rate(http_requests_total[5m]) — RPS per endpoint
  • histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) — p95 latency
  • rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) — error rate
Panel Metric Source Importance
RPS rate(http_requests_total[5m]) API High — shows load
p95 latency histogram_quantile(0.95, ...) API Critical — affects UX
Error rate ... / rate(...) API High
Active connections pg_stat_activity_count postgres_exporter Medium
Queue lag redis_queue_length redis_exporter Medium

How to set up alerting?

Grafana Alerting or Alertmanager — we configure thresholds for PagerDuty/Telegram/Slack. Minimum set of alerts for a mobile backend:

# alerting/rules.yml
groups:
  - name: api
    rules:
      - alert: HighErrorRate
        expr: rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) > 0.05
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Error rate > 5% on {{ $labels.job }}"

      - alert: HighLatency
        expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 1
        for: 5m
        annotations:
          summary: "p95 latency > 1s"

for: 2m — avoid firing alerts for short spikes, only for sustained degradation. Step-by-step instructions:

  1. Install Alertmanager and configure receivers (Telegram, Slack).
  2. Create a rules.yml file with the rules described.
  3. Add the file to prometheus.yml under rule_files.
  4. Validate rules with promtool check rules rules.yml.
  5. Set up routing: critical alerts go to Telegram/PagerDuty, warning to Slack.
  6. Test alerts by generating a spike (e.g., hit a 500 endpoint) to verify delivery.

Typical mistakes in monitoring setup

  • High cardinality labels — do not include user_id or session_id in path.
  • Missing alert thresholds — without them you learn about problems only from users.
  • Ignoring p99 latency — average hides occasional slowdowns.
  • Wrong scrape_interval — too infrequent collection misses short spikes.

Deliverables

  • Docker Compose or Kubernetes manifests for Prometheus, Grafana, Alertmanager, all pre-configured.
  • Instrumentation of the API server (Python / Go / Node.js / Java) with custom business metrics.
  • Connection of exporters: postgres_exporter, redis_exporter, node_exporter, and any other needed.
  • Custom Grafana dashboards tailored to your application, with annotations for deployments.
  • Alert setup with routing to Telegram / Slack / PagerDuty, plus test scenarios.
  • Documentation of all metrics, alert thresholds, and troubleshooting guide.
  • Training session for your team (up to 2 hours) on using dashboards and modifying alerts.
  • 30 days of post-setup support for fine-tuning.

Timeline and cost

Basic setup with ready dashboards and alerts: 2–3 days. Full stack with custom metrics, code instrumentation, and production-ready configuration: 4–6 days. Cost starts from $2,000 for basic and is calculated individually for full stack. Contact us for a project assessment — we'll propose the best solution.

Order monitoring setup — we'll evaluate your project in 1 day. Get a consultation on choosing metrics and alert thresholds.

Mobile App Analytics: Firebase, Amplitude, AppsFlyer and Attribution

Our team regularly encounters projects where analytics is already "set up" but yields no real insights. A typical example is a startup with 50k DAU: tracking dozens of events without a single answer to the question "why don't users reach payment?". In two weeks we built a basic funnel and found that 70% of users drop off at the phone number verification screen. After fixing the bug, retention increased by 12%. The takeaway: analytics should start with specific questions, not tracking everything indiscriminately.

Why Event Taxonomy is the Foundation of Mobile App Analytics?

Firebase Analytics, Amplitude, Mixpanel — technically similar. The difference lies in what you put into them. A common mistake: events like screen_view, button_tap_1, button_tap_2 without context. A month later, no one remembers what button_tap_2 means.

Proper taxonomy: object + action + context. product_viewed, checkout_started, payment_completed with parameters product_id, category, price, source. This allows building funnels, cohort analysis, and retention without additional tracking.

We document the naming convention in a tracking plan — a document (Google Sheet or Amplitude Data Catalog) describing every event, its parameters, and triggering conditions. The tracking plan is synced with the analytics team before development begins, not after. This approach ensures that data remains interpretable months later and doesn't become a dump. Experience from 50+ projects confirms: without a tracking plan, analytics maintenance costs increase 2-3 times due to rework.

What Should You Choose for Mobile App Analytics: Firebase, Amplitude, or Mixpanel?

The table below highlights key differences between the three popular platforms. Choice depends on budget, traffic, and tasks.

Criteria Firebase Analytics Amplitude Mixpanel
Free limit Unlimited (Spark plan) Up to 10M events/month Up to 1K MTU/month (Special)
Data latency Up to 24 hours (standard) Minutes (real-time) Minutes (real-time)
Funnels and cohorts Basic funnels, limited count Deep funnels, Journeys, cohorts Funnels, Retention, Insights
BigQuery export Yes (free, raw data) Yes (subscription) Yes (Enterprise)
Session Replay No Yes (iOS/Android SDK) No
Ad integration Google Ads (native) Via Universal Links Via partners

Firebase Analytics — free, deep integration with Google Ads, BigQuery export for raw data. Limitations: data latency up to 24 hours, limited funnels. For startups with Google Ads traffic, it's the first choice.

Amplitude — product analytics focused on cohorts and user journeys. Journeys (formerly Pathfinder) shows actual paths between events — not assumed funnels but real routes. Session Replay records sessions for UX analysis. The free tier up to 10M events/month is enough for most products at launch.

Mixpanel — close to Amplitude, stronger in real-time segmentation. Insights, Funnels, Retention cover 90% of product analysts' tasks.

How to Solve Multi-Channel Attribution with AppsFlyer?

Knowing where a user came from is a separate task. Firebase Attribution works only within the Google ecosystem. For multi-channel attribution (Facebook Ads, TikTok, Apple Search Ads, programmatic), an MMP (Mobile Measurement Partner) is needed.

AppsFlyer is the market leader. OneLink — universal deep link working on iOS and Android, correctly attributing installs from any channel. Protect360 — built-in fraud protection (fake installs, click injection on Android). Adjust and Branch are competitors with similar features. Branch excels in deep linking; Adjust is popular in gaming.

According to Apple, with iOS 14.5, apps must obtain user permission via ATT before collecting IDFA for tracking. AppsFlyer uses probabilistic matching (IP + user agent + timing) for these users — accuracy is lower but better than nothing. SKAdNetwork and Privacy Preserving Attribution provide aggregated data from Apple with a 24-72 hour delay.

How to Set Up Crash Analytics to Not Miss Bugs?

Firebase Crashlytics is the standard for crash reporting. It automatically groups crashes by stack trace, shows affected users %, and sends velocity alerts when crash rate increases by more than 10% per hour.

Important: symbolication. On iOS, .dSYM files must be automatically uploaded with each build — via Fastlane upload_symbols_to_crashlytics or Xcode Cloud built-in. Without symbols, crashes in Crashlytics appear as memory addresses. This happens more often than expected when switching to a new CI — in one project with 500k users, we found that 40% of crashes remained unsymbolicated due to a missing CI/CD step. After automation, bug response time dropped from 3 hours to 15 minutes.

For React Native and Flutter, @sentry/react-native and sentry_flutter provide additional context: breadcrumbs, network requests before the crash, Redux/Provider state.

Below is a comparison of popular crash analytics tools to choose according to your needs.

Criteria Firebase Crashlytics Sentry Instabug
Free limit Unlimited (Spark) 5k events/month 250 MAU
Grouping By stack trace + parameters By fingerprint By stack trace + metadata
Symbolication Automatic (via file) Automatic (via CLI) Automatic
Velocity alerts Yes (by % change) Yes (by count) Yes (by threshold)
Extra context Logs, Keys, Custom Keys Breadcrumbs, User, Tags User steps, network requests
Price Free (in Firebase) Paid plans available Paid plans available

Environment Setup

Three environments with separate Firebase projects: dev, staging, production. Mixing analytics from test sessions and production is a common mistake that skews all metrics. On iOS via GoogleService-Info.plist per scheme, on Android via google-services.json in each flavor folder.

Timelines: basic analytics with Firebase + Crashlytics — 3-5 days. Full tracking plan + Amplitude/Mixpanel with funnels and cohorts — 2-3 weeks. Attribution via AppsFlyer with deep linking and fraud protection — 1-2 weeks. Cost is calculated individually based on integration complexity.

What Is Included in Our Work

As part of analytics implementation, we provide:

  • Development and approval of a tracking plan with product and marketing teams.
  • SDK integration (Firebase, Amplitude, Mixpanel, AppsFlyer) considering your stack (Swift/Kotlin/Flutter/React Native).
  • Setup of funnels, cohorts, dashboards, and alerts.
  • Automation of symbolication and .dSYM upload via Fastlane.
  • Documentation of events and parameters.
  • Team training on the analytics platform.
  • Two weeks of post-release support and tracking adjustments.

Our experience: 7 years of analytics implementation and over 80 successful projects in mobile development. We guarantee data correctness and transparency at every stage.

Contact us for a consultation on setting up analytics for your app. Request an audit of your current analytics — and we will show you which metrics you are losing.