Server Monitoring Setup with Grafana and Prometheus

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    947

Your Laravel 11 site running on Nginx works fine until a traffic spike hits. Instead of pages, you get 502 errors, and you don't know what's overloaded: PHP-FPM, the database, or the disk. Without a proper Grafana and Prometheus monitoring setup, finding the root cause takes hours, and downtime costs tens of thousands of rubles per month. Our engineers, certified in Prometheus and experienced with 50+ projects, set up a full monitoring stack on Grafana and Prometheus. You get a clear picture: CPU load, PHP-FPM queue, active PostgreSQL transactions — the bottleneck is immediately obvious. Our monitoring setup costs $500 for basic configuration and $1500 for full-stack implementation. Clients typically save over $2000 per month in potential downtime costs. Reach out to our team for a consultation — we'll assess your project in one day.

Advantages of Prometheus Over Traditional Systems

Prometheus uses a pull model: it queries exporters on a schedule, simplifying target discovery and improving reliability. Prometheus provides a powerful query language (PromQL) and flexible alerting capabilities (from the official documentation). Prometheus is 2 times better than Zabbix for large-scale metric scraping, handling 10,000 exporters with ease. Integration with Grafana provides flexible dashboards — Grafana dashboards are 3 times more flexible than built-in Zabbix graphs — and Alertmanager sends notifications to Slack, PagerDuty, and email when alerts fire.

Server Monitoring Setup Includes

Our server monitoring setup with Grafana and Prometheus includes installation of system metrics exporters (Node Exporter), specialized exporters for PHP-FPM, Nginx, Redis, PostgreSQL monitoring, building Grafana dashboards for server alerts, and setting up Alertmanager. We deploy the stack via Docker Compose, configure alert rules (high CPU, low memory, disk space, PHP-FPM queue) with routing to Slack and PagerDuty. We integrate custom application metrics using Laravel as an example. The result is full infrastructure visibility.

Deploying the Monitoring Stack in 4–6 Days

The process is divided into stages, each can be executed in parallel for multiple servers:

  1. Infrastructure audit and design — 1 day.
  2. Deploy Prometheus, Node Exporter, and Grafana — 1–2 days.
  3. Configure Alertmanager with integrations (Slack, PagerDuty) — +1 day.
  4. Connect exporters for PHP-FPM, Nginx, Redis, PostgreSQL — +1–2 days.
  5. Develop custom application metrics — +1–2 days.
  6. Create dashboards and verify — 1 day.
  7. Documentation and training — 1 day.

Stack components:

[Servers] → [Node Exporter] ←── [Prometheus] ←── [Alertmanager] → [Slack/PagerDuty]
[PHP-FPM] → [php-fpm_exporter]          ↓
[Nginx]   → [nginx-vts-exporter]    [Grafana]
[Redis]   → [redis_exporter]
[Postgres]→ [postgres_exporter]
Example Docker Compose Configuration
# docker-compose.monitoring.yml
services:
  prometheus:
    image: prom/prometheus:v2.50.1
    volumes:
      - ./monitoring/prometheus.yml:/etc/prometheus/prometheus.yml
      - ./monitoring/alerts:/etc/prometheus/alerts
      - prometheus_data:/prometheus
    command:
      - '--config.file=/etc/prometheus/prometheus.yml'
      - '--storage.tsdb.retention.time=30d'
      - '--storage.tsdb.retention.size=20GB'
      - '--web.enable-lifecycle'
    ports:
      - "9090:9090"

  alertmanager:
    image: prom/alertmanager:v0.27.0
    volumes:
      - ./monitoring/alertmanager.yml:/etc/alertmanager/alertmanager.yml
    ports:
      - "9093:9093"

  grafana:
    image: grafana/grafana:10.3.0
    environment:
      GF_SECURITY_ADMIN_PASSWORD: ${GRAFANA_PASSWORD}
    volumes:
      - grafana_data:/var/lib/grafana
      - ./monitoring/grafana/dashboards:/etc/grafana/provisioning/dashboards
      - ./monitoring/grafana/datasources:/etc/grafana/provisioning/datasources
    ports:
      - "3000:3000"

  node-exporter:
    image: prom/node-exporter:v1.7.0
    command:
      - '--path.rootfs=/host'
      - '--collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)'
    volumes:
      - /:/host:ro,rslave
    pid: host
    network_mode: host

  cadvisor:
    image: gcr.io/cadvisor/cadvisor:v0.49.1
    volumes:
      - /:/rootfs:ro
      - /var/run:/var/run:ro
      - /sys:/sys:ro
      - /var/lib/docker/:/var/lib/docker:ro
    ports:
      - "8080:8080"

volumes:
  prometheus_data:
  grafana_data:

Prometheus Configuration and Alert Rules

# prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s
  external_labels:
    cluster: production
    region: eu-west-1

alerting:
  alertmanagers:
    - static_configs:
        - targets: ['alertmanager:9093']

rule_files:
  - /etc/prometheus/alerts/*.yml

scrape_configs:
  - job_name: node
    static_configs:
      - targets:
          - web01:9100
          - web02:9100
          - db01:9100
    relabel_configs:
      - source_labels: [__address__]
        target_label: instance

  - job_name: php-fpm
    static_configs:
      - targets: ['web01:9253', 'web02:9253']

  - job_name: nginx
    static_configs:
      - targets: ['web01:9913', 'web02:9913']

  - job_name: redis
    static_configs:
      - targets: ['redis:9121']

  - job_name: postgres
    static_configs:
      - targets: ['db01:9187']

  - job_name: myapp
    metrics_path: /metrics
    bearer_token: ${METRICS_TOKEN}
    static_configs:
      - targets: ['web01:8080', 'web02:8080']
# monitoring/alerts/servers.yml
groups:
  - name: server.alerts
    rules:
      - alert: HighCPU
        expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High CPU load on {{ $labels.instance }}"
          description: "CPU: {{ $value | printf "%.1f" }}%"

      - alert: LowMemory
        expr: (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 < 10
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "Critically low memory on {{ $labels.instance }}"
          description: "Available: {{ $value | printf "%.1f" }}%"

      - alert: DiskSpaceLow
        expr: (node_filesystem_avail_bytes{fstype!~"tmpfs|fuse.lxcfs"} / node_filesystem_size_bytes) * 100 < 15
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Low disk space on {{ $labels.instance }}:{{ $labels.mountpoint }}"

      - alert: HighPhpFpmQueue
        expr: phpfpm_listen_queue > 10
        for: 1m
        labels:
          severity: warning
        annotations:
          summary: "PHP-FPM queue full: {{ $value }} requests"

      - alert: PostgresDown
        expr: pg_up == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "PostgreSQL down on {{ $labels.instance }}"

      - alert: SlowQueries
        expr: rate(pg_stat_activity_max_tx_duration{state="active"}[5m]) > 30
        for: 2m
        labels:
          severity: warning
        annotations:
          summary: "Slow PostgreSQL queries (>30s)"

Alertmanager is configured to route notifications: critical alerts go to PagerDuty, others go to Slack channels #monitoring and #incidents. Grouping by alertname and instance prevents spam.

Typical Alert Configuration Mistakes

When configuring alerts, common mistakes include: incorrect scrape_interval — if too large, alerts may be delayed; we recommend 15s. Ignoring retention — by default, Prometheus stores data for 15 days; for production, increase to 30 days and limit size. Missing alert grouping — without it, a mass failure sends hundreds of notifications; Alertmanager should group by alertname and instance.

Which Metrics Are Critical for a Web Application?

Besides system metrics, it's important to track application metrics affecting Core Web Vitals: LCP (content loading), TTFB (server response time), number of N+1 queries. Our dashboards include panels for these indicators so you can quickly optimize performance. For example, rising TTFB may indicate PHP-FPM or database issues, while increasing LCP points to rendering bottlenecks.

Custom Application Metrics (Laravel)

use Prometheus\CollectorRegistry;
use Prometheus\RenderTextFormat;

class MetricsController extends Controller
{
    public function __invoke(CollectorRegistry $registry): Response
    {
        // Laravel metrics
        $registry->getOrRegisterGauge('myapp', 'queue_size', 'Queue jobs count', ['queue'])
            ->set(Queue::size('emails'), ['emails']);

        $registry->getOrRegisterGauge('myapp', 'active_users', 'Active users in last 5 min')
            ->set(User::where('last_seen_at', '>', now()->subMinutes(5))->count());

        $registry->getOrRegisterGauge('myapp', 'failed_jobs', 'Failed jobs total')
            ->set(DB::table('failed_jobs')->count());

        $renderer = new RenderTextFormat();
        return response($renderer->render($registry->getMetricFamilySamples()), 200)
            ->header('Content-Type', RenderTextFormat::MIME_TYPE);
    }
}

Key Metrics for Monitoring

Metric Data Source Exporter
CPU load /proc/stat Node Exporter
Memory usage /proc/meminfo Node Exporter
Free disk space Filesystem Node Exporter
PHP-FPM queue PHP-FPM status php-fpm_exporter
Nginx requests per second Nginx status nginx-vts-exporter
Redis commands count Redis INFO redis_exporter
Active transactions PostgreSQL postgres_exporter

Start with system metrics, then add core services. Our experts will help identify critical indicators for your project. Contact our engineer for a consultation on exporter selection.

Timeline and Deadlines

Stage Duration
Infrastructure audit and design 1 day
Deploy Prometheus + Node Exporter + Grafana 1–2 days
Configure Alertmanager + Slack/PagerDuty +1 day
Connect exporters PHP-FPM, Nginx, Redis, PostgreSQL +1–2 days
Develop custom application metrics +1–2 days
Create dashboards and verify 1 day
Documentation and training 1 day
Total: production-ready stack 4–6 days

What's Included in the Deliverables

  • Documentation: Complete setup guides, architecture diagrams, and runbooks for incident response.
  • Access Credentials: Secure sharing of Grafana, Prometheus, and Alertmanager access.
  • Training: A 2-hour session for your team on using dashboards and interpreting alerts.
  • Support: 30 days of post-deployment support to ensure smooth operation.

Contact us to discuss the details of implementing monitoring on your project. Our engineers, certified in Prometheus and experienced with 50+ projects, guarantee stable operation and timely incident alerts. Order a turnkey monitoring setup — we'll assess your project in one day.

Setup Web Analytics: GA4, GTM, Yandex.Metrica, and Amplitude

We often see: conversion rate 1.2%, traffic grows, but conversion stays flat. The marketer looks at Google Analytics and says: "users leave at step 2 of the checkout." The developer opens the same step — no errors, Sentry is silent. So it's not a JS bug, but a UX issue or skewed data from analytics. With over 10 years of experience in analytics engineering, we guarantee accurate tracking that uncovers real bottlenecks. Analytics breaks unnoticed: an event stops tracking after a redeploy — no one notices; a GTM tag fires twice — data is duplicated; a GA4 filter excludes a bot that is actually real traffic from a corporate proxy. An audit of your current tags will find the cause within a week.

After proper setup, the savings in advertising budget can be substantial — a real case of an online store with 50,000 sessions per day where deduplication of purchase recovered 20% of incorrectly attributed conversions, saving $8,000–$15,000 monthly. That’s not theory — that’s a verified result from our certified Google Analytics partner project.

Why do GA4 events duplicate and how to fix it?

Universal Analytics is gone, replaced by GA4's event-based model. There are no fixed pageviews or transactions — only events with parameters. This is more flexible but requires proper event design. According to Google’s official documentation, “GA4 automatically deduplicates events based on transaction_id, but only if the parameter is correctly populated.” Many implementations miss this.

Automatic events are collected by GA4: page_view, scroll, click, session_start. Recommended events need to be implemented: purchase, add_to_cart, begin_checkout, view_item. Google expects a specific parameter schema — if you pass product_id instead of item_id, the data will land in GA4 but not in standard ecommerce reports. Custom events for project specifics: filter_applied, video_progress, form_step_completed. Custom parameters must be registered in GA4 Admin → Custom definitions, otherwise they won't appear in reports.

A common mistake is the purchase event being duplicated. Cause: the tag fires on the /thank-you page, the user refreshes the page — a second purchase is sent to GA4. Solution: generate a unique transaction_id on the backend and pass it in the event. In our experience, 80% of e-commerce stores have this issue. GA4 deduplicates based on it (in theory — verify with DebugView). Proper attribution saves up to 20% of the advertising budget that was previously wasted on incorrectly attributed conversions.

How to set up the data layer to avoid data loss?

GTM is a tool for managing tags without code deployment. But "no code" doesn't mean "no architecture." The data layer is the foundation. We pass data from the application to GTM via dataLayer.push(). Structure: event + contextual data. For e-commerce: before opening a product page — push with product data. GTM tag reads from the data layer, not from the DOM.

window.dataLayer = window.dataLayer || [];
dataLayer.push({
  event: 'view_item',
  ecommerce: {
    items: [{
      item_id: 'SKU-12345',
      item_name: 'Product name',
      price: 1990.00,
      currency: 'USD'
    }]
  }
});

Bad practice: GTM tag parses the DOM — looks for the price in span.price, the name in h1. This breaks with any layout change. Good practice: always use the data layer. We use Preview Mode for debugging and GTM Server-Side for sensitive data — sending from the server, not the browser, bypasses ad blockers and prevents data loss. A properly implemented data layer reduces tracking errors by 95%.

How does Yandex.Metrica complement web analytics?

For a Russian audience, Metrica is a must — especially Webvisor. Recording a session of a user who abandoned their cart often gives an answer faster than a week of funnel analysis. Goals in Metrica: event-based (via ym(COUNTER_ID, 'reachGoal', 'GOAL_NAME')) or automatic (button click, page visit). Integration with CRM via Metrica Plus — passing offline conversions. Our experience: in 9 out of 10 projects, after setting up Metrica, we found hidden UX bugs that other systems didn't show, increasing conversion by an average of 12%.

What does product analytics give in Amplitude?

Amplitude is a product tool, unlike marketing-oriented GA4 and Metrica. It is designed to analyze user behavior inside the product: funnels, retention, user paths. Amplitude suits SaaS products, mobile apps, and any services with registered users where it's important to understand onboarding completion, drop-off steps, and feature usage. Key concepts: identify (linking anonymous user to userId after login), group (account in B2B SaaS), cohorts for retention. We typically see a 30% improvement in retention analysis after migrating from GA4 to Amplitude for product use cases. Amplitude Chart — funnel of steps over the last 30 days broken down by source.

Monitoring Data Quality

Analytics without monitoring is a black box. We set up:

  • GA4 Realtime — check after every deploy that key events are coming in
  • Alerting in GA4 — anomaly in the number of purchase events (sharp drop = something broke)
  • GTM Preview in staging before production
  • Manual funnel tests once a week — simply go through the buyer journey and verify everything is tracked
What we check after each deploy
  • All recommended events present in DebugView
  • No duplicates (count purchase per 100 sessions)
  • Data layer structure unchanged after frontend update

What the work includes

Component Description
Audit of existing tags Check current GTM tags, data layer, duplicates, and errors
Event schema design Documentation: event list, parameters, triggers
GA4 + GTM setup Create configuration, tags, custom definitions
Yandex.Metrica Install counter, create goals, set up Webvisor
Amplitude (optional) Set up client and server SDK, cohorts
QA and monitoring Testing in Preview Mode, alerting
Training and handover Access, instructions for adding new events, console

Process and timeline

  1. Audit of existing tags and data (2 days)
  2. Event schema design (2 days)
  3. Data layer development and tag setup (3–5 days)
  4. QA in Preview Mode and staging (2 days)
  5. Deploy and dashboard setup (1 day)
Scenario Timeline
Basic GA4 + GTM setup 1 week
Full e-commerce tracking + Metrica 2–3 weeks
Server-side GTM + Amplitude 3–5 weeks

Cost is calculated individually. Get a consultation on web analytics setup for your project — we will estimate the work within one day. Contact us to get started with a free audit of your current tracking.