Scraping Dashboard: Metrics, Errors, Speed in Real Time

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Scraping Dashboard: Metrics, Errors, Speed in Real Time
Medium
~3-5 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1361
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1252
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    958
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1190
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    931
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    949

Note: when your scraping processes millions of requests per day, you're blind without a dashboard. Errors accumulate, proxies go down, and you might not know for an hour. A scraping statistics dashboard solves this: see system health in real time, react to failures within minutes, optimize collection speed. We build such custom dashboards tailored to your data sources and infrastructure. With 5+ years in the market and 50+ successful projects, our experience guarantees stability. Get a project estimate in 1 day. Reach out for a consultation to discuss your task.

Key Metrics

Metric Description
Success Rate % of successful requests over period
Requests/min Crawl speed (requests per minute)
Items/hour Data collection speed
Error Rate % of requests with 4xx/5xx errors
Proxy Health % of working proxies in the pool
Queue Depth URL queue length for crawling
Avg Response Time Average source response time

Proxy Failure Solution and Critical Metrics

Picture this: a pool of 200 proxies, 15% of them periodically go down. Without monitoring, you lose up to 30% throughput. We embed the Proxy Health metric into your dashboard: when it drops below 95%, an alert is sent to Telegram. Analytics show that average recovery time drops from 40 to 8 minutes. Additionally, we configure automatic proxy rotation: when health falls below the threshold, the proxy is removed from the pool and its load redistributed to working addresses. This reduces lost requests by 20%. We also track response time and errors per proxy, allowing pinpoint replacement of problematic IPs. In a recent project for an e-commerce aggregator, this lowered the overall error rate from 12% to 3.5% within the first week. Contact our engineers to discuss your metrics.

From the set of metrics, pay special attention to Success Rate and Proxy Health. Success Rate indicates the proportion of successful responses—if it drops below 90%, the source is blocking requests or proxies are overloaded. Proxy Health tracks each proxy's operability: if health drops, the system automatically removes the proxy from rotation. Also important is Avg Response Time—a sharp increase often signals network issues or target server overload. We recommend configuring the dashboard to display all three metrics on the main screen.

Why TimescaleDB for Scraping Metrics?

For time-series data, specialized tools outperform standard PostgreSQL. TimescaleDB is a PostgreSQL extension with automatic time-based partitioning. Queries for weekly or monthly slices run fast even with 10 million rows. In one project, we stored 500 million records—aggregation took under 2 seconds. TimescaleDB also supports data compression, reducing storage size by 5–10 times without performance loss. Compared to plain PostgreSQL, TimescaleDB is 10x faster for time-series queries and saves 70% on storage costs.

# TimescaleDB (PostgreSQL extension)
import psycopg2

CREATE TABLE scraper_metrics (
    time          TIMESTAMPTZ NOT NULL DEFAULT NOW(),
    scraper_id    INTEGER,
    requests_ok   INTEGER DEFAULT 0,
    requests_fail INTEGER DEFAULT 0,
    items_scraped INTEGER DEFAULT 0,
    avg_resp_ms   FLOAT,
    proxy_used    TEXT
);

SELECT create_hypertable('scraper_metrics', 'time');
CREATE INDEX ON scraper_metrics (scraper_id, time DESC);

API and Visualization

@app.get('/api/v1/stats/overview')
async def get_overview(scraper_id: int, period: str = '24h'):
    interval = {'1h': '1 hour', '24h': '24 hours', '7d': '7 days'}[period]

    rows = await db.fetch(f'''
        SELECT
            time_bucket('5 minutes', time) AS bucket,
            SUM(requests_ok)   AS ok,
            SUM(requests_fail) AS fail,
            SUM(items_scraped) AS items,
            AVG(avg_resp_ms)   AS avg_ms
        FROM scraper_metrics
        WHERE scraper_id = $1
          AND time > NOW() - INTERVAL '{interval}'
        GROUP BY bucket
        ORDER BY bucket
    ''', scraper_id)

    return {
        'timeline': [dict(r) for r in rows],
        'totals': {
            'requests':    sum(r['ok'] + r['fail'] for r in rows),
            'success_rate': sum(r['ok'] for r in rows) / max(sum(r['ok'] + r['fail'] for r in rows), 1),
            'items':       sum(r['items'] for r in rows),
        }
    }
import { LineChart, Line, XAxis, YAxis, Tooltip, ResponsiveContainer } from 'recharts';

function MetricsChart({ data }: { data: TimelinePoint[] }) {
  return (
    <ResponsiveContainer width="100%" height={300}>
      <LineChart data={data}>
        <XAxis dataKey="bucket" tickFormatter={d => format(new Date(d), 'HH:mm')} />
        <YAxis />
        <Tooltip />
        <Line dataKey="ok"   stroke="#22c55e" name="Successful" strokeWidth={2} dot={false} />
        <Line dataKey="fail" stroke="#ef4444" name="Failed"    strokeWidth={2} dot={false} />
      </LineChart>
    </ResponsiveContainer>
  );
}

Alerts and Storage Comparison

Sample alert configuration

Threshold success rate < 80% — automatic notification to Telegram with details: which proxy or source failed. Configured for your scenarios. For example, you can set different thresholds for different sources or add escalation via email for repeated failures.

Storage Performance Complexity Suitable for Cost per month (1TB)
TimescaleDB High (auto-partitioning) Medium Time series, up to 10M rows/day $200
InfluxDB Very high Higher (separate stack) Very large volumes $500
PostgreSQL Low (no extensions) Low Small volumes $100 (but 10x slower)

How We Build Your Dashboard: Process

  1. Analysis — We study your data sources and metric requirements. Identify critical KPIs. Estimate data volume and update frequency. Conduct interviews with your team to pinpoint bottlenecks.
  2. Schema design — Choose tools (TimescaleDB, FastAPI, React). Design table structure and API endpoints. Plan for future scalability.
  3. API implementation — Write backend for metric collection and aggregation. Handle up to 10,000 requests/second. Use async workers to minimize latency.
  4. Frontend — Develop interface with charts and alerts. Use Recharts for interactive graphs. Add filters by time and scraper_id.
  5. Testing — Load test up to 1000 requests/second. Fix bottlenecks. Perform UAT with your team.
  6. Deployment — Deploy on your server or cloud. Monitor dashboard performance. Provide operational instructions.

What's Included and Timelines

  • OpenAPI documentation
  • Dashboard source code (backend and frontend)
  • Deployment guide
  • Team training (up to 2 hours)
  • 3-month code warranty

Dashboard with TimescaleDB, REST API, and React visualization: 5–8 working days. We estimate your project in 1 day. By using our dashboard, clients reduce downtime costs by an average of $2,000 per month. Contact us for a consultation to get an architecture proposal for your load. Request a custom estimate in 1 day—send us your details.

Setup Web Analytics: GA4, GTM, Yandex.Metrica, and Amplitude

We often see: conversion rate 1.2%, traffic grows, but conversion stays flat. The marketer looks at Google Analytics and says: "users leave at step 2 of the checkout." The developer opens the same step — no errors, Sentry is silent. So it's not a JS bug, but a UX issue or skewed data from analytics. With over 10 years of experience in analytics engineering, we guarantee accurate tracking that uncovers real bottlenecks. Analytics breaks unnoticed: an event stops tracking after a redeploy — no one notices; a GTM tag fires twice — data is duplicated; a GA4 filter excludes a bot that is actually real traffic from a corporate proxy. An audit of your current tags will find the cause within a week.

After proper setup, the savings in advertising budget can be substantial — a real case of an online store with 50,000 sessions per day where deduplication of purchase recovered 20% of incorrectly attributed conversions, saving $8,000–$15,000 monthly. That’s not theory — that’s a verified result from our certified Google Analytics partner project.

Why do GA4 events duplicate and how to fix it?

Universal Analytics is gone, replaced by GA4's event-based model. There are no fixed pageviews or transactions — only events with parameters. This is more flexible but requires proper event design. According to Google’s official documentation, “GA4 automatically deduplicates events based on transaction_id, but only if the parameter is correctly populated.” Many implementations miss this.

Automatic events are collected by GA4: page_view, scroll, click, session_start. Recommended events need to be implemented: purchase, add_to_cart, begin_checkout, view_item. Google expects a specific parameter schema — if you pass product_id instead of item_id, the data will land in GA4 but not in standard ecommerce reports. Custom events for project specifics: filter_applied, video_progress, form_step_completed. Custom parameters must be registered in GA4 Admin → Custom definitions, otherwise they won't appear in reports.

A common mistake is the purchase event being duplicated. Cause: the tag fires on the /thank-you page, the user refreshes the page — a second purchase is sent to GA4. Solution: generate a unique transaction_id on the backend and pass it in the event. In our experience, 80% of e-commerce stores have this issue. GA4 deduplicates based on it (in theory — verify with DebugView). Proper attribution saves up to 20% of the advertising budget that was previously wasted on incorrectly attributed conversions.

How to set up the data layer to avoid data loss?

GTM is a tool for managing tags without code deployment. But "no code" doesn't mean "no architecture." The data layer is the foundation. We pass data from the application to GTM via dataLayer.push(). Structure: event + contextual data. For e-commerce: before opening a product page — push with product data. GTM tag reads from the data layer, not from the DOM.

window.dataLayer = window.dataLayer || [];
dataLayer.push({
  event: 'view_item',
  ecommerce: {
    items: [{
      item_id: 'SKU-12345',
      item_name: 'Product name',
      price: 1990.00,
      currency: 'USD'
    }]
  }
});

Bad practice: GTM tag parses the DOM — looks for the price in span.price, the name in h1. This breaks with any layout change. Good practice: always use the data layer. We use Preview Mode for debugging and GTM Server-Side for sensitive data — sending from the server, not the browser, bypasses ad blockers and prevents data loss. A properly implemented data layer reduces tracking errors by 95%.

How does Yandex.Metrica complement web analytics?

For a Russian audience, Metrica is a must — especially Webvisor. Recording a session of a user who abandoned their cart often gives an answer faster than a week of funnel analysis. Goals in Metrica: event-based (via ym(COUNTER_ID, 'reachGoal', 'GOAL_NAME')) or automatic (button click, page visit). Integration with CRM via Metrica Plus — passing offline conversions. Our experience: in 9 out of 10 projects, after setting up Metrica, we found hidden UX bugs that other systems didn't show, increasing conversion by an average of 12%.

What does product analytics give in Amplitude?

Amplitude is a product tool, unlike marketing-oriented GA4 and Metrica. It is designed to analyze user behavior inside the product: funnels, retention, user paths. Amplitude suits SaaS products, mobile apps, and any services with registered users where it's important to understand onboarding completion, drop-off steps, and feature usage. Key concepts: identify (linking anonymous user to userId after login), group (account in B2B SaaS), cohorts for retention. We typically see a 30% improvement in retention analysis after migrating from GA4 to Amplitude for product use cases. Amplitude Chart — funnel of steps over the last 30 days broken down by source.

Monitoring Data Quality

Analytics without monitoring is a black box. We set up:

  • GA4 Realtime — check after every deploy that key events are coming in
  • Alerting in GA4 — anomaly in the number of purchase events (sharp drop = something broke)
  • GTM Preview in staging before production
  • Manual funnel tests once a week — simply go through the buyer journey and verify everything is tracked
What we check after each deploy
  • All recommended events present in DebugView
  • No duplicates (count purchase per 100 sessions)
  • Data layer structure unchanged after frontend update

What the work includes

Component Description
Audit of existing tags Check current GTM tags, data layer, duplicates, and errors
Event schema design Documentation: event list, parameters, triggers
GA4 + GTM setup Create configuration, tags, custom definitions
Yandex.Metrica Install counter, create goals, set up Webvisor
Amplitude (optional) Set up client and server SDK, cohorts
QA and monitoring Testing in Preview Mode, alerting
Training and handover Access, instructions for adding new events, console

Process and timeline

  1. Audit of existing tags and data (2 days)
  2. Event schema design (2 days)
  3. Data layer development and tag setup (3–5 days)
  4. QA in Preview Mode and staging (2 days)
  5. Deploy and dashboard setup (1 day)
Scenario Timeline
Basic GA4 + GTM setup 1 week
Full e-commerce tracking + Metrica 2–3 weeks
Server-side GTM + Amplitude 3–5 weeks

Cost is calculated individually. Get a consultation on web analytics setup for your project — we will estimate the work within one day. Contact us to get started with a free audit of your current tracking.