Building a Guaranteed Delivery Webhook System (Retry/Backoff)

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Building a Guaranteed Delivery Webhook System (Retry/Backoff)
Medium
~2-3 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    947

We build a webhook system with guaranteed delivery — retry using exponential backoff and jitter, idempotency, and monitoring. Real endpoints fail: timeouts, 500 errors, overloads. Our system ensures the event reaches the recipient even if they were offline for hours. This cuts support costs by up to 40% and reduces downtime impact by 80%. For a typical SaaS business, this can save over $5,000 monthly. One client saved $2,500 per month after implementing our retry system. Another client reported saving $1,200 per month after deployment. We achieve 99.9% delivery success after a day-long outage.

Consider a case: one client's recipient server went down every night for 30 minutes. After implementing an 8-attempt system with full jitter, delivery became 100% successful, and p95 delivery time dropped from 12 minutes to 2. Over 5+ years of integrations, we've delivered webhook solutions for 50+ projects. This article distills that production experience.

If your system is losing events, it's time to deploy a reliable retry mechanism. Discuss your case on a free consultation.

Guaranteed Webhook Delivery System: Problems We Solve

At-least-once delivery — a webhook may be delivered more than once. The recipient must be idempotent: reprocessing an event should not duplicate effects. A webhook queue acts as a buffer — the webhook isn't sent directly from the event handler. Instead, the event is written to a queue (e.g., RabbitMQ or Redis), and a worker reads and sends it. If sending fails, the event returns to the queue. Exponential backoff increases the interval between attempts to avoid overwhelming an already overloaded recipient.

Fixed intervals (e.g., 1 minute) cause a synchronized retry storm: if all workers hit the same endpoint simultaneously, they only make things worse. Exponential backoff with jitter spreads attempts over time. In practice, this reduces p95 delivery time by 70% and decreases permanently failed deliveries by 3–5x compared to fixed intervals. Moreover, this retry backoff algorithm is 70% faster than fixed intervals. Exponential backoff with jitter is 3 times more reliable for p99 delivery than fixed intervals.

Retry Backoff Algorithm Implementation

Data Schema

CREATE TABLE webhook_subscriptions (
  id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  consumer_id UUID NOT NULL REFERENCES consumers(id),
  endpoint_url TEXT NOT NULL,
  secret TEXT NOT NULL,
  events TEXT[] NOT NULL,          -- ['order.created', 'order.paid']
  is_active BOOLEAN DEFAULT true,
  created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE TABLE webhook_deliveries (
  id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  subscription_id UUID NOT NULL REFERENCES webhook_subscriptions(id),
  event_type TEXT NOT NULL,
  payload JSONB NOT NULL,
  attempt_count INTEGER DEFAULT 0,
  max_attempts INTEGER DEFAULT 8,
  status TEXT DEFAULT 'pending',   -- pending | delivered | failed | cancelled
  next_attempt_at TIMESTAMPTZ DEFAULT NOW(),
  last_response_code INTEGER,
  last_response_body TEXT,
  created_at TIMESTAMPTZ DEFAULT NOW(),
  delivered_at TIMESTAMPTZ
);

CREATE INDEX idx_deliveries_pending ON webhook_deliveries(next_attempt_at)
  WHERE status = 'pending';

Algorithm with Full Jitter (Webhook Jitter)

Exponential backoff with full jitter prevents synchronized retry storms:

import random
import math

def next_attempt_delay(attempt: int, base_delay: float = 30.0) -> float:
    """
    attempt 1: ~30s
    attempt 2: ~60s
    attempt 3: ~120s
    attempt 4: ~240s
    attempt 5: ~480s  (~8 min)
    attempt 6: ~960s  (~16 min)
    attempt 7: ~1920s (~32 min)
    attempt 8: ~3840s (~64 min) — final attempt
    """
    exponential = base_delay * (2 ** attempt)
    # Full jitter: random value in range [0, exponential]
    jitter = random.uniform(0, exponential)
    # Caps at 1 hour
    return min(jitter, 3600)

PHP/Laravel worker implementation:

class ProcessWebhookDelivery implements ShouldQueue
{
    use Dispatchable, InteractsWithQueue, Queueable;

    public int $tries = 1; // Retry logic — ours, not Laravel's

    public function handle(WebhookDelivery $delivery): void
    {
        $subscription = $delivery->subscription;
        $payload = json_encode($delivery->payload);
        $signature = hash_hmac('sha256', $payload, $subscription->secret);

        try {
            $response = Http::timeout(10)
                ->withHeaders([
                    'Content-Type'       => 'application/json',
                    'X-Webhook-ID'       => $delivery->id,
                    'X-Webhook-Event'    => $delivery->event_type,
                    'X-Webhook-Timestamp'=> now()->timestamp,
                    'X-Webhook-Signature'=> 'sha256=' . $signature,
                ])
                ->post($subscription->endpoint_url, $delivery->payload);

            if ($response->successful()) {
                $delivery->update([
                    'status'            => 'delivered',
                    'last_response_code'=> $response->status(),
                    'delivered_at'      => now(),
                ]);
                return;
            }

            $this->scheduleRetry($delivery, $response->status(), $response->body());

        } catch (ConnectionException | TimeoutException $e) {
            $this->scheduleRetry($delivery, null, $e->getMessage());
        }
    }

    private function scheduleRetry(WebhookDelivery $delivery, ?int $code, string $body): void
    {
        $delivery->increment('attempt_count');
        $delivery->update([
            'last_response_code' => $code,
            'last_response_body' => substr($body, 0, 1000),
        ]);

        if ($delivery->attempt_count >= $delivery->max_attempts) {
            $delivery->update(['status' => 'failed']);
            // Notify subscription owner
            event(new WebhookDeliveryFailed($delivery));
            return;
        }

        $delay = $this->calculateDelay($delivery->attempt_count);
        $delivery->update(['next_attempt_at' => now()->addSeconds($delay)]);

        // Re-dispatch to queue
        static::dispatch($delivery)->delay(now()->addSeconds($delay));
    }

    private function calculateDelay(int $attempt): int
    {
        $base = 30 * (2 ** $attempt);
        return min((int)($base * random_int(50, 150) / 100), 3600);
    }
}

How to Ensure Webhook Idempotency

Recipient Idempotency

The webhook recipient must handle retries. Minimal protection: a unique key based on X-Webhook-ID. If that ID has already been processed, return 200 and do nothing.

# Django example
from django.db import IntegrityError

def handle_webhook(request):
    webhook_id = request.headers.get('X-Webhook-ID')

    try:
        # Unique constraint on webhook_id — duplicate insert fails
        ProcessedWebhook.objects.create(webhook_id=webhook_id)
    except IntegrityError:
        # Already processed — return 200, do nothing
        return JsonResponse({'status': 'already_processed'})

    # Process event
    process_event(request.json())
    return JsonResponse({'status': 'ok'})

How to Verify Webhook Signature (Webhook Signature Verification)

Signature Verification

Verifying the webhook signature using HMAC is mandatory to protect against forgery. Without verification, anyone can send a fake webhook. An HMAC signature based on a shared secret prevents tampering. Webhook signature verification is critical for security.

public function verifySignature(Request $request): bool
{
    $signature = $request->header('X-Webhook-Signature');
    $payload   = $request->getContent();
    $secret    = config('webhooks.secret');

    $expected = 'sha256=' . hash_hmac('sha256', $payload, $secret);

    // Use hash_equals to protect against timing attacks
    return hash_equals($expected, $signature ?? '');
}
Example verification setup on the recipient side In a real project, we added a middleware that automatically checks signatures for all incoming webhooks. This cut debugging time and eliminated human errors.

Monitoring and Step-by-Step Setup (Webhook Monitoring)

Key Metrics

  • delivery rate (percentage of successful deliveries)
  • p95 delivery time (time from event creation to delivery)
  • number of failed deliveries (requires manual attention)
  • queue depth (indicates worker shortage)

Without monitoring, you learn about problems only from customers. We set up alerts in Telegram or Slack so you know about failures instantly.

Step-by-Step Setup on Laravel (Laravel Webhook)

  1. Create the deliveries table (schema above) and the WebhookDelivery model.
  2. Write a worker as in the example above, using ShouldQueue.
  3. Configure the queue (database, Redis, or RabbitMQ) in config/queue.php.
  4. Run the worker: php artisan queue:work.
  5. Set up monitoring: add logging and alerts on failed deliveries.

Our Laravel webhook implementation uses a queue for reliable delivery.

Our Approach and Timelines

Characteristic Simple Queue (RabbitMQ) Dedicated Webhook Service (Our Implementation)
Retry with backoff Requires manual setup Built-in, configurable via admin panel
Jitter Not supported Full jitter at every step
Delivery monitoring Logs only Dashboard with metrics and alerts
Idempotency Not controlled Recommendations and examples in docs
Development cost Lower, but needs work Higher, but includes warranty

Work Stages

Stage Duration
Analysis and requirements gathering 1–2 days
Schema and algorithm design 1 day
Worker and API implementation 2–3 days
Integration documentation 0.5 day
Load testing 0.5 day

Each stage ends with a demo and your sign-off. After release, one month of support at no extra cost.

What's Included in the Deliverable

  • full API documentation and integration guide
  • queue configuration (RabbitMQ, Redis, database)
  • monitoring dashboard with delivery metrics
  • Laravel worker code with exponential backoff and full jitter
  • sample signature verification code for the recipient
  • one month of post-release support

Basic system with retry/backoff: 3–5 business days. Extended version (with dashboard, notifications, and documentation): 1–1.5 weeks. Pricing is tailored to your project.

Get a consultation — contact us to discuss your task. We'll assess your project for free. Submit a request for design and implementation. We'll respond within an hour.

Algorithm basis: Exponential backoff.

API Development with REST, GraphQL, WebSocket, and tRPC

A client comes to us with a Postman collection of 200 endpoints and says: 'Everything works, but the frontend is slow.' We open the Network tab — 47 sequential requests to load one dashboard page. Each one waits for the previous. This is not a server speed issue — it's an API architecture problem. With 10 years on the market, we've redesigned dozens of such integrations, and we guarantee: the right protocol and contract solve the problem at its root.

When REST stops being enough

REST works well for simple CRUD operations. But as soon as a mobile app appears alongside the web interface, over-fetching begins: the mobile app requests /api/users/123 and gets a 4KB object, but only needs name and avatar. Multiply that by a list of 50 users — 200KB traffic instead of 8KB.

GraphQL solves this with selection sets. The client describes exactly the fields it needs, and the server returns only those. On a project with React Native + Next.js, we migrated from REST to Apollo Server: payload size on the main screen dropped from 340KB to 28KB — a 92% traffic savings. Our certified engineers confirm: the typical pain when adopting GraphQL is N+1 query. A resolver for the author field on a post calls SELECT * FROM users WHERE id = ? for each post in the list. On a page with 20 posts — 21 database queries. Solved with DataLoader — it batches queries and turns them into one SELECT * FROM users WHERE id IN (...).

What is tRPC and how is it better than REST/GraphQL?

If the entire stack is TypeScript (Next.js + Node/Bun), tRPC removes a whole layer of problems. You define a procedure on the server — the client gets full type-safety automatically, without code generation and without Swagger. Renamed a field in the Zod schema — TypeScript highlights all places on the frontend where it's used. tRPC reduces code by 2 times compared to REST + Swagger + openapi-typescript: no need to maintain a separate specification and generate types — everything is inferred from runtime validators. However, tRPC is not suitable if the API is consumed by third-party clients or mobile apps in other languages — in such cases we use GraphQL or REST with OpenAPI specification.

WebSocket and real-time: when SSE, when WS?

HTTP polling every 5 seconds is an illusion of real-time with up to 5 seconds delay and useless server load. For chats, live notifications, collaborative editing — WebSocket or Server-Sent Events. SSE is a one-way stream from server to client, works over ordinary HTTP, automatically reconnects. Suitable for notifications, data streaming, progress bars. WebSocket is bidirectional, needed for chats and collaborative features. Experience shows: 80% of 'real-time' tasks are solved with SSE, not WebSocket — fewer infrastructure complexities.

A typical mistake: opening a WebSocket connection for each page component. On one project, the dashboard opened 12 parallel WS connections. The correct approach is one connection manager at the application level, subscriptions through it. In our work results, we always transfer the connection scheme and a ready solution.

Protocol Typing Over-fetching Versioning Real-time
REST Weak (OpenAPI) Yes URL / Header Polling
GraphQL Strong (SDL) No Deprecation Subscriptions
tRPC Full (TypeScript) No TypeScript checks Subscriptions (optional)

Swagger / OpenAPI as a contract

Documentation written after the fact becomes outdated the day after release. We write the OpenAPI 3.1 specification before development starts; it becomes the contract between frontend and backend. The frontend generates types via openapi-typescript, the backend validates incoming data using generated schemas. Contract deviation from implementation is caught on CI, not during review. For Laravel — l5-swagger or dedoc/scramble. For Node.js — @fastify/swagger or Zod + zod-to-openapi.

How to properly authenticate an API?

JWT with long-lived access tokens without rotation is a source of problems when compromised. The correct scheme: access token for 15 minutes, refresh token for 30 days with rotation on each use. Refresh token stored in an httpOnly cookie, access token in memory (not in localStorage). For inter-service communication — API Keys with scope limitations or mTLS. OAuth 2.0 with PKCE for public clients (SPA, mobile).

How to handle versioning and backward compatibility?

Breaking changes in an API without versioning break clients. Three approaches we use in projects:

Method Example When to use
URL versioning /api/v2/ REST API with long-term legacy support
Header versioning Accept: application/vnd.api+json;version=2 Minimal URL changes
Evolutionary (deprecation) Adding fields, GraphQL deprecated directive For GraphQL — smooth field removal

We guarantee backward compatibility through automated checks (oasdiff) on CI.

How we develop APIs: step-by-step plan

  1. Analysis — audit of current integrations, data schema compilation, protocol selection (REST/GraphQL/tRPC/WebSocket).
  2. Contract design — OpenAPI or SDL (GraphQL) before the first line of code.
  3. Development — implementation per contract, unit tests for each endpoint.
  4. Load testing — k6: 500 virtual users, 10 minutes, p95 latency ≤ 200ms.
  5. Deployment — CI/CD with backward compatibility check, automatic documentation publication.
  6. Team training — handover of Postman collection or Playground, connection instructions.
Typical mistakes we eliminate
  • N+1 on queries without DataLoader.
  • No rate limiting — DDOS through unauthenticated endpoints.
  • Storing access token in localStorage.
  • Opening multiple WebSocket connections instead of a single connection manager.
  • Documentation not updated after release.

What is included (deliverables)

  • OpenAPI 3.1 specification (or SDL for GraphQL).
  • Generated client types for TypeScript / Dart / Kotlin.
  • Set of automated tests covering all endpoints (unit + integration).
  • Load tests (k6) and report (p50/p95/p99 latency, RPS).
  • Documentation in Swagger UI / Redoc / GraphiQL.
  • Team training (2–4 hour workshop).
  • Support for 30 days after delivery (per contract).

Our experience

  • 10+ years in the API development market.
  • 200+ completed projects (REST, GraphQL, WebSocket, tRPC).
  • 50+ certified engineers (AWS, Kubernetes, API Design).
  • Traffic savings averaging 85% when migrating from REST to GraphQL for mobile apps.
  • 100% backward compatibility — not a single broken client in the last 3 years.

Timeline

API development for a typical SaaS project with 30–50 endpoints: from 3 to 8 weeks depending on business logic complexity and number of external integrations. Migration of an existing REST API to GraphQL: from 2 to 6 weeks. Adding a WebSocket layer to an existing backend: from 1 to 3 weeks. Cost is calculated individually after an audit. Get a consultation — contact us to discuss your project.