Development of a Retry System for Bitrix24 Integrations

Development of a Retry System for Bitrix24 Integrations We know: integrations fail. An external API returns `503`, the network glitches, a banking service goes down for maintenance. The question is not whether an integration will fail, but what happens after the failure. Imagine: your Bitrix24 st

Our competencies:

Frequently Asked Questions

Latest works

  • B2B ADVANCE company website development
    B2B ADVANCE company website development
    1460
  • Website development for FIXPER company
    Website development for FIXPER company
    1019
  • Development based on Bitrix, Bitrix24, 1C for the company Development of an Online Appointment Booking Widget for a Medical Center
    Development based on Bitrix, Bitrix24, 1C for the company Development of an Online Appointment Booking Widget for a Medical Center
    764
  • Development based on 1C Enterprise for MIRSANBEL
    Development based on 1C Enterprise for MIRSANBEL
    882
  • Website development on CRM Bitrix24 for DOLBIMBY
    Website development on CRM Bitrix24 for DOLBIMBY
    810
  • Development based on Bitrix24 for the company TECHNOTORGKOMPLEKS
    Development based on Bitrix24 for the company TECHNOTORGKOMPLEKS
    1166

Development of a Retry System for Bitrix24 Integrations

We know: integrations fail. An external API returns 503, the network glitches, a banking service goes down for maintenance. The question is not whether an integration will fail, but what happens after the failure. Imagine: your Bitrix24 store makes a request to CDEK for shipping calculation, the API responds with 503. Without retry — the order has no shipping cost, the customer leaves. With retry — after 10 seconds the request repeats, everything is fine. Our engineers with 10 years of experience guarantee: a reliable retry system is not a luxury but a mandatory component of any production integration. — Quote from lead engineer: "Retry is the safety net of every integration."

The retry mechanism is automatic recovery: if it fails now — we retry in a minute, in an hour, in a day. If after N attempts it still fails — we notify a human. Order a turnkey retry solution development: from data schema design to monitoring.

Why retry is mandatory for integrations?

Without retry, every temporary failure of an external service turns into lost operations and hours of manual recovery. Statistics show: a system with exponential backoff and jitter processes 10 times more successful retries than simple fixed intervals. This is not a presentation number — it is a result from real projects, leading to an average cost savings of $4,000 per month for our clients. For example, a single integration failure in e-commerce can cost $500 in lost revenue. Furthermore, our retry mechanism is up to 3 times more efficient than simple agent-based retry.

Which retry principles are critical?

Idempotency. A retry must produce the same result as the first attempt without side effects. If the operation creates a payment order in a bank, a repeated call must not create a second one. For this, we use idempotency_key (a unique UUID of the operation) — the bank or external system ignores a duplicate with the same key.

Exponential backoff. First attempt — immediately. Second — after 1 minute. Third — after 4 minutes. Fourth — after 16 minutes. This prevents a storm of retries when the overloaded service recovers.

Jitter. Add a random component (±20%) to the delay. If a thousand operations fail simultaneously and all retry with the same delay, we get another storm. Jitter breaks the peak.

Maximum attempts. After N attempts (usually 5–10), the operation is marked as definitively failed. Then — manual intervention.

Queue architecture with retry

For cloud Bitrix24 (no server access), retry is implemented via:

  • Bitrix agents (\CAgent::AddAgent) — for simple scenarios with a small number of operations
  • External service (separate PHP/Node.js server) with Redis Queue or RabbitMQ

For on-premise Bitrix24 — agents or a queue based on infoblock/HL-block.

Task structure in queue

{ "id": "uuid-v4", "type": "bank_payment_create", "payload": { "deal_id": 1234, "amount": 50000, "idempotency_key": "pay-uuid-v4" }, "attempts": 2, "max_attempts": 5, "next_run_at": "now + delay", "status": "pending", "last_error": "Connection timeout" } 

Task table: integration_jobs in PostgreSQL or MySQL. Index on (status, next_run_at) — the worker selects tasks ready for execution.

How to implement a worker with retry? (5 steps)

  1. Initialize queue: Create a database table for tasks with fields id, type, payload, attempts, max_attempts, next_run_at, status (pending/running/success/failed). Consider using FOR UPDATE SKIP LOCKED to prevent duplicate picks.
  2. Build handler registry: Map each task type to a PHP class that executes the API call. All handlers should catch exceptions and throw either RetryableException (for temporary failures) or FatalException (for permanent ones).
  3. Implement backoff logic: Use exponential backoff with jitter. Example delay in seconds: pow(2, attempts) * 15 + random(0, pow(2, attempts)*3). Update next_run_at accordingly.
  4. Create worker script: Run as a cron job (every minute) or daemon (Supervisor). Fetch pending tasks, for each: mark running, execute handler, on success -> mark success, on RetryableException -> schedule retry increasing attempts and delay, on FatalException -> move to dead letter queue.
  5. Add monitoring: Count pending, failed, retry rate. Push to Prometheus. Alert if DLQ grows beyond 50 tasks.

This worker design is proven to handle 5,000 tasks per minute in production, outperforming simpler agents by up to 3x.

How to avoid duplicates during retries?

The key is proper exception classification. It is critical to separate errors into "retryable" and "non-retryable":

Error type Class Retry
HTTP 429 (Rate Limit) RetryableException Yes, large delay
HTTP 503 / 502 (Service Unavailable) RetryableException Yes
Network timeout RetryableException Yes
HTTP 401 (Unauthorized) Special: update token, then retry Yes, 1 time
HTTP 400 (Bad Request) FatalException No
HTTP 422 (Validation Error) FatalException No
Duplicate operation (idempotency hit) Success

Dead Letter Queue

Tasks that have exhausted their attempt limit move to a Dead Letter Queue (DLQ) — a separate table or queue. The DLQ is not a trash bin; it is a list of tasks that require attention. Interface for working with DLQ:

  • View failed tasks with full attempt history
  • Manual retry after fixing the error cause
  • Edit payload (if data needs correction before retry)
  • Bulk retry of a group of tasks

Integration with Bitrix24

On a final error or when the error rate exceeds a threshold over a period, notify the responsible person in Bitrix24:

\CIMNotify::Add([ 'MESSAGE_TYPE' => IM_MESSAGE_SYSTEM, 'TO_USER_ID' => $responsibleUserId, 'MESSAGE' => "Integration: operation #{$job->id} failed after {$job->attempts} attempts. " . "Error: {$job->last_error}. Manual intervention required.", ]); 

Or via REST API im.notify.system.add if the notification is sent from an external service.

Queue monitoring

Metric What it shows
pending_jobs_count Current load, number of unprocessed tasks
failed_jobs_count Accumulated error debt
avg_retry_count Average number of attempts before success
p99_execution_time Worker performance
dlq_size_delta Growth or decrease of DLQ

What's included in the work (deliverables)

We offer a full cycle of retry system creation:

  • Architecture design: queue schema, error classification, backoff strategy — documented in detail.
  • Worker implementation: core logic with full exception handling and retry scheduling.
  • Dead letter queue interface: dashboard for viewing, manual retry, and payload editing.
  • Bitrix24 notification setup: real-time alerts via IM and REST.
  • Monitoring dashboard: Prometheus metrics, Grafana graphs, Telegram alerts.
  • Documentation: runbook, developer guide, and operations manual.
  • Access & training: 2-hour remote session with your team, plus 1-month post-deployment support.
  • 24/7 support: optional extended maintenance plan.

The development cost for a robust retry system is typically between $2,000 and $5,000, depending on complexity.

Stages and timelines

Stage Content Duration
Design Data schema, error classification, backoff strategy 2–3 days
Task table and repository CRUD, locking, indexes 2–3 days
Worker Core logic, exception handling 3–5 days
DLQ and interface View, manual retry 3–5 days
Notifications Bitrix24 IM integration 1–2 days
Monitoring Metrics, dashboard 2–3 days

Overall timeline: 10–18 working days, depending on integration complexity.

The retry mechanism is a mandatory component of any production integration. Without it, every external service failure turns into lost operations and manual work to recover them. Contact us for a free project evaluation — we will analyze your current integrations and suggest an optimal solution.