Implementing Saga Pattern for Distributed Transactions
Imagine an e-commerce site with hundreds of thousands of orders per day. After a successful payment, the delivery service goes down, and data gets out of sync: money deducted but order not shipped. Without proper management of distributed transactions, such scenarios lead to financial losses and customer churn. The Saga Pattern solves this by breaking a business transaction into local steps with compensations on failure. We have been implementing sagas in microservice systems for over 5 years — contact us for a turnkey solution with data integrity guarantee.
Two Types of Saga
Choreography — services react to each other's events without a central coordinator. Each service publishes events to a broker (e.g., Kafka) and subscribes to relevant ones. If payment fails, the Inventory Service rolls back the reservation via an inverse event.
Orchestration — a central Saga Orchestrator (e.g., Temporal) explicitly manages steps and compensations. Code is clearer, easier to debug, but requires a separate service.
| Criteria | Orchestration | Choreography |
|---|---|---|
| Coordination | Central coordinator | Event bus (Kafka, RabbitMQ) |
| Development complexity | Medium (needs coordinator service) | Low at start, high with many services |
| Debugging | Easy (coordinator logs) | Hard (event tracing) |
| Reliability | Depends on coordinator | Decentralized |
| Performance | One step at a time | Parallel steps (less control) |
How to Choose Between Orchestration and Choreography?
If you have up to 5 services and simplicity of debugging is crucial — choose orchestration. For large systems with 10+ services and high load, choreography via Kafka provides better scalability. We often combine both: orchestration for critical chains, choreography for background processes. Our engineers will select the right approach for your project — get in touch for a consultation.
What Is Temporal and Why Do You Need It?
Temporal is a production-ready engine for long-running workflows. It automatically retries activities, stores execution history, and lets you inspect sagas via UI. It guarantees compensation execution even if a service crashes. Over 95% of sagas with Temporal complete without manual intervention. According to our data, orchestration via Temporal is 2-3 times more reliable than choreography without a coordinator. Example orchestration with Temporal:
import { proxyActivities, sleep } from '@temporalio/workflow'; const { reserveStock, chargePayment, createShipment, releaseStock, refund } = proxyActivities({ startToCloseTimeout: '10 seconds' }); export async function createOrderWorkflow(input: CreateOrderInput): Promise<void> { let stockReserved = false; let paymentCharged = false; try { await reserveStock({ orderId: input.orderId, items: input.items }); stockReserved = true; await chargePayment({ orderId: input.orderId, amount: input.amount }); paymentCharged = true; await createShipment({ orderId: input.orderId, address: input.address }); } catch (error) { // Temporal ensures compensations execute if (paymentCharged) { await refund({ orderId: input.orderId }); } if (stockReserved) { await releaseStock({ orderId: input.orderId }); } throw error; } } Persistent Saga with State
A saga must survive service restarts. State is stored in a database (PostgreSQL, MySQL). We use a table with statuses (running, completed, failed, compensating) and context. When a service crashes, it re-reads incomplete sagas and continues from the last step.
interface SagaState { sagaId: string; sagaType: string; status: 'running' | 'completed' | 'failed' | 'compensating'; currentStep: number; context: Record<string, unknown>; completedSteps: string[]; failedStep?: string; createdAt: Date; updatedAt: Date; } class PersistentSagaOrchestrator { async startSaga(sagaType: string, context: unknown): Promise<string> { const sagaId = uuidv4(); await this.sagaRepo.save({ sagaId, sagaType, status: 'running', currentStep: 0, context, completedSteps: [] }); await this.executeSaga(sagaId); return sagaId; } } Choreography via Kafka
Example event handling in the Inventory service:
// Order Service publishes an event await kafka.producer.send({ topic: 'order.events', messages: [{ key: orderId, value: JSON.stringify({ type: 'OrderCreated', orderId, items, customerId })}] }); // Inventory Service listens and reserves kafka.consumer.subscribe({ topic: 'order.events' }); kafka.consumer.run({ eachMessage: async ({ message }) => { const event = JSON.parse(message.value.toString()); if (event.type !== 'OrderCreated') return; try { await inventoryService.reserveStock(event.orderId, event.items); await kafka.producer.send({ topic: 'inventory.events', messages: [{ key: event.orderId, value: JSON.stringify({ type: 'StockReserved', orderId: event.orderId })}] }); } catch { await kafka.producer.send({ topic: 'inventory.events', messages: [{ key: event.orderId, value: JSON.stringify({ type: 'StockReservationFailed', orderId: event.orderId })}] }); } } }); Common Problems and Solutions
| Problem | Solution |
|---|---|
| Non-idempotent operations | Check existing state before creating (example above) |
| Loss of saga state on crash | Persist status and context in DB |
| Infinite retries and system overload | Exponential backoff and retry limit (usually 3-5) |
| Lack of monitoring | Tools like Jaeger, Grafana to visualize saga progress |
What’s Included in the Work
The implementation process covers:
- Analysis: define business transactions, service boundaries, failure points.
- Design: choose between orchestration and choreography, prepare compensation plan.
- Implementation: write saga code, integrate with Temporal or Kafka, ensure idempotency.
- Testing: unit tests, integration tests, failure scenario tests (chaos engineering).
- Deployment: deploy in Docker/Kubernetes, set up monitoring (Jaeger, Grafana).
Deliverables: saga documentation, repository access, team training, and 2 weeks of post-launch support.
Example of a complex saga with multiple compensations
In real projects, one saga can involve dozens of services. For example, an order with pre-order and delivery: warehouse reservation, payment, supplier order creation, shipping setup. Compensations run in reverse order, and Temporal guarantees execution even after several restarts.
Timelines and Guarantees
- Simple orchestration (2-3 services, without Temporal) — 1 to 2 weeks.
- Orchestration with Temporal + monitoring — 2 to 3 weeks.
- Choreography via Kafka with idempotent handlers — 2 to 4 weeks.
Cost is determined individually after analysis. We offer a 6-month code guarantee. Over 100 delivered projects in microservices (5+ years in the market). Contact us for a free project assessment and an optimal solution.
Further reading: Saga pattern (Wikipedia).







