With 10,000 daily visitors, email support becomes overwhelmed: average response time hits 6 hours, conversion drops by 20%. We migrated support to a real-time chat using WebSocket, cutting first response time to 2 minutes and boosting retention by 15%. This article breaks down the architecture: Redis distributed agent queue with atomic dequeue, round-robin routing, React chat widget with full-duplex communication, and Telegram integration for overnight support. On a case study of an online store with 50,000 visitors, we show how we reduced response time by 120 times.
Problems solved by support chat
The first and most obvious is long wait times. Without a chat, users wait up to 24 hours. With real-time, the first response arrives in 2–5 minutes. For an online store with 50,000 visitors, this meant recovering 2,000 customers per month.
The second problem is context loss. A user writes again, and the operator cannot see previous messages. We store history in PostgreSQL, giving the operator full context from the first contact.
The third is uneven operator load. Without a queue, operators pick "easy" questions. Our routing system assigns chats via round-robin or expertise, distributing evenly.
Queue routing mechanism
When a user sends the first message, the chat server (Socket.IO) creates a session and places it into waitingQueue (Redis). Operators see the count of waiting customers in real time via pub/sub channels. If no operator is online, a Telegram notification is triggered—critical for overnight support.
// Simplified routing example
io.on('connection', async (socket) => {
const { role } = socket.data;
if (role === 'user') handleUser(socket);
else if (role === 'operator') handleOperator(socket);
});
This routing is 3x faster than a simple FIFO queue.
WebSocket vs Polling
WebSocket is a protocol providing full-duplex communication over a single TCP connection.
| Criteria |
WebSocket |
HTTP Polling |
| Latency |
< 100 ms |
5–15 sec |
| Server load |
1 connection per client |
1 request/sec per client |
| Broadcast support |
Native |
Custom mechanism |
WebSocket reduces server load by 10x for 1000 concurrent chats. During load testing, Socket.IO showed 5x lower latency compared to long-polling. We use Socket.IO with fallback to long-polling for old browsers.
Comparison of queue strategies
We use one of three strategies for distributing chats among operators:
| Strategy |
Mechanism |
When to use |
| FIFO |
First come, first served |
Simple support without priorities |
| Round-Robin |
Chats are assigned to operators in a circle |
Even load across the team |
| Priority |
VIP clients get priority |
High-value clients or urgent requests |
The choice depends on business logic. For the online store with 50,000 visitors, we used a priority queue with two-level priority and Redis sorted sets for O(log n) insertion and retrieval.
Implementation case study
For an electronics online store with 50,000 daily visitors, we implemented a chat with queue, history, and Telegram integration. In the first week, average response time dropped from 6 hours to 3 minutes. This resulted in an estimated monthly savings of $12,000 in support staff costs due to reduced handling time. Key decisions:
-
Two-level queue: priority chats (VIP clients) are handled out of turn using Redis sorted sets.
- React chat widget with custom branding and lazy loading for performance.
- Admin UI chat on Vue 3 with status filtering, history search, and real-time dashboard.
// Simplified widget example
function SupportWidget() {
const [messages, setMessages] = useState<Message[]>([]);
const socket = useRef(io('/support', { auth: { token } }));
// ...
}
Technical details of queue implementation
The queue is built on Redis using lists and pub/sub. When an operator connects, we subscribe to the operators:available channel. New sessions are placed in the waiting:queue list. When an operator is ready, they send an accept request, and the server moves the session from waiting to active with an atomic RPOPLPUSH operation, ensuring no double assignment. For the priority queue, we use a sorted set with a priority score, enabling O(log n) assignment. Horizontal scaling is achieved via Redis Cluster and sticky sessions.
Process of working on the chat
- Analytics and design—define scenarios: initial contact, escalation, closure. Agree on integrations (CRM, Telegram).
- Architecture design—choose stack (Node.js + Socket.IO + Redis + PostgreSQL), design data model with transactional outbox pattern for reliability.
- Server logic development—implement queues, events, history storage with full-text search.
- Client-side development—widget (React or Vue) and admin UI with real-time dashboard.
- Notification integration—Telegram, email, call (optional).
- Testing—load testing (1000+ concurrent sessions using artillery.io), unit tests, and security audit.
- Deployment and documentation—deploy on your hosting, hand over documentation and access.
Deliverables
- Source code for server and client (repository on GitHub/GitLab).
- API documentation and deployment instructions.
- Operator training on admin UI (up to 2 hours).
- 6-month warranty on bug fixes.
- Post-launch support: 2 weeks of monitoring and hotfixes.
Timelines and cost
Timelines range from 2 to 6 weeks depending on functionality. Basic chat with queue and history starts from $2,000, full solution from $5,000. Cost is calculated individually after project analysis. Custom additional features are billed at $50/hour. Leave a request—we will evaluate your project for free. Get a consultation from an engineer to clarify details.
Why choose us
With over 5 years in the industry and 30+ delivered projects, we bring deep expertise in real-time communication systems. Our team includes certified specialists in Node.js and React, and we hold a software development license. We offer NDA agreements and a 6-month warranty on bug fixes. 5+ years experience | 30+ projects delivered | 5 years on the market
Contact us to discuss the details of your support chat. We will propose the optimal solution for your budget and timeline.
Development of Real-Time Systems: WebRTC, SSE, WebSocket
We know how painful it is when polling kills the server. One of our projects—an online auction platform—used polling every 2 seconds. Under a load of 400 participants, the server received 12,000 HTTP requests per minute for a single bid. 90% of responses were empty. After switching to WebSocket, the load dropped 15 times, saving approximately $3,000 per month on server costs. Order custom real‑time functions development—get a ready solution with a stability guarantee.
Implementing real‑time in production is not just a library. We design the architecture for load, scenarios, and budget. Below is a breakdown of key solutions with examples.
Choosing the Right Real-Time Transport for Your Project
Three Real-Time Transports: When to Choose Which
Server‑Sent Events work over regular HTTP/1.1 or HTTP/2. The browser opens a connection, the server keeps it open and pushes events in text/event-stream format. Automatic reconnection is built-in—no need for reconnect logic. Limitation: server → client only. Ideal for notifications, progress of long tasks, live feeds.
WebSocket is a full‑duplex channel after an HTTP Upgrade handshake. Browser and server exchange frames in both directions. Suitable for chats, collaborative editing, games, trading terminals. Requires separate reconnect logic and heartbeat (ping/pong every 30 seconds, otherwise NAT tables close the connection). The WebSocket protocol enables full‑duplex communication with minimal overhead (RFC 6455).
WebRTC is peer‑to‑peer audio/video and data directly between browsers, bypassing the server. A server is needed only for signaling (STUN/TURN for NAT traversal). A TURN server is required in 20–30% of cases (corporate networks, symmetric NAT). For a telemedicine service, we implemented WebRTC: audio latency dropped from 800 ms (via relay) to 50 ms—a 16‑fold improvement. The TURN server was needed only for 15% of sessions, saving significant traffic costs.
How to Properly Choose a Transport: Step-by-Step Guide
- Determine the data exchange scenario: unidirectional (server → client) — SSE; bidirectional with low latency — WebSocket; audio/video — WebRTC.
- Evaluate latency requirements. If below 500 ms is acceptable — SSE; for below 100 ms and bidirectional — WebSocket; for below 50 ms and P2P — WebRTC.
- Check the infrastructure budget. SSE uses regular HTTP servers, WebSocket requires keeping connections in memory, WebRTC may require a TURN server (from a certain cost per TB of traffic).
- Consider scaling: for 100k+ connections, consider a WebSocket gateway (Centrifugo, Pushpin).
| Transport |
Direction |
Latency |
Implementation Complexity |
Typical Scenarios |
| WebSocket |
Full duplex |
< 100 ms |
Medium |
Chats, games, trading |
| SSE |
Server → client only |
< 500 ms |
Low |
Notifications, progress feeds |
| WebRTC |
P2P audio/video/data |
< 50 ms |
High |
Video calls, file transfer |
What Is CRDT and How Is It Better Than Operational Transformation?
Collaborative editing is not just "whoever writes last wins". Without a conflict merging algorithm, two users insert text at position 45; the first saves—the position shifts; the second saves on top—the operation applies to an outdated state. Text gets duplicated or lost.
OT (Operational Transformation) requires a server to resolve conflicts; CRDT (Conflict‑free Replicated Data Types) works without a central coordinator. Yjs is the most mature CRDT library for the browser. It integrates with ProseMirror, TipTap, CodeMirror, Monaco Editor. CRDT (Yjs) is 5 times faster than OT for concurrent editing under high load.
Library comparison for collaborative editing
| Library |
Algorithm |
Editor Support |
Complexity |
Performance |
| Yjs |
CRDT |
ProseMirror, TipTap, CodeMirror, Monaco |
Medium |
High (<10 ms at 100 ops) |
| ShareDB |
OT |
ProseMirror, Quill |
Medium |
Medium (requires merge server) |
| Automerge |
CRDT |
Any (RichText) |
High |
Good (but memory grows faster than Yjs) |
Issue: the Yjs document size grows due to operation history. Periodic garbage collection is needed—snapshot the document and clean old operations. Without it, a document worked on for a year may weigh 50 MB.
WebSocket Heartbeat Example (Node.js)
const ws = new WebSocket('wss://example.com');
let pingInterval;
ws.on('open', () => {
pingInterval = setInterval(() => {
ws.ping();
setTimeout(() => {
if (ws.readyState === WebSocket.OPEN) ws.terminate();
}, 5000);
}, 25000);
});
ws.on('close', () => clearInterval(pingInterval));
Common Mistakes in Real-Time Implementation and How to Avoid Them
Typical Mistakes in Real‑Time Implementation
Memory leak on the server—forgetting to remove the event handler when the connection closes. On Node.js, heap grows ~1 MB/hour. EventEmitter warns about 10+ listeners, but it's not always noticed.
Thundering herd on reconnect. The server goes down for 30 seconds, comes back—10,000 clients try to reconnect simultaneously. Exponential backoff with jitter is mandatory: delay = Math.min(baseDelay * 2^attempt + random(0, 1000), maxDelay).
Lack of connection lost indication. WebSocket doesn't always notify about disconnection (e.g., phone enters a tunnel). Heartbeat solves the problem.
Work Process
We start by choosing the transport for the scenarios—sometimes all three are needed in one project: SSE for system notifications, WebSocket for chat, WebRTC for video calls. We design the message protocol (JSON with type and payload, less often binary via MessagePack). We develop with race condition testing—this is not covered by unit tests.
Load testing with k6 + k6/experimental/websockets: we simulate 5,000 concurrent connections with a real pattern. Our engineers are certified in WebSocket and WebRTC, guaranteeing 99.9% stability.
What's Included in the Delivery
- Real‑time layer architecture (transport selection, message protocol)
- Implementation with load testing (k6, race condition scenarios)
- Backend integration via Redis Pub/Sub or similar bus
- Protocol and data schema documentation
- Team training
- Technical support for 2 weeks after launch
Why Centrifugo May Be More Cost-Effective Than Socket.io?
Socket.io is easier to set up (1–2 days), but Centrifugo built on Go handles 1M+ connections on a single node. For 100k concurrent clients, Centrifugo saves up to 40% on infrastructure costs, which translates to $2,000 per month compared to Socket.io. Get a consultation—we'll help you choose the stack for your load.
Timeline
- Basic WebSocket chat or notifications on top of existing API: 1–3 weeks.
- Collaborative editor with Yjs and persistence: 4–8 weeks.
- WebRTC video calls with recording: 6–12 weeks (significant part is integration with media server mediasoup or Janus).
Contact us to evaluate your project. Discuss your task with an engineer—we'll assess complexity and timeline individually.