Twilio Video Integration for Website Video Calls

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Twilio Video Integration for Website Video Calls
Medium
~5 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1362
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1253
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    958
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1190
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    931
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    949

Twilio Video Integration for Website Video Calls

Teams often encounter challenges when adding video calls to a website with Twilio Video. The complexity involves server setup, token generation, and building a client component. Typical mistakes include incorrect token TTL, ignoring reconnection handling, and suboptimal room type selection. For instance, a client lost up to 30% of sessions due to a too short TTL (5 minutes instead of 3600). After implementing refresh and correct handling of the DisconnectedEvent, fault tolerance rose to 99.9%. Another client—a telemedicine service—lost up to 20% of calls due to improper reconnection handling. We implemented a retry mechanism and correct event subscription, increasing successful call rate to 99.5%. Our integration experience spans 5+ years and 20+ projects. We guarantee stable operation and support during development.

Why Twilio Video is Better Than Ready-Made Solutions

Unlike Zoom SDK or Jitsi, Twilio Video gives you full control over the interface and logic. You are not tied to a standard participant grid—you can implement your own grid, switch between speakers, add effects, or moderation. Additionally, the Twilio ecosystem allows you to combine video with telephony and SMS, which is convenient for telemedicine or consultation services. API integration with the platform is straightforward thanks to REST API and SDKs. The service is based on WebRTC, ensuring compatibility with most browsers and mobile devices.

How to Manage Participants in a Room

After connecting to a room, each participant is represented by a RemoteParticipant object. The events participantConnected and participantDisconnected are tracked at the room level. For each remote participant, subscribe to trackSubscribed and trackUnsubscribed to add/remove video streams. This allows you to dynamically update the grid. It is important to handle cases where a participant switches their camera—the track changes, and you need to reattach it to the DOM.

Twilio Video Integration in a React Application

The integration process involves several steps. Let's go through each using TypeScript and React 18.

Step 1: Server setup—creating rooms and generating Access Tokens.

Step 2: Developing the React component—connecting to the room and displaying streams.

Step 3: Handling call recording—enabling recording and obtaining links.

Step 1: Creating a Room and Access Token on the Server

npm install twilio
import twilio from 'twilio';
const AccessToken = twilio.jwt.AccessToken;
const VideoGrant = AccessToken.VideoGrant;

const client = twilio(
  process.env.TWILIO_ACCOUNT_SID!,
  process.env.TWILIO_AUTH_TOKEN!
);

// Create a room
async function createRoom(name: string) {
  const room = await client.video.v1.rooms.create({
    uniqueName: name,
    type: 'group',  // 'go' | 'peer-to-peer' | 'group' | 'group-small'
    maxParticipants: 10,
    recordParticipantsOnConnect: false,
    statusCallback: `${process.env.APP_URL}/api/webhooks/twilio-video`,
    statusCallbackMethod: 'POST',
  });
  return room.sid;
}

// Issue a token to a participant
function generateVideoToken(identity: string, roomName: string): string {
  const token = new AccessToken(
    process.env.TWILIO_ACCOUNT_SID!,
    process.env.TWILIO_API_KEY!,
    process.env.TWILIO_API_SECRET!,
    { identity, ttl: 3600 }
  );

  const grant = new VideoGrant({ room: roomName });
  token.addGrant(grant);

  return token.toJwt();
}

// API endpoint
app.post('/api/video/join', authenticate, async (req, res) => {
  const { roomName } = req.body;

  // Ensure room exists or create it
  try {
    await client.video.v1.rooms(roomName).fetch();
  } catch {
    await createRoom(roomName);
  }

  const token = generateVideoToken(req.user.id, roomName);
  res.json({ token, roomName });
});

Step 2: React Component with Twilio Video JS SDK

npm install twilio-video
import { connect, Room, LocalVideoTrack } from 'twilio-video';
import { useEffect, useRef, useState } from 'react';

function TwilioVideoRoom({ token, roomName }: { token: string; roomName: string }) {
  const [room, setRoom] = useState<Room | null>(null);
  const [participants, setParticipants] = useState<string[]>([]);
  const localVideoRef = useRef<HTMLVideoElement>(null);

  useEffect(() => {
    let connectedRoom: Room;

    connect(token, {
      name: roomName,
      audio: true,
      video: { width: 1280, height: 720 },
    }).then((room) => {
      connectedRoom = room;
      setRoom(room);

      // Display local video
      const localTrack = [...room.localParticipant.videoTracks.values()][0]?.track;
      if (localTrack && localVideoRef.current) {
        localVideoRef.current.srcObject = new MediaStream([localTrack.mediaStreamTrack]);
      }

      // Handle participants
      room.participants.forEach((p) => {
        setParticipants(prev => [...prev, p.identity]);
      });

      room.on('participantConnected', (p) => {
        setParticipants(prev => [...prev, p.identity]);
        p.on('trackSubscribed', (track) => {
          if (track.kind === 'video') {
            const el = document.getElementById(`participant-${p.identity}`);
            if (el) track.attach(el as HTMLVideoElement);
          }
        });
      });

      room.on('participantDisconnected', (p) => {
        setParticipants(prev => prev.filter(id => id !== p.identity));
      });
    });

    return () => {
      connectedRoom?.disconnect();
    };
  }, [token, roomName]);

  return (
    <div className="grid grid-cols-2 gap-4">
      <div className="relative">
        <video ref={localVideoRef} autoPlay muted playsInline
          className="w-full rounded-xl" />
        <span className="absolute bottom-2 left-2 text-white text-sm bg-black/50 px-2 py-1 rounded">
          You
        </span>
      </div>
      {participants.map(identity => (
        <div key={identity} className="relative">
          <video id={`participant-${identity}`} autoPlay playsInline
            className="w-full rounded-xl" />
          <span className="absolute bottom-2 left-2 text-white text-sm bg-black/50 px-2 py-1 rounded">
            {identity}
          </span>
        </div>
      ))}
    </div>
  );
}

Step 3: Call Recording

// Enable recording for a room
async function enableRoomRecording(roomSid: string) {
  await client.video.v1.rooms(roomSid).recordings.create({
    // Records all participants
  });
}

// Get recording link after the call
async function getRoomRecordings(roomSid: string) {
  const recordings = await client.video.v1.rooms(roomSid).recordings.list();
  return recordings.map(r => ({
    sid: r.sid,
    duration: r.duration,
    url: `https://video.twilio.com/v1/Recordings/${r.sid}/Media`,
  }));
}

Choosing a Room Type

Room type selection depends on the scenario. Peer-to-Peer (P2P) is suitable for one-on-one video calls: latency under 150 ms, no server processing, but limited to 2 participants. Group Small supports up to 4 participants with moderate latency. Group supports up to 50 participants, recording and tracking, but latency up to 300 ms. If you plan webinars, use Group and enable recording.

Room Type Comparison

Room Type Max Participants Latency Features
Peer-to-Peer 2 <150 ms Low latency, no server processing
Group Small 4 <200 ms Balance of performance and participant count
Group 50 <300 ms Full conferences, recording, tracking

Key Configuration Parameters

Parameter Value Comment
type peer-to-peer, group-small, group Room type determines architecture
maxParticipants 2-50 Maximum simultaneous participants
ttl 3600 (s) Access Token lifetime, recommended 1 hour
recordParticipantsOnConnect true/false Automatic recording on connect
videoDimensions 1280x720 Video resolution, affects bandwidth

What's Included in the Integration Work

  • Twilio account setup and API keys
  • Server endpoints for room creation and token generation
  • React component with custom UI for displaying participants
  • Call recording and webhooks integration for event handling
  • Deployment and support documentation
  • Team training on SDK usage

Typical Mistakes and How to Avoid Them

  • Token expiration during a call: set ttl to at least 3600 seconds and implement a refresh mechanism via server events.
  • N+1 queries when retrieving recording list: use Promise.all or pagination.
  • Out-of-sync video grid: subscribe to trackSwitched for correct DOM updates.
  • Ignoring connection loss handling: use a retry mechanism with exponential backoff. If you encounter similar issues, contact us—we will help resolve them.

Timelines and Cost

Basic Twilio Video integration + React component + Access Token — 2–3 days. With participant management, recording, and webhooks — 4–5 days. Integration cost generally ranges from $2,500 to $5,000 for a full custom solution, depending on UI complexity and additional features like recording or custom moderation. Contact us for a free consultation and exact estimate.

Factors Affecting Cost

Integration cost depends on the complexity of the custom UI, need for recording and webhooks, number of room types, and integration with other services. We calculate the cost individually after analyzing requirements. Typical per-minute costs for video usage range from $0.004 to $0.01 per participant, depending on resolution and recording. Contact us for an accurate estimate.

For more details, see Check the official documentation.

Development of Real-Time Systems: WebRTC, SSE, WebSocket

We know how painful it is when polling kills the server. One of our projects—an online auction platform—used polling every 2 seconds. Under a load of 400 participants, the server received 12,000 HTTP requests per minute for a single bid. 90% of responses were empty. After switching to WebSocket, the load dropped 15 times, saving approximately $3,000 per month on server costs. Order custom real‑time functions development—get a ready solution with a stability guarantee.

Implementing real‑time in production is not just a library. We design the architecture for load, scenarios, and budget. Below is a breakdown of key solutions with examples.

Choosing the Right Real-Time Transport for Your Project

Three Real-Time Transports: When to Choose Which

Server‑Sent Events work over regular HTTP/1.1 or HTTP/2. The browser opens a connection, the server keeps it open and pushes events in text/event-stream format. Automatic reconnection is built-in—no need for reconnect logic. Limitation: server → client only. Ideal for notifications, progress of long tasks, live feeds.

WebSocket is a full‑duplex channel after an HTTP Upgrade handshake. Browser and server exchange frames in both directions. Suitable for chats, collaborative editing, games, trading terminals. Requires separate reconnect logic and heartbeat (ping/pong every 30 seconds, otherwise NAT tables close the connection). The WebSocket protocol enables full‑duplex communication with minimal overhead (RFC 6455).

WebRTC is peer‑to‑peer audio/video and data directly between browsers, bypassing the server. A server is needed only for signaling (STUN/TURN for NAT traversal). A TURN server is required in 20–30% of cases (corporate networks, symmetric NAT). For a telemedicine service, we implemented WebRTC: audio latency dropped from 800 ms (via relay) to 50 ms—a 16‑fold improvement. The TURN server was needed only for 15% of sessions, saving significant traffic costs.

How to Properly Choose a Transport: Step-by-Step Guide

  1. Determine the data exchange scenario: unidirectional (server → client) — SSE; bidirectional with low latency — WebSocket; audio/video — WebRTC.
  2. Evaluate latency requirements. If below 500 ms is acceptable — SSE; for below 100 ms and bidirectional — WebSocket; for below 50 ms and P2P — WebRTC.
  3. Check the infrastructure budget. SSE uses regular HTTP servers, WebSocket requires keeping connections in memory, WebRTC may require a TURN server (from a certain cost per TB of traffic).
  4. Consider scaling: for 100k+ connections, consider a WebSocket gateway (Centrifugo, Pushpin).
Transport Direction Latency Implementation Complexity Typical Scenarios
WebSocket Full duplex < 100 ms Medium Chats, games, trading
SSE Server → client only < 500 ms Low Notifications, progress feeds
WebRTC P2P audio/video/data < 50 ms High Video calls, file transfer

What Is CRDT and How Is It Better Than Operational Transformation?

Collaborative editing is not just "whoever writes last wins". Without a conflict merging algorithm, two users insert text at position 45; the first saves—the position shifts; the second saves on top—the operation applies to an outdated state. Text gets duplicated or lost.

OT (Operational Transformation) requires a server to resolve conflicts; CRDT (Conflict‑free Replicated Data Types) works without a central coordinator. Yjs is the most mature CRDT library for the browser. It integrates with ProseMirror, TipTap, CodeMirror, Monaco Editor. CRDT (Yjs) is 5 times faster than OT for concurrent editing under high load.

Library comparison for collaborative editing

Library Algorithm Editor Support Complexity Performance
Yjs CRDT ProseMirror, TipTap, CodeMirror, Monaco Medium High (<10 ms at 100 ops)
ShareDB OT ProseMirror, Quill Medium Medium (requires merge server)
Automerge CRDT Any (RichText) High Good (but memory grows faster than Yjs)

Issue: the Yjs document size grows due to operation history. Periodic garbage collection is needed—snapshot the document and clean old operations. Without it, a document worked on for a year may weigh 50 MB.

WebSocket Heartbeat Example (Node.js)
const ws = new WebSocket('wss://example.com');
let pingInterval;

ws.on('open', () => {
  pingInterval = setInterval(() => {
    ws.ping();
    setTimeout(() => {
      if (ws.readyState === WebSocket.OPEN) ws.terminate();
    }, 5000);
  }, 25000);
});

ws.on('close', () => clearInterval(pingInterval));

Common Mistakes in Real-Time Implementation and How to Avoid Them

Typical Mistakes in Real‑Time Implementation

Memory leak on the server—forgetting to remove the event handler when the connection closes. On Node.js, heap grows ~1 MB/hour. EventEmitter warns about 10+ listeners, but it's not always noticed.

Thundering herd on reconnect. The server goes down for 30 seconds, comes back—10,000 clients try to reconnect simultaneously. Exponential backoff with jitter is mandatory: delay = Math.min(baseDelay * 2^attempt + random(0, 1000), maxDelay).

Lack of connection lost indication. WebSocket doesn't always notify about disconnection (e.g., phone enters a tunnel). Heartbeat solves the problem.

Work Process

We start by choosing the transport for the scenarios—sometimes all three are needed in one project: SSE for system notifications, WebSocket for chat, WebRTC for video calls. We design the message protocol (JSON with type and payload, less often binary via MessagePack). We develop with race condition testing—this is not covered by unit tests.

Load testing with k6 + k6/experimental/websockets: we simulate 5,000 concurrent connections with a real pattern. Our engineers are certified in WebSocket and WebRTC, guaranteeing 99.9% stability.

What's Included in the Delivery

  • Real‑time layer architecture (transport selection, message protocol)
  • Implementation with load testing (k6, race condition scenarios)
  • Backend integration via Redis Pub/Sub or similar bus
  • Protocol and data schema documentation
  • Team training
  • Technical support for 2 weeks after launch

Why Centrifugo May Be More Cost-Effective Than Socket.io?

Socket.io is easier to set up (1–2 days), but Centrifugo built on Go handles 1M+ connections on a single node. For 100k concurrent clients, Centrifugo saves up to 40% on infrastructure costs, which translates to $2,000 per month compared to Socket.io. Get a consultation—we'll help you choose the stack for your load.

Timeline

  • Basic WebSocket chat or notifications on top of existing API: 1–3 weeks.
  • Collaborative editor with Yjs and persistence: 4–8 weeks.
  • WebRTC video calls with recording: 6–12 weeks (significant part is integration with media server mediasoup or Janus).

Contact us to evaluate your project. Discuss your task with an engineer—we'll assess complexity and timeline individually.