System Design Lab›System Design Questions›Design Video Conferencing (Zoom)
Design Video Conferencing (Zoom)
HardReal-time Systemsvideoreal-timestreamingwebrtc
Question Overview
Design a video conferencing platform like Zoom where participants join meetings with live audio, video, and screen sharing. The interesting parts are choosing between mesh, SFU, and MCU topologies, traversing NATs, adapting quality to each participant's bandwidth, and placing media servers close to users.…
Sign up to see the full question and AI interviewer
Requirements
- Join meetings by link with live audio, video, and screen sharing for up to 500 participants
- One-way latency under 150 ms in-region; audio stays clear even when video degrades
- Per-participant quality adaptation for links from 300 kbps to 10 Mbps
- Work behind NATs and firewalls, falling back to TCP or TLS on port 443 when UDP is blocked
- Host controls such as mute and remove, in-meeting chat, and cloud recording
- Survive a media server failure with a brief reconnect rather than a dropped meeting
Back-of-the-envelope numbers
- Concurrency: 10M meetings/day × 45 min ≈ 7.5M meeting-hours ÷ 24 h ≈ 312K concurrent meetings on average; the 1M peak is ≈ 3×
- Uplink per participant with simulcast: 720p at 1.5 Mbps + 360p at 0.5 Mbps + 180p at 0.15 Mbps + ~32 kbps Opus audio ≈ 2.2 Mbps
- Mesh vs SFU: in a 10-person mesh each client uploads 9 × 1.5 Mbps = 13.5 Mbps; through an SFU it uploads ~2.2 Mbps regardless of meeting size
- SFU traffic: 5M × 2.2 Mbps ≈ 11 Tbps in; each viewer receives ~2 Mbps (720p speaker, three 180p thumbnails, audio) → 10 Tbps out
- SFU fleet: 10 Tbps ÷ ~5 Gbps per server (packet-rate bound at ~500K packets/s of ~1,200 bytes) ≈ 2,000 servers across ~20 regions before headroom
- Signaling: 50M joins/day ÷ 86,400 s ≈ 580 joins/s average, but top-of-the-hour spikes reach ~10× (≈ 6K joins/s)
- Recording: 5% × 7.5M = 375K hours/day × 675 MB/hour (1.5 Mbps × 3,600 s) ≈ 253 TB/day, ≈ 7.6 PB at 30-day retention
Key components
- Signaling service over WebSocket: roster and meeting state, host controls, SDP offer/answer and ICE candidate exchange, with a stateful per-meeting controller
- ICE with STUN and TURN: STUN discovers public addresses, TURN relays media when direct UDP fails, and SFUs on public IPs also accept TCP or TLS on 443
- SFU media servers: forward RTP packets without decoding and choose which layer each subscriber receives, far cheaper and lower latency than an MCU
- Simulcast or SVC (VP9/AV1): each sender publishes several resolutions or layers so the SFU can downshift a weak receiver without affecting others
- Congestion control and loss recovery: transport-wide bandwidth estimation, NACK retransmits within the jitter-buffer window, keyframe requests, and Opus in-band FEC
- Placement and cascading: route users to the nearest region via geo-DNS or anycast, and link SFUs across regions so each stream crosses an ocean once
- Recording: a hidden recorder participant subscribes to streams, writes them to object storage, and composites and transcodes after the meeting
Common mistakes
- Sending media over TCP or WebSockets, where head-of-line blocking turns one lost packet into a latency spike; real-time media wants UDP
- Using a full mesh for large meetings, where every client's upload grows linearly with the number of participants
- Choosing an MCU for every meeting, paying to decode and re-encode every stream and adding latency, when an SFU only forwards packets
- Forwarding every participant's video to everyone in a 500-person meeting instead of only visible tiles and the few loudest speakers' audio
- Hosting each meeting in one fixed region, so participants on other continents pay 150 ms or more of extra round-trip time
- Retransmitting every lost packet even after its playout deadline, when late media is useless and FEC or concealment works better
Likely follow-ups
- How would you host a 20,000-attendee webinar where only a few people speak?
- How would you add end-to-end encryption when the SFU still needs to route packets and select layers?
- A participant's bandwidth drops to 300 kbps mid-meeting; what exactly happens to what they send and receive?
- How do you cascade SFUs for a meeting with participants in Mumbai, London, and San Francisco?
- What happens to a live meeting when its SFU crashes?
- How would you add live captions or server-side noise suppression without breaking the latency target?
No community solutions yet
Be the first to publish your solution
Practice ‘Design Video Conferencing (Zoom)’ with an AI Interviewer
Get scored feedback on your diagram, scalability approach, and trade-offs. Free while we grow — up to 3 full interviews a day.