How Does Zoom Work? Architecture Deep Dive

Millions of people click a Zoom link every day without thinking about what makes it work. Let's break it down.
How Does Zoom Work Architecture Deep Dive cover image

When Zoom first appeared, it seemed almost magical. Before Zoom, video conferencing was a hassle. Skype was all frozen frames, and Google Hangouts degraded fast past three or four participants.

Zoom just worked. Without even needing to log in, you could click a link and be in a meeting a few seconds later. You didn't need an account or a plugin, and you didn't even need the other person to have added you as a contact first. And most importantly, the call itself held up, even with twenty video tiles on screen and someone joining from hotel wifi.

It was magical, but none of it was magic.

Zoom made a series of specific architectural decisions about how meetings exist, how video works, and where that work is done.

Here, we want to walk through those decisions: how a meeting can sit in Zoom's system for months without any media infrastructure running, how your video reaches everyone else, and how the same design works from a two-person call to a webinar with thousands of viewers.

How a Meeting Exists Before Anyone Joins

When you schedule a Zoom meeting for a week next Thursday, what actually exists between now and then? There are two ways of thinking about this:

  • A persistent-room model. Here, the room exists from the moment the meeting is set up. Think of a Discord voice channel, which is always there whether anyone's in it or not, holding its members and settings while people drift in and out.
  • A pure meeting model. Then, the answer is nothing. The call exists while people are connected and disappears when the last person hangs up.
With Zoom, you get both. Take a weekly team standup. You schedule it once, and the same meeting ID sits in everyone's calendar for months. That ID is a real, long-lived object. It stores the schedule, the passcode, the waiting room setting, and the host.

But between Tuesdays, nothing is running anywhere. There's no idle room on a server waiting for you. When the first person clicks the link on Tuesday morning, Zoom builds the live session on the spot, and when the last person leaves, it's torn down again. Next Tuesday's standup is a brand-new session that happens to map to the same ID.

Zoom breaks it down like this:

Object Lifetime What it holds
Scheduled or recurring meeting Until it expires or is deleted The meeting ID, schedule, settings, and policy
Personal Meeting ID Long-lived, but expires after 365 days of non-use A fixed ID and link tied to your account
Live meeting instance From the first join to the last leave Participants, roles, and media routing state

You can see the split directly in Zoom's API. A recurring meeting keeps a single reusable meeting ID, but each occurrence that runs gets its own meeting UUID. A call to that endpoint for a recurring meeting returns something like this:

GET https://api.zoom.us/v2/past_meetings/81402630914/instances

{
  "meetings": [
    { "uuid": "hDzXcQrDTiqDGDIeastGXg==", "start_time": "2026-08-04T14:00:11Z" },
    { "uuid": "K2f9tPZ0QQiJgpXlHzy4tA==", "start_time": "2026-08-11T14:00:36Z" },
    { "uuid": "0hUKUcnkQFCTL4vgg0HSyg==", "start_time": "2026-08-18T13:59:52Z" }
  ]
}

One meeting ID in the URL, and one UUID for each session that actually happened. The distinction carries through the rest of the API, so pulling the recording or the participant report for a specific Tuesday means querying by UUID, because those belong to an occurrence rather than the schedule.

Nothing on the media side exists until someone joins. The lifecycle of a live instance looks like this:

  1. The client talks to Zoom's control services over HTTPS. It authenticates, resolves the meeting ID, and gets checked against the meeting's settings and admission rules.
  2. The control layer picks media resources for this occurrence, routing you to a nearby data center and a lightly loaded media server. Zoom calls these servers multimedia routers, or MMRs.
  3. The client establishes its real-time media connection to that MMR while keeping its control connection alive. The live instance now exists, with a roster, roles, and forwarding state.
  4. When the last participant leaves, all of that state is torn down, including forwarding subscriptions, per-viewer quality state, and any recording attachments. What survives is the meeting's metadata sitting in a database.

A scheduled meeting is a cheap database state, so Zoom can hold millions of them indefinitely. Media infrastructure is the expensive part, and it only gets allocated to meetings that are actually running.

How Video Gets From One Participant to Everyone Else

In a five-person meeting, your camera feed has to end up on four other screens within a few hundred milliseconds. Every video calling system has to make those copies somewhere. Zoom's answer is a selective forwarding unit, or SFU, though the Zoom version is the MMR above.

An MMR does subscription-based switching with no server-side transcoding or mixing. It reads packet headers, checks who has subscribed to which streams, and copies packets toward them. It never decodes anyone's video, so the expensive codec work stays on the endpoints (independent researchers have confirmed this with packet captures).

The MMR sits in the middle of a meeting:

  1. Your client encodes your camera locally. This doesn't just happen once. Zoom clients can send up to four simultaneous H.264 streams of the same camera at different resolutions, so a full-screen viewer and a thumbnail viewer can be served from different encodings.
  2. Your client sends each stream just once to your assigned MMR. This happens over UDP where the network allows it, falling back to TCP or TLS where it doesn't.
  3. The MMR consults its subscription state and forwards the matching packets to each person who currently has your tile visible, at the resolution their view calls for.
  4. Each receiver decodes only the streams it's actually displaying, not every camera in the meeting.
Diagram showing Alice's client publishing once to the MMR, which forwards subscription-based streams at different resolutions to Bob's and Carol's clients

The result is that your client sends the same amount of video whether two people are watching you or two hundred. The copying still uses plenty of bandwidth, but that load now sits on the MMR in a data center, not on your home connection.

Zoom also chooses which MMR your client connects to. MMRs are grouped into Meeting Zones, zone controllers manage capacity within each zone, and a global controller coordinates across them. When you join, that hierarchy assigns you a nearby server with spare capacity. Participants in different regions can connect to different MMRs, with cascaded links carrying media between them. And a plain two-person call that qualifies skips the MMR entirely and runs peer-to-peer.

One thing to remember, though. Zoom's web client can run over standard WebRTC, but the native apps use a proprietary protocol, and researchers who reverse-engineered the traffic found a custom RTP variant with Zoom-specific headers rather than anything browser-standard. It may be SFU, but it isn't using an open protocol like the one you'd get from real-time video providers.

How Zoom Scales From a 1:1 Call to a Webinar With Thousands of Viewers

At the small end, Zoom doesn't have a scaling problem. A two-person call can run peer-to-peer, and an ordinary meeting runs through an MMR, with the biggest meetings topping out at 1,000 interactive participants with the Large Meeting add-on.

But a webinar with 100,000 viewers is a different problem, and Zoom handles them with two separate concepts:

  1. Adding servers. A large or geographically spread meeting isn't confined to a single MMR. Each participant connects to a nearby server, and the MMRs pass streams to one another over cascaded paths, carrying only the sources that viewers on the other end have subscribed to. No single machine has to handle everyone.
  2. Restricting who can send. This is the bigger move. In a meeting, every participant can share audio, video, and their screen. In a webinar, only the host and panelists are full participants; attendees are view-and-listen-only unless the host deliberately promotes them. However large the audience gets, the media entering the system remains small.
Zoom architecture diagram showing the Global Cloud Controller routing traffic between data centers, zone controllers, and multimedia routers (MMRs)

(Source: Zoom Architecture)

Webinar licenses scale to hundreds of thousands of attendees, with single-use events up to a million attendees. Understandably, Zoom asks organizers to coordinate with it in advance for anything past 200,000. At that scale, a webinar is closer to a broadcast with a small interactive core than to a call.

There is a final, somewhat funny, issue with the size of Zoom calls: what if everyone is in the same building?

Zoom use exploded during COVID and remote work, but when 5,000 employees RTO and watch the same town hall in one place, the building's internet connection carries 5,000 copies of an identical stream. Zoom Mesh handles this within the corporate network. Zoom's cloud picks capable desktop clients as parents, sends them the stream, and has them redistribute it to other clients over the local network. Each parent can serve up to 50 children, and the roles get reassigned during the event as machines slow down or drop off.

Ready to integrate? Our team is standing by to help you. Contact us today and launch tomorrow!

How Zoom Adapts Quality for Each Participant

No two people in a meeting are watching you under the same conditions. One has you full-screen on office fiber, another has you in a thumbnail-sized tile, and a third is on a phone with two bars of signal. A single version of your video can't serve all three, so Zoom runs a set of feedback loops for the whole call.

The client-side piece is called the Reactive QoS layer. It monitors bandwidth, packet loss, latency, and jitter, along with the machine's CPU, memory, and network I/O, and adjusts its encoding to fit. Even the codec is chosen this way, with AV1 used when the device can afford it and H.264 otherwise.

On the receiving side, each app requests streams sized to what it's actually displaying, and the MMR forwards each subscriber whichever version they asked for, the simulcast-style setup from the SFU section.

Trigger Response Where it happens
The CPU is the bottleneck Frame rate drops, Zoom's example being 720p at 30 fps falling to 15 fps Sender
The network is the bottleneck Resolution and bitrate drop instead Sender
A viewer has you in a small tile Their app requests a smaller encoding Receiver
One viewer's connection degrades That viewer gets forwarded a smaller version; everyone else is unaffected MMR
Severe packet loss Video is sacrificed, and audio kept Whole path

This is roughly how the adaptation loops fit together:

Diagram of Zoom's Reactive QoS loop, showing network and device signals feeding the encoder, multiple encodings, MMR forwarding, and reception feedback

The result is that one bad connection stays as one person's problem. The person on the hotel wifi gets a smaller, choppier picture while everyone else keeps the full-quality version, and the meeting never gets dragged down by the worst connection. When conditions get truly bad, video goes first. Zoom prioritizes audio over video and claims sessions stay usable at roughly 45% packet loss.

How Screen Sharing Differs From Camera Video

Camera video and screen content are nearly opposite problems, so Zoom handles them as two different kinds of stream.

Camera video Screen share
Typical content Continuous natural motion Mostly static, with small high-contrast text
What matters most Smooth movement Legible detail
Default priority Frame rate, some blur is fine Resolution, low frame rate is fine
Transport UDP, lost packets skipped or concealed "Reliable UDP", with a setting to force TCP

The capture side reflects the same split. The client exposes five capture modes. Auto picks one for you, and the rest exist for specific situations:

  • The two advanced capture modes add motion detection, which notices when you drag a window or play a movie. They differ only in whether Zoom's own windows appear in the share.
  • Plain capture with window filtering skips motion detection and hides the Zoom windows.
  • Secure share, on Windows, shares only the selected window's content, so nothing in the background leaks through while you move or resize things.
  • Legacy mode exists for older operating systems and video drivers, where the modern capture path can leave viewers looking at a blank screen.
Zoom desktop app Advanced settings screen showing screen capture mode options, including Capture with window filtering

The priorities flip back when the shared item is itself a video. The optimize for video clip option trades sharpness for frame rate, and Zoom warns against leaving it on for ordinary text and slides because they turn fuzzy.

How Host Controls and Breakout Rooms Work

Muting someone, holding them in a waiting room, or moving them into a breakout room. These are all critical elements of video conferencing that have nothing to do with the actual video. They are changes to the meeting state that the control plane makes and the media servers then follow. Zoom defines roles (host, co-host, alternative host, participant) that are really permission sets over that state.

The sequence is:

  1. The host's client sends the command, say "mute Alice," to Zoom's control services.
  2. The control plane checks that the sender's role allows it.
  3. The authoritative meeting state changes.
  4. Alice's client and the media servers are notified, and her audio stops.

Alice gets angry. Unmuting is different because it is a privacy issue. A host can mute anyone at any time, but remotely unmuting someone is a request that the participant must accept, unless they granted standing permission in advance. Host privileges don't extend to switching on another person's microphone.

The waiting room is a holding state before someone actually joins the meeting. A person sitting in it is connected to Zoom just enough to appear on the host's list with a name, but they can't see or hear anything in the meeting itself. This is the state a participant moves through, all managed by the control plane:

State diagram showing a participant moving through JoinRequest, WaitingRoom, MainMeeting, and BreakoutRoom states

Breakout rooms work the same way. They are sessions split off from the main meeting, but it's still one meeting underneath, with the control plane deciding who exchanges audio and video with whom:

  1. When Alice is sent to breakout A, her subscriptions are updated so she can see and hear only the people in that room.
  2. When the host drops in on breakout B, that's the same operation, the host changing their own room membership.
  3. When the host broadcasts a message to all rooms, it gets pushed into each group separately.
  4. When the rooms close, everyone's membership flips back to the main meeting.

Once membership is a control-plane concept, splitting a meeting into rooms is just bookkeeping. The MMRs keep doing the one thing they do, forwarding each person the streams they're subscribed to.

Why Zoom Felt Like Magic

None of the pieces in this article are exotic. A database holding meeting settings, servers that copy packets, clients that adjust their own encoders, and a permission system deciding who can do what.

What made Zoom different was the placement of each piece. A meeting's identity is cheap and permanent, while its media session is expensive and temporary, so the expensive part exists only while people are talking. The servers in the middle forward rather than process, so the same machinery serves a two-person call and a hundred-thousand-viewer webinar. And anything slow, from transcoding a recording to generating a transcript, runs after the call, where nobody is waiting on it.

What feels like magic is good architectural design. A set of decisions about where work happens, applied consistently enough that nobody clicking a meeting link ever has to think about them.

Add Chat to my app: getstream.io/SKILL.md

The fastest way to build with Stream. Start a new project or improve an existing one. Full CLI and documentation integration out of the box.


Ask your agent:

/stream Build me a Social App with Feeds and Moderation.
/stream Any livestream calls running?