# Architecture & Benchmark

Stream's Moderation API is designed to moderate text, images, audio, and video at high volume while maintaining low latency. Here we explain the architecture behind the API and share benchmarks from production traffic that show how it performs under load.

## Architecture overview

The Moderation API integrates with Stream Chat, Video, and Feeds and also serves as a standalone product for any app or platform that supports user-generated content.

A few key things about the architecture:

- **Edge network:** Your backend connects to the nearest Stream edge server. TLS termination, authentication, and rate limiting happen at the edge.
- **Text, images, audio, and video:** One API and one policy cover all content types, so you do not need a separate vendor for each.
- **Excellent performance:** English text typically completes in about 35 ms. High-volume text classification handles 1,500+ requests per second with API latency under 50 ms. Image and live video API calls remain within the same latency range.
- **Highly redundant infrastructure:** Moderation runs on AWS and GCP, with failover across availability zones.

### Benchmark at a glance

Stream Moderation is measured continuously on production traffic. As you can see in the chart, text classification reaches over 1,500 requests per second at the daily peak while API latency stays at about 50 ms or lower. Latency rises slightly as load increases because more text needs translation at peak times. It returns to about 35 ms overnight. There is no degradation as traffic moves through the day.

![Text classification load and latency](https://getstream.io/docs-assets/images/f815df09acb9.png)

Image and live-video traffic show the same pattern in the API: hundreds of requests per second, with latency between about 40 ms and 50 ms. Scoring a frame takes about 300 ms, so live-stream apps send frames without waiting for the model.

![Image and live-video load and latency](https://getstream.io/docs-assets/images/0d0765fc40e6.png)

## APIs and content types

One policy applies across three APIs and five content types.

| API                               | What you send                                                                                                     |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| `POST /api/v2/moderation/check`   | Text, images, audio files, and video files from Chat, Feeds, or your own entities. Returns keep, flag, or remove. |
| `POST /api/v2/moderation/labels`  | High-volume text. Returns labels without creating a review item for every message.                                |
| `POST /api/v2/moderation/analyze` | Images and live-video keyframes, with optional text.                                                              |

| Content         | How it runs                                                                                                                                            |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Text**        | Use check or labels. English text on the send path typically completes in about 35 ms.                                                                 |
| **Images**      | Use analyze or check. The API responds in about 40 ms. Scoring a photo takes about 300 ms.                                                             |
| **Live video**  | Your app samples keyframes and sends them to analyze. The same image model that scores still images scores the frames. You decide how often to sample. |
| **Video files** | Use check. The API accepts the upload, and scoring runs in the background.                                                                             |
| **Audio files** | Use check. The file is transcribed and then run through the same text engines. analyze does not accept audio.                                          |

## The tech behind the Stream Moderation API

Stream uses Go as the backend language for moderation.

### Edge network

When you call the API, the request goes to the nearest edge server. This is similar to how a CDN works, with more processing at the edge. Edge servers support HTTP/3 and HTTP/2. TLS terminates at the edge. Traffic between the edge and the origin servers is encrypted and uses long-lived HTTP/2 connections. Authentication, rate limiting, and CORS handling also run at the edge.

For resilience, the edge infrastructure uses circuit breakers, request retries, consistent-hash routing, and locality-aware load balancing with failover to another availability zone when needed.

### API performance

English text typically completes in about 35 ms, with a p95 of about 70 ms. Stream achieves this by caching policies and blocklists in memory and in Redis, and by keeping its own processing to a few milliseconds around the model call.

Each API fits a different workload. Chat and Feeds call check on the send path. Games and other high-volume text apps call labels. Images and live-video keyframes go to analyze. Uploaded video and audio files stay on check and finish in the background.

### Infra & testing

A large Go integration test suite, production smoke tests, and QA run before each deploy. The infrastructure runs on AWS and GCP and is highly available. Enterprise customers can add a 99.999% uptime SLA.

## Benchmark

We measure live production traffic, which includes the policies and models that customers actually run. The numbers therefore include real model time and do not come from a synthetic replay that hides it. The charts on this page cover a global snapshot of seven days of production traffic from all regions, including Europe, US East, and US West.

English text on the Chat and Feeds send path has a p50 of about 35 ms, a p95 of about 70 ms, and a p99 of about 100 ms.

![English text latency percentiles](https://getstream.io/docs-assets/images/623f3f78d3c7.png)

### check: text (English)

|                        | p50    | p95    | p99     |
| ---------------------- | ------ | ------ | ------- |
| Chat / Feeds send path | ~35 ms | ~70 ms | ~100 ms |

### labels: high-volume text

| Load         | p50     |
| ------------ | ------- |
| 1,500+ req/s | < 50 ms |

### analyze: images and live-video keyframes

| Load                       | API p50 | Time to score a frame |
| -------------------------- | ------- | --------------------- |
| Hundreds / sec (peak ~700) | ~40 ms  | ~300 ms               |

### check: video files and audio files

Both file types are accepted by the request and scored in the background. Video files go through a full-file job. Audio files are transcribed and then moderated as text. These paths are not inline, so we do not publish a millisecond p50 for them.

**Key takeaways:**

- **check** (English text): about 35 ms p50, 70 ms p95, 100 ms p99.
- **labels**: 1,500+ requests per second at under 50 ms p50.
- **analyze**: API under 50 ms, and a frame score in about 300 ms.
- **Video files and audio files**: processed in the background on **check**.
- API latency stays at about 50 ms or lower through the daily traffic peak.

For model quality (precision and recall across content categories), see the [AI Moderation Benchmark Report](https://getstream.io/assets/stream-moderation-benchmark-17074005.pdf). This page is how the API is built and how it behaves under load.

---

For the most recent version of this documentation, visit [https://getstream.io/moderation/docs/ruby/architecture-and-benchmark/](https://getstream.io/moderation/docs/ruby/architecture-and-benchmark/).