Scaling Moderation to 1,500+ Requests per Second: Stream's Latest Benchmarks

Scaling Moderation to 1,500+ Requests per Second: Stream's Latest Benchmarks

Stream AI Moderation checks text, images, audio, and video through one API, and its performance is measured continuously on live production traffic. Over seven days of that traffic, the API handled more than 1,500 text classification requests per second at the daily peak, with API latency at about 50 ms or lower. English text typically completes in about 35 ms.

Stream Moderation Benchmarks at a Glance

These numbers come from production. The policies and models are the ones customers run, so the results include real model time.

Text classification load and latency over seven days of production traffic

Key results include:

  • About 35 ms p50 for text in most languages, including English, on the Chat and Feeds send path (a few languages, such as Arabic and Turkish, take longer)
  • 1,500+ text classification requests per second at peak, with API latency at about 50 ms or lower
  • Hundreds of image and live-video requests per second, with API latency of about 40 to 50 ms
  • About 300 ms to score an image or a live-video frame
  • Seven days of production traffic measured

API latency rises slightly as traffic builds toward the daily peak and stays at about 50 ms or lower.

Image and live-video load and latency over seven days of production traffic

Why We Measure Moderation Latency

Moderation runs in the path of the content it checks. When someone sends a chat message, the check has to finish quickly, or the message is delayed for the sender and for everyone else in the conversation. API latency is therefore the number that users experience directly.

Latency also differs by content type. Text can be checked in tens of milliseconds. Scoring an image takes about 300 ms. Video files are scored as full files, and audio files are transcribed before they are moderated, so both take longer and cannot run inline.

For this reason, we report latency separately for each type of content.

One Policy, Three Endpoints, Five Content Types

Stream Moderation applies one policy across every content type. You can use it on its own with any app or platform that moderates user-generated content, or add it to Stream Chat, Video, or Feeds.

Stream's own integrations use the check endpoint. Your own backend can use the same endpoint, or use labels for high-volume text and analyze images and live video frames.

Content API How it runs
Text check or labels English text on the send path typically completes in about 35 ms.
Images analyze or check The API responds in about 40 ms. Scoring a photo takes about 300 ms.
Live video analyze Your app samples keyframes and sends them. The image model scores them. You decide how often to sample.
Video files check The API accepts the upload, and scoring runs in the background.
Audio files check The file is transcribed and then run through the same text engines. analyze does not accept audio.
  • check: Accepts text, images, audio files, and video files. Returns the action set on the policy, including keep, flag, remove, bounce, shadow, and mask. Chat and Feeds call it on the send path.
  • labels: Accepts high-volume text, such as game chat. Returns labels without creating a review item for every message.
  • analyze: Accepts images and live-video keyframes, with optional text. Your app samples keyframes from a live stream and sends them. You decide how often.

Enterprise-Ready by Design

Fast responses depend on how the API is built.

  • Edge network: Requests go to the nearest edge server. TLS, authentication, and rate limiting happen there. Edge servers support HTTP/2.
  • Caching: Blocklists are cached in memory and in Redis. Policies are cached in Redis, so each request does not reload them from the database.
  • Resilience: Circuit breakers, retries, consistent-hash routing, and locality-aware load balancing spread requests across availability zones and shift traffic away from a zone that fails its health checks.
  • Infrastructure: The Go backend runs on AWS and GCP. Integration tests, production smoke tests, and QA run before each deploy.

Enterprise customers can add a 99.999% uptime SLA.

Published Benchmarks, Measured on Production Traffic

Performance numbers mean the most when they come from real traffic and real policies. Stream publishes them, along with the underlying architecture, so engineering and product teams can evaluate Moderation before building on it.

For model quality, including precision and recall across content categories, see the AI Moderation Benchmark Report. If your app handles user-generated text, images, audio, or video, try Stream Moderation.