Stream's AI Moderation Roadmap: From Detection → Proof

Trust and safety are changing. Our Moderation API is changing with it. Here's a look at what's coming.
Stream's AI Moderation Roadmap: From Detection → Proof

When we first built our Moderation API, trust and safety meant one thing. Catch bad content before it reaches users.

That's still our baseline, but the landscape has changed. Now, automated systems and AI handle most of the catching. The new challenge is proving whether the decision was accurate, and we're tackling this head-on over the next several quarters.

Our roadmap moves through three stages:

  • We're giving teams more control over how their policies work.
  • We're building the reporting and feedback loops that let teams verify accuracy for themselves.
  • We're extending moderation beyond single messages and single accounts, so coordinated patterns become visible.

1. Configurable by Design

Policy Revamp

We're redesigning the policy creation flow with:

  • Engine selection at the policy level, so customers can switch their LLM or NLP engine themselves
  • Policy and configuration versioning, so teams can roll back to a setup that worked better
  • Default policy templates built for specific niches: marketplace, gaming, dating
  • Direct JSON editing for LLM labels and context
  • A redesigned dashboard UI

LLM as a Judge

When a team flags labels as high-stakes, you'll soon see a fallback that routes those decisions through an additional round of model review before they're finalized. You can think of it as a second check on the calls that matter most.

Vertical Policy Packs

We're creating preconfigured policy sets for prediction markets, dating, and social platforms, applied from the dashboard in one click. Teams in these categories will start with a policy set built for their vertical rather than from scratch.

2. Proving It Works

Model Performance Reporting

We're building per-customer reporting that breaks down:

  • Precision, recall, and false-positive rates
  • By policy, language, and harm category
  • With deltas tracked before and after tuning

Instead of a single accuracy number, teams get visibility into where a model performs well and where it doesn't, specific to their own content.

Model Feedback Automation

Soon, teams will be able to flag false positives and false negatives directly from the Stream dashboard, with a dedicated screen tracking each item as open or resolved. Teams will no longer need multiple support tickets and follow-up emails to confirm a model mistake was fixed.

Workforce Tooling

As moderation teams grow, especially outsourced teams handling high volumes, the tooling supporting the humans in the loop should scale with them. For our product, that means:

  • Blind QA sampling
  • A moderator performance API
  • Configurable escalation limits per agent
  • Cross-app rollup reporting

We see this as foundational for any team running moderation at meaningful headcount.

The Stream Moderation Benchmark

We're publishing a public, versioned accuracy report:

  • Per-class precision and recall at fixed false-positive rates
  • Benchmarked against named baselines
  • Published methodology
  • Re-run quarterly

Anyone evaluating us, whether they're a current customer or comparing us to other options, should have access to real numbers, a documented process for how we arrived at them, and a fresh version every quarter.

3. Reasoning Beyond the Message

Account Linking and Coordinated Abuse

Some of the most damaging content shows up as a pattern across accounts rather than a single bad message. This feature set will cover:

  • Cross-account correlation
  • Linked-account detection
  • Templated-message clustering
  • Campaign bundling in the review queue

A moderator will see the campaign instead of fifty disconnected reports.

We're replacing static allowlist maintenance with:

  • URL reputation scoring
  • Redirect-chain resolution
  • Newly-registered-domain scoring
  • QR code extraction

This adapts as scam tactics change, rather than relying on lists that fall out of date.

Context-Aware Enforcement

Not every violation deserves the same response, and the same content doesn't always mean the same thing depending on when it appears in a conversation. For example, a flagged image as the first message reads differently than the same image fifteen messages into an established exchange.

We're building stateful, contextual severity:

  • Tiered strikes that decay over time
  • Severity shaped by conversation context, not content in isolation

This has come up repeatedly with social and gaming communities, where message-by-message enforcement misses the pattern.

Where This Is Headed

We're watching closely as moderation shifts from a single decision made in isolation to an ongoing, provable, and increasingly connected system.

Configurable policies give teams control over how decisions get made. Performance reporting and the public benchmark give them a way to verify those decisions are correct. Cross-account intelligence and contextual enforcement extend the picture beyond any single message.

That's the direction we're building toward.

We'll keep you looped in as it ships.

Interested in AI Moderation or have questions about your existing setup?

Contact Moderation GTM Lead Kenzie Wilson and you will receive a reply shortly.

Add Chat to my app: getstream.io/SKILL.md

The fastest way to build with Stream. Start a new project or improve an existing one. Full CLI and documentation integration out of the box.


Ask your agent:

/stream Build me a Social App with Feeds and Moderation.
/stream Any livestream calls running?