How Do AI Image Moderation Filters Work?
TL;DR
- AI image moderation filters normalize, hash, classify, and OCR each upload, then compare the scores against per-category thresholds to allow, flag, or block it.
- The classifier gives a confidence score. Calibrated thresholds decide what action fires for each category.
- Context is what these models get wrong most. The same pixels can be news, medicine, art, or a violation depending on where they came from.
- Human reviewers handle the middle band, high-severity content, and appeals, and their overturns become the next round of training data.
One of the worst things that can happen to an app is a user uploading a photo nobody should ever see, and it just sits there in the feed until somebody reports it. Nudity, gore, someone's driver's license, a meme with a slur across it.
Nobody can moderate images by hand. Too much user-generated content comes in, and people expect their photo to show up the second they hit send. So a machine has to look first.
That's what an AI image moderation filter is. AI image moderation is a model that looks at each image and scores how likely it is to contain explicit content, violence, a hate symbol, or other inappropriate content. You decide what score is high enough to block, what's worth sending to a content moderator, and what to let through.
What Happens to an Image When It Reaches a Moderation Filter?
It gets normalized, hashed against a database of known illegal images, scored by a multi-label computer vision classifier, OCR'd for any text, and then a policy layer compares the scores to your thresholds and picks an action. Everything up to the classifier is cheap enough to run on every upload. The expensive model, if there is one, only sees the images the classifier couldn't settle.
Say someone posts a photo from a boozy night out. Here's what happens to it:
- Normalize. The file is decoded, oriented, and resized into one standard copy that every later stage works from. Broken files and oversized ones that exist only to crash the decoder get rejected here.
- Hash matching. The normalized copy gets a perceptual hash, which is compared against hashes of previously identified illegal images, mostly child sexual abuse material cataloged by NCMEC. PhotoDNA is the standard one. A perceptual hash survives resizing, recompression, and small edits, unlike a cryptographic hash, where one changed pixel breaks the match, but it only finds images that are already in the database.
- Classification. This is the artificial intelligence part. A computer vision model fine-tuned on labeled moderation examples gives one score per category. It's multi-label, because an image can be nude, violent, and contain a hate symbol at the same time. On Stream, the image engine returns a confidence score from 1 to 100 for each of ten categories, covering explicit and adult content, weapons and gore, drugs, alcohol, gambling, and rude gestures, so that the night-out photo might come back with alcohol at 92, swimwear at 12, and everything else near zero.
- OCR. Any text in the image, such as a caption, a chat screenshot, or the overlay on a meme, gets pulled out with optical character recognition and run through the text moderation pipeline like any other message, so profanity, slurs, and threats get caught the same way they would in chat. A classifier trained on visual content can't read, and without text detection, a screenshot of a threat is just a picture.
- Policy layer. Each score is compared against the threshold you set for that category, and the action attached to that threshold fires. On Stream, that's allow, flag into the review queue, or block. Alcohol at 92 in a bar-finder app is nothing. The same score in an app for 13-year-olds is a block, and the model never needs to know the difference. Your content policies live outside the model, so you can move a threshold in the dashboard without retraining anything.
The big platforms add a stage for the images the classifier left in the middle. A vision-language model gets the image, the OCR text, and the surrounding post, and is asked the actual policy question, something like "is this nudity medical or sexual?" It's slow and expensive per call, so it only ever sees the leftovers.

On Stream, the classifier and OCR run asynchronously. The message is accepted immediately, the image is scored in the background, and if a score crosses a threshold, the action is applied to the message afterward. The upload path never blocks on inference, so the user experience is the same as with no moderation at all. Video moderation works the same way on keyframes, and for live streams the keyframes are scored while the call is running, with actions that escalate from a warning to a blur to a kick.
How Are Image Moderation Models Trained?
The classifier starts from a general-purpose vision model pretrained on hundreds of millions of web images. Then it gets fine-tuned on images that human reviewers have labeled against a specific policy. Pretraining gives it a broad sense of what things look like. Fine-tuning is where the policy comes in.
Take a CLIP-style model, which is where many image safety classifiers start. CLIP was trained on 400 million image-and-caption pairs pulled from the web, with one task: figuring out which caption goes with which image. After that, it has a decent idea of what almost anything looks like, including plenty of things nobody ever labeled for moderation, and it can score an image against a text description with no further training.
Fine-tuning turns that into a moderation model. This involves:
- Labeled examples. Human moderators mark real user-generated images against the platform's content policies, and the model trains on those decisions. Most of the set is images marked as fine because that's what real traffic looks like, so the loss gets weighted toward the rare categories. Otherwise, the model learns to call everything safe and still scores well.
- One output per category. The classification head is multi-label, so each category gets its own independent score, and the model is never forced to pick one.
- AI-generated images. Classifiers trained only on photographs do worse on generated images, and generated images are now a normal part of what people upload, so they go into the training set on purpose.

(Source: OpenAI)
For a brand-new category, say a hate symbol that showed up last month, a CLIP-style model can score a text description of it with zero labeled examples. That's good enough to get a rough detector running the same day. Then you collect labels and fine-tune properly, because similarity to a text prompt isn't a calibrated probability that your policy was broken.
The labeling is the hard part. "Nudity" on its own is a useless instruction to a human moderator. They need rules for medical images, breastfeeding, classical art, cartoons, swimwear, and what to do when they can't tell someone's age. When two reviewers disagree, the model learns the disagreement, and borderline images get borderline scores.
How Does a Confidence Score Turn Into a Decision?
The policy layer compares the score against a threshold you set per category, and the threshold decides the action. You end up with a threshold for each category, and often one for each action within a category, so the score that earns a warning label is lower than the score that gets an image blocked.
The usual pattern is three bands per category:
- Below the low threshold, the image is allowed, and the score is just logged.
- Between the two thresholds, it goes to a moderator, or gets friction like a blur or a warning label while it waits.
- Above the high threshold, the configured action fires with nobody looking.
Meta runs the same split. Its AI models handle clear cases, and the rest is escalated and ranked for reviewers by severity, virality, and likelihood of violating.
The bands should move as needed. A warning label can eat a lot of false positives; an account ban can't, so the block threshold sits well above the flag threshold, and if you automate bans at all, that threshold sits higher still.
They move with the audience too. A marketplace scanning product listings can auto-block anything that scores 70 on drugs. A harm-reduction forum can't, because drug paraphernalia is what people are there to talk about.
On Stream, the threshold for each category runs from 1 to 100, and the recommended starting points are:
| Threshold | When to use it |
|---|---|
| 80 to 90 | Categories you enforce strictly, where a false positive is expensive |
| 60 to 70 | Balanced detection |
| 40 to 50 | Maximum sensitivity, where over-flagging is fine because a person reviews everything flagged |
A raw classifier output isn't a probability, and neural networks run overconfident, so a 95 from the model doesn't mean it's right 95 times in 100 until someone has calibrated it against held-out real traffic. Serious deployments do that before setting any threshold.
They also redo it whenever the model changes, because a threshold tuned on last quarter's model doesn't mean the same thing on this quarter's. Day to day, that means your trust and safety team watching the review queue and moving thresholds when the overturn rate says to.
What Do AI Image Moderation Filters Get Wrong?
Context, more than anything. A classifier scores pixels, and the same pixels can be news, medicine, art, or inappropriate content depending on where the image came from and what's written next to it. After context, the main problems are edits that push a score below the threshold, images that don't look like the training set, and error rates that change depending on who's in the picture.
The ones you'll see in practice:
- Context. A gory image can be war reporting, a surgical photo, or something posted to glorify violence, and the pixels don't say which. A photo of weapons can be a hunting store's stock or a threat. A swastika in a museum photograph and one in propaganda look the same to a classifier. Google's SafeSearch returns a separate "medical" likelihood next to "adult" partly for this reason, and medical and artistic nudity still throw false positives at a rate you'll notice.
- Memes. A harmless photo with a hateful caption is hateful content, and social media is full of them. OCR gets the caption into text moderation, but the harm is in the pairing, and neither the image model nor the text model sees both halves together.
- Simple edits. Cropping, rotating, adding a border, recompressing, or pasting the prohibited part into a corner of a bigger image can all drop the score under the threshold. Perceptual hashes hold up to resizing, but not to someone deliberately probing for a version that slips past.
- Distribution shift. Every new phone camera, filter trend, or image generator changes what uploads look like, and the AI models were trained on last year's uploads. A filter can hold its benchmark score while its recall on live traffic drops.
- Uneven error rates. Aggregate accuracy hides how the filter does on particular groups. Skin tone, clothing norms, religious dress, and regional symbols all move the error rate because machine learning algorithms learn whatever the training data contains, and a filter trained mostly on one market's data over-flags or under-flags another's.

The fix is redundancy. Run several AI models with different weaknesses, score transformed copies of high-risk uploads, keep a review band instead of a single cutoff, and send a random sample of images the model passed to reviewers anyway. Without that last one, the only images a person ever looks at are the ones the model already caught, and your false negatives stay invisible.
Where Do Human Reviewers Fit In?
Human moderators take the images the model couldn't settle, and their decisions become the labels the next version of the model is trained on. The filter's job is to keep the obvious cases at both ends away from them so that they can focus on the middle band, high-severity content, and appeals.
Meta says every decision its reviewers make goes back into training its systems, which is how the automated share of its enforcement grows over time. The same content moderation loop works at any size. If your content moderators overturn the model on the same kind of image often enough, that's your fine-tuning data.
On Stream, an image that crosses a flag threshold lands in the Media Queue in the moderation dashboard with its label and score. A moderator confirms or overturns the call and takes an action, like deleting the message or banning the user.
Reviewers also do two things the model can't:
- They resolve context. A person can tell whether a nude image is medical, whether a violent one is news, and whether a symbol is being condemned or promoted.
- They handle appeals. A user contesting a removal is the best source of false positives you'll get, and a climbing overturn rate on one category tells you that threshold or that model has a problem well before a benchmark would.
Known illegal material follows its own path. In the US, providers that become aware of apparent child sexual abuse material on their service have to report it to NCMEC under federal law, and other countries have their own rules. Hash matches for that material go to a specialized trust and safety workflow with tighter access controls, never into the general queue next to the swimwear flags.
If you're building on Stream, the thresholds, the Media Queue, and appeals all live in the same dashboard as the automated image moderation itself, so you configure the whole system in one place instead of assembling it from pieces.
