Tag: Content Moderation

Instagram shrinks reach of profiles hiding AI faces

Instagram will limit the reach of profiles that hide AI-generated people.

Anthropic’s older Claude models slip past its own content rules

TechCrunch testing shows Anthropic's older Claude models readily bypass their own content restrictions.

Nine months of AI abuse ads slipped through Meta’s ad review

A watchdog found more than 50 paid ads with AI-generated child abuse imagery on Meta platforms.

Mistral opens 3B Shieldstral model for adaptive content safety

Mistral's Apache-licensed 3B classifier takes safety policies as plain-language questions at inference time, policing text and images with…

Reddit Turns to LLMs to Combat the AI Spam Problem

Reddit developed LLM-powered moderation tools to fight AI-generated spam, blocking 23 million spam views per day in an…

The Clippening Is Faking Sex Scenes with Actors’ Faces, and Nobody Cares

Micro drama studios are using generative AI to paste actors' faces onto explicit body doubles for fake sex…